Recent studies have shown that tuning prediction models increases prediction accuracy and that Random Forest can be used to construct prediction intervals. However, to our best knowledge, no study has investigated the need to, and the manner in which one can, tune Random Forest for optimizing prediction intervals { thi…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study on random matrices in deep neural networks with IID entries.
Random projections simplify complex data for classification.
Enhances uncertainty modeling in random PDEs using PINNs and generative models.
New method uses random convex polytopes to measure representation quality.
New model calculates logarithmic surface diameter.
This work employs some techniques in order to filter random noise from the information provided by minimum spanning trees obtained from the correlation matrices of international stock market indices prior to and during times of crisis. The first technique establishes a threshold above which connections are considered a…
Paper improves neural network robustness certification with tighter radii estimates.
Random walk speed on Teichmüller space is a proper function.
New algorithms improve combinatorial linear semi-bandits for clustered feature vectors.
The paper examines how random seed affects model stability and proposes ASWA and NASWA techniques to improve model robustness.
Randomized neural networks improve deep RL agents' generalization.
We consider the binary classification problem when data are large and subject to unknown but bounded uncertainties. We address the problem by formulating the nonlinear support vector machine training problem with robust optimization. To do so, we analyze and propose two bounding schemes for uncertainties associated to …
Study on random matrices in deep neural networks using Gaussian data.
An important problem in training deep networks with high capacity is to ensure that the trained network works well when presented with new inputs outside the training dataset. Dropout is an effective regularization technique to boost the network generalization in which a random subset of the elements of the given data …
ORCCA improves CCA performance with randomized features.
Graph connection Laplacian (GCL) is a modern data analysis technique that is starting to be applied for the analysis of high dimensional and massive datasets. Motivated by this technique, we study matrices that are akin to the ones appearing in the null case of GCL, i.e the case where there is no structure in the datas…
We study a maturity randomization technique for approximating optimal control problems. The algorithm is based on a sequence of control problems with random terminal horizon which converges to the original one. This is a generalization of the so-called Canadization procedure suggested by Carr [Review of Financial Studi…
As a typical dimensionality reduction technique, random projection can be simply implemented with linear projection, while maintaining the pairwise distances of high-dimensional data with high probability. Considering this technique is mainly exploited for the task of classification, this paper is developed to study th…
In recent studies, the generalization properties for distributed learning and random features assumed the existence of the target concept over the hypothesis space. However, this strict condition is not applicable to the more common non-attainable case. In this paper, using refined proof techniques, we first extend the…
This paper investigates the theory of robustness against adversarial attacks. It focuses on the family of randomization techniques that consist in injecting noise in the network at inference time. These techniques have proven effective in many contexts, but lack theoretical arguments. We close this gap by presenting a …
Recent theoretical work has identified random projection as a promising dimensionality reduction technique for learning mixtures of Gausians. Here we summarize these results and illustrate them by a wide variety of experiments on synthetic and real data.
The performance of classification algorithms with a massive and highly imbalanced data stream depends upon efficient balancing strategy. Some techniques of balancing strategy have been applied in the past with Batch data to resolve the class imbalance problem. This paper proposes a new incremental data balancing framew…
Study compares classification techniques to predict customer churn in banking.
New ensemble SVM model reduces prediction error without choosing best kernel.
Study random walks on groups with superlinear divergent geodesics.
Gradient estimation techniques applied to programs with randomness in high energy physics.
Bounds on Gaussian approximation for neural networks with novel smoothing techniques.
Random forest (RF) methodology is one of the most popular machine learning techniques for prediction problems. In this article, we discuss some cases where random forests may suffer and propose a novel generalized RF method, namely regression-enhanced random forests (RERFs), that can improve on RFs by borrowing the str…
Orthogonal random features approximate a Bessel kernel, offering sharper bounds than random Fourier features.
Extends random feature analysis to spectral methods and improves learning rates.
The paper presents a practical method for evaluating investment projects using real options.
Random feature maps are ubiquitous in modern statistical machine learning, where they generalize random projections by means of powerful, yet often difficult to analyze nonlinear operators. In this paper, we leverage the "concentration" phenomenon induced by random matrix theory to perform a spectral analysis on the Gr…
Randomized algorithm solves vector-valued regression problems with low-rank operators.
This study evaluates data pre-processing techniques for class imbalance in biomedical data.
ECV method optimizes ensemble parameters for randomized ensembles.
Microarray cancer gene expression data comprise of very high dimensions. Reducing the dimensions helps in improving the overall analysis and classification performance. We propose two hybrid techniques, Biogeography - based Optimization - Random Forests (BBO - RF) and BBO - SVM (Support Vector Machines) with gene ranki…
Lecture notes on kernel functions and Random Fourier Features.
UniNet efficiently learns network representations from large graphs.
Method improves treatment effect estimation in randomized experiments.
XGBoost outperforms other boosting techniques in training speed and generalization performance.
Project classifies Hinglish social content on platforms like Twitter, Reddit.
Missing data is an expected issue when large amounts of data is collected, and several imputation techniques have been proposed to tackle this problem. Beneath classical approaches such as MICE, the application of Machine Learning techniques is tempting. Here, the recently proposed missForest imputation method has show…
We propose a non-parametric regression methodology, Random Forests on Distance Matrices (RFDM), for detecting genetic variants associated to quantitative phenotypes representing the human brain's structure or function, and obtained using neuroimaging techniques. RFDM, which is an extension of decision forests, requires…
Develops precise expressions for random projections for better machine learning tasks.
Recent breakthroughs in the field of deep learning have led to advancements in a broad spectrum of tasks in computer vision, audio processing, natural language processing and other areas. In most instances where these tasks are deployed in real-world scenarios, the models used in them have been shown to be susceptible …
New technique improves time series forecasting with less data.
Random Forest variable importance is improved by class balancing techniques.