The paper analyzes LOCV for high-dimensional risk estimation, proving error bounds.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New method for estimating out-of-sample R² from gene expression data.
Estimates error for robust M-estimators with convex penalties.
Paper introduces DOO models to outperform SAA out-of-sample.
Improves test set performance and reduces out-of-sample disappointment for unstable models.
We identify and validate a model for PCR in high dimensions, improving prediction guarantees.
Optimal number of voters for a voting ensemble can be estimated from the distribution of classifier errors.
Paper presents a new way to analyze machine learning generalization without probabilistic assumptions.
Many popular dimensionality reduction procedures have out-of-sample extensions, which allow a practitioner to apply a learned embedding to observations not seen in the initial training sample. In this work, we consider the problem of obtaining an out-of-sample extension for the adjacency spectral embedding, a procedure…
Study improves prediction accuracy and uncertainty for mobile sensor data using randomized neural networks.
Dimensionality reduction methods are very common in the field of high dimensional data analysis. Typically, algorithms for dimensionality reduction are computationally expensive. Therefore, their applications for the analysis of massive amounts of data are impractical. For example, repeated computations due to accumula…
When the in-sample Sharpe ratio is obtained by optimizing over a k-dimensional parameter space, it is a biased estimator for what can be expected on unseen data (out-of-sample). We derive (1) an unbiased estimator adjusting for both sources of bias: noise fit and estimation error. We then show (2) how to use the adjust…
Paper studies M-estimators with derivatives and residual distribution for robust adaptive tuning.
Fast, reliable, and error-bounded option pricing with neural networks
Nonlinear kernels can be approximated using finite-dimensional feature maps for efficient risk minimization. Due to the inherent trade-off between the dimension of the (mapped) feature space and the approximation accuracy, the key problem is to identify promising (explicit) features leading to a satisfactory out-of-sam…
New bounds for DeepONets reduce out-of-sample error without width dependence.
Proposes a new framework to optimize portfolios with reduced estimation errors.
Diffusion maps are a nonlinear manifold learning technique based on harmonic analysis of a diffusion process over the data. Out-of-sample extensions with computational complexity , where is the number of points comprising the manifold, frustrate applications to online learning applications requiring…
Sparse modeling improves portfolio optimization by reducing errors in complex market systems.
Random feature mapping (RFM) is a popular method for speeding up kernel methods at the cost of losing a little accuracy. We study kernel ridge regression with random feature mapping (RFM-KRR) and establish novel out-of-sample error upper and lower bounds. While out-of-sample bounds for RFM-KRR have been established by …
Under the framework of spectral clustering, the key of subspace clustering is building a similarity graph which describes the neighborhood relations among data points. Some recent works build the graph using sparse, low-rank, and -norm-based representation, and have achieved state-of-the-art performance. Howeve…
A new deep learning model improves asset pricing predictions.
The paper tackles stock prediction models by improving their generalizability to out-of-sample domains using causal representation learning.
Paper proposes a method to improve prediction intervals for neural networks.
The paper improves ALO for -regularized models.
A new framework for time series forecasting that adapts to varying patterns.
Let be a data set in , where is the training set and is the test one. Many unsupervised learning algorithms based on kernel methods have been developed to provide dimensionality reduction (DR) embedding for a given training set $Φ: \mathbf{X} \to \mat…
This work gives a simultaneous analysis of both the ordinary least squares estimator and the ridge regression estimator in the random design setting under mild assumptions on the covariate/response distributions. In particular, the analysis provides sharp results on the ``out-of-sample'' prediction error, as opposed to…
Variant of mSSA improves time series prediction error.
A new hyperparameter optimization method reduces overfitting.
We study the out-of-sample properties of robust empirical optimization problems with smooth -divergence penalties and smooth concave objective functions, and develop a theory for data-driven calibration of the non-negative "robustness parameter" that controls the size of the deviations from the nominal model. Bu…
Robustifies Markowitz portfolios to reduce transaction costs and improve performance.
The paper optimizes asset selection for index trackers and enhanced trackers with varying cardinality constraints.
This paper presents an alternative approach to p-values in regression settings. This approach, whose origins can be traced to machine learning, is based on the leave-one-out bootstrap for prediction error. In machine learning this is called the out-of-bag (OOB) error. To obtain the OOB error for a model, one draws a bo…
The only input to attain the portfolio weights of global minimum variance portfolio (GMVP) is the covariance matrix of returns of assets being considered for investment. Since the population covariance matrix is not known, investors use historical data to estimate it. Even though sample covariance matrix is an unbiased…
Manifold learning has been successfully applied to a variety of medical imaging problems. Its use in real-time applications requires fast projection onto the low-dimensional space. To this end, out-of-sample extensions are applied by constructing an interpolation function that maps from the input space to the low-dimen…
This paper presents an out-of-sample prediction comparison between major machine learning models and the structural econometric model. Over the past decade, machine learning has established itself as a powerful tool in many prediction applications, but this approach is still not widely adopted in empirical economic stu…
We study model evaluation and model selection from the perspective of generalization ability (GA): the ability of a model to predict outcomes in new samples from the same population. We believe that GA is one way formally to address concerns about the external validity of a model. The GA of a model estimated on a sampl…
The paper analyzes prediction error in nonstationary settings using weighted risk minimization.
Alpha-based performance evaluation may fail to capture correlated residuals due to model errors. This paper proposes using the Generalized Information Ratio (GIR) to measure performance under misspecified benchmarks. Motivated by the theoretical link between abnormal returns and residual covariance matrix, GIR is deriv…
Paper analyzes high-dimensional portfolio risks and finds empirical out-of-sample relative loss is more reliable.
In the past decade many researchers have proposed new optimal portfolio selection strategies to show that sophisticated diversification can outperform the naïve 1/N strategy in out-of-sample benchmarks. Providing an updated review of these models since DeMiguel et al. (2009b), I test sixteen strategies across six empir…
This paper proposes an out-of-sample extension framework for a global manifold learning algorithm (Isomap) that uses temporal information in out-of-sample points in order to make the embedding more robust to noise and artifacts. Given a set of noise-free training data and its embedding, the proposed framework extends t…
We consider the problem of vertex classification for graphs constructed from the latent position model. It was shown previously that the approach of embedding the graphs into some Euclidean space followed by classification in that space can yields a universally consistent vertex classifier. However, a major technical d…
Study improves understanding of non-differentiable penalties in high-dimensional settings.
Graph embeddings, a class of dimensionality reduction techniques designed for relational data, have proven useful in exploring and modeling network structure. Most dimensionality reduction methods allow out-of-sample extensions, by which an embedding can be applied to observations not present in the training set. Appli…
The paper proposes a new model for predicting and analyzing economic variables.
CASTLE learns causal DAG to improve model generalization.