Paper analyzes high-dimensional portfolio risks and finds empirical out-of-sample relative loss is more reliable.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper proposes an out-of-sample extension framework for a global manifold learning algorithm (Isomap) that uses temporal information in out-of-sample points in order to make the embedding more robust to noise and artifacts. Given a set of noise-free training data and its embedding, the proposed framework extends t…
We consider the problem of vertex classification for graphs constructed from the latent position model. It was shown previously that the approach of embedding the graphs into some Euclidean space followed by classification in that space can yields a universally consistent vertex classifier. However, a major technical d…
New method for estimating out-of-sample R² from gene expression data.
Graph embeddings, a class of dimensionality reduction techniques designed for relational data, have proven useful in exploring and modeling network structure. Most dimensionality reduction methods allow out-of-sample extensions, by which an embedding can be applied to observations not present in the training set. Appli…
Paper introduces DOO models to outperform SAA out-of-sample.
Many popular dimensionality reduction procedures have out-of-sample extensions, which allow a practitioner to apply a learned embedding to observations not seen in the initial training sample. In this work, we consider the problem of obtaining an out-of-sample extension for the adjacency spectral embedding, a procedure…
OTSL improves structure learning accuracy with out-of-sample and resampling strategies.
Dimensionality reduction methods are very common in the field of high dimensional data analysis. Typically, algorithms for dimensionality reduction are computationally expensive. Therefore, their applications for the analysis of massive amounts of data are impractical. For example, repeated computations due to accumula…
Paper proposes a diagnostic tool for evaluating model performance out-of-sample.
We identify and validate a model for PCR in high dimensions, improving prediction guarantees.
Study the impact of overfitting on linear predictive models' performance.
New framework for learning policies that converge in out-of-sample regions.
Optimizes decisions without knowing the true distribution using historical data.
Paper develops a method to predict spatial point processes with guarantees.
Bayesian neural networks show good correlation between out-of-sample performance and Bayesian evidence.
Improves test set performance and reduces out-of-sample disappointment for unstable models.
Paper presents a new way to analyze machine learning generalization without probabilistic assumptions.
Under the framework of spectral clustering, the key of subspace clustering is building a similarity graph which describes the neighborhood relations among data points. Some recent works build the graph using sparse, low-rank, and -norm-based representation, and have achieved state-of-the-art performance. Howeve…
Diffusion maps are a nonlinear manifold learning technique based on harmonic analysis of a diffusion process over the data. Out-of-sample extensions with computational complexity , where is the number of points comprising the manifold, frustrate applications to online learning applications requiring…
Enhances supervised visualization for unseen data using autoencoders and random forest.
The paper uses machine learning to forecast macroeconomic outcomes with high-dimensional data.
Proposes a new model to maximize out-of-sample Sharpe ratios by forecasting tangency portfolios.
Let be a data set in , where is the training set and is the test one. Many unsupervised learning algorithms based on kernel methods have been developed to provide dimensionality reduction (DR) embedding for a given training set $Φ: \mathbf{X} \to \mat…
Understanding if classifiers generalize to out-of-sample datasets is a central problem in machine learning. Microscopy images provide a standardized way to measure the generalization capacity of image classifiers, as we can image the same classes of objects under increasingly divergent, but controlled factors of variat…
EB improves asset pricing by mining large strategies without lookahead bias.
We study the problem of out-of-sample risk estimation in the high dimensional regime where both the sample size and number of features are large, and can be less than one. Extensive empirical evidence confirms the accuracy of leave-one-out cross validation (LO) for out-of-sample risk estimation. Yet, a un…
Optimal data-driven formulations are found for learning and decision-making with historical data.
Several popular graph embedding techniques for representation learning and dimensionality reduction rely on performing computationally expensive eigendecompositions to derive a nonlinear transformation of the input data space. The resulting eigenvectors encode the embedding coordinates for the training samples only, an…
We propose a Genetic Programming architecture for the generation of foreign exchange trading strategies. The system's principal features are the evolution of free-form strategies which do not rely on any prior models and the utilization of price series from multiple instruments as input data. This latter feature consti…
Estimates error for robust M-estimators with convex penalties.
New model optimizes portfolios over multiple periods using predictive control.
Random feature mapping (RFM) is a popular method for speeding up kernel methods at the cost of losing a little accuracy. We study kernel ridge regression with random feature mapping (RFM-KRR) and establish novel out-of-sample error upper and lower bounds. While out-of-sample bounds for RFM-KRR have been established by …
Endogenous randomness emerges from adversarial market learning.
RandALO speeds up risk estimation for large datasets.
A new deep learning model improves asset pricing predictions.
Non-linear manifold learning enables high-dimensional data analysis, but requires out-of-sample-extension methods to process new data points. In this paper, we propose a manifold learning algorithm based on deep learning to create an encoder, which maps a high-dimensional dataset and its low-dimensional embedding, and …
Validates policies using past observational data with guarantees about out-of-sample performance.
Paper proposes a method to improve prediction intervals for neural networks.
We consider the multi-class classification problem when the training data and the out-of-sample test data may have different distributions and propose a method called BCOPS (balanced and conformal optimized prediction sets). BCOPS constructs a prediction set as a subset of class labels, possibly empty. It tries …
We study the out-of-sample properties of robust empirical optimization problems with smooth -divergence penalties and smooth concave objective functions, and develop a theory for data-driven calibration of the non-negative "robustness parameter" that controls the size of the deviations from the nominal model. Bu…
High-performing equity factor with Sharpe ratio above 13 out-of-sample.
Bayesian method predicts future network configurations from past snapshots.
Downsampling can improve generalization in ridgeless linear regression, especially with optimal sketching size.
We study the profitability of optimal mean reversion trading strategies in the US equity market. Different from regular pair trading practice, we apply maximum likelihood method to construct the optimal static pairs trading portfolio that best fits the Ornstein-Uhlenbeck process, and rigorously estimate the parameters.…
Two strategies for embedding new data points from proximity data are explored.
Robustifies Markowitz portfolios to reduce transaction costs and improve performance.
Meta-GLAR combines global deep representations with local adaptation for improved forecasting accuracy.