This paper proposes an out-of-sample extension framework for a global manifold learning algorithm (Isomap) that uses temporal information in out-of-sample points in order to make the embedding more robust to noise and artifacts. Given a set of noise-free training data and its embedding, the proposed framework extends t…
The paper proves limit theorems for graph embeddings out-of-sample.
problem Proving limit theorems for graph embeddings out-of-sample.
method Least-squares and maximum-likelihood objectives for adjacency and Laplacian spectral embeddings.
result Out-of-sample extensions based on these objectives obey central limit theorems and concentration inequalities.
Paper analyzes mathematical theory behind out-of-sample DR extensions.
problem Developing a solid mathematical foundation for out-of-sample DR extensions.
method Utilizes RKHS theory to treat DR extension as an extension of the identity on RKHS defined on X.
result Shows Nyström-type DR extension as an orthogonal projection and provides conditions for exact DR extension.
Two approaches extend graph embedding to unseen vertices.
problem Extending graph embedding to new data points.
method Least-squares and maximum-likelihood formulations.
result Both methods estimate latent positions with the same error rate under latent position models.
Dimensionality reduction methods are very common in the field of high dimensional data analysis. Typically, algorithms for dimensionality reduction are computationally expensive. Therefore, their applications for the analysis of massive amounts of data are impractical. For example, repeated computations due to accumula…
Landmark Diffusion Maps reduce manifold learning complexity for high-volume data streams.
problem Complexity of out-of-sample extensions in manifold learning techniques.
method Landmark Diffusion Maps (L-dMaps) using pruned spanning trees or k-medoids to select landmark points.
result Up to 50-fold speedups in out-of-sample extension with less than 4% errors in manifold reconstruction.
We consider the problem of vertex classification for graphs constructed from the latent position model. It was shown previously that the approach of embedding the graphs into some Euclidean space followed by classification in that space can yields a universally consistent vertex classifier. However, a major technical d…
This paper presents a new method for dimensionality reduction and out-of-sample extension.
problem Dimensionality reduction and out-of-sample extension in high-dimensional data.
method Adaptive non-linear embedding using positive semi-definite kernel eigenvectors.
result The embedding method is more robust to outliers compared to spectral embedding.
Enhances supervised visualization for unseen data using autoencoders and random forest.
problem Lack of generalization to unseen test sets in supervised dimensionality reduction.
method Combines autoencoder and random forest proximities for out-of-sample extension.
result 40% reduction in training time with 10% of training data, achieving consistent quality.
Several popular graph embedding techniques for representation learning and dimensionality reduction rely on performing computationally expensive eigendecompositions to derive a nonlinear transformation of the input data space. The resulting eigenvectors encode the embedding coordinates for the training samples only, an…
Non-linear manifold learning enables high-dimensional data analysis, but requires out-of-sample-extension methods to process new data points. In this paper, we propose a manifold learning algorithm based on deep learning to create an encoder, which maps a high-dimensional dataset and its low-dimensional embedding, and …
RandALO speeds up risk estimation for large datasets.
problem Estimating out-of-sample risk for large, high-dimensional models.
method RandALO: a randomized approximate leave-one-out estimator.
result RandALO is a computationally efficient risk estimator in high dimensions.
The paper analyzes LOCV for high-dimensional risk estimation, proving error bounds.
problem Estimating out-of-sample prediction error in high-dimensional settings.
method Theoretical analysis of leave-one-out cross validation (LOCV) in penalized regression.
result Finite sample upper bounds on LOCV error, showing it converges to zero as n,p → ∞.
Under the framework of spectral clustering, the key of subspace clustering is building a similarity graph which describes the neighborhood relations among data points. Some recent works build the graph using sparse, low-rank, and ℓ2-norm-based representation, and have achieved state-of-the-art performance. Howeve…
This research extends the Pareto/NBD model using neural networks for better out-of-sample predictions.
problem The limitations of the Pareto/NBD model in predicting out-of-sample data.
method A neural network-based extension of the Pareto/NBD model.
result The proposed method shows extraordinary predictability on repeat purchases at individual and aggregate levels.
Paper proposes a method to improve prediction intervals for neural networks.
problem Improving prediction intervals for neural network models.
method Adapting extremely randomized trees to neural networks to create ensembles.
result The method yields gains in out-of-sample accuracy and is superior to existing methods.
Paper proves bounds for shallow networks without dimensionality issues.
problem Avoiding curse of dimensionality and estimating approximation degree.
method Abstract theorem for G-networks on compact metric measure spaces. result Dimension independent bounds for approximation by shallow networks.
Manifold learning has been successfully applied to a variety of medical imaging problems. Its use in real-time applications requires fast projection onto the low-dimensional space. To this end, out-of-sample extensions are applied by constructing an interpolation function that maps from the input space to the low-dimen…
New method improves feature selection in tree-based models.
problem Previous feature selection methods in tree-based models lack sufficient regularization and sub-optimal performance.
method Developed a new gain penalization approach for tree-based models that allows for flexible feature-specific importance weights.
result The new method improves out-of-sample performance, especially with correlated features.
High-dimensional big data appears in many research fields such as image recognition, biology and collaborative filtering. Often, the exploration of such data by classic algorithms is encountered with difficulties due to `curse of dimensionality' phenomenon. Therefore, dimensionality reduction methods are applied to the…
Proposes a new method for manifold alignment using geometry-regularized twin autoencoders.
problem Traditional MA methods lack out-of-sample extension and real-world applicability.
method Guided representation learning with geometry-regularized twin autoencoders.
result Improves cross-domain generalization and robustness while maintaining alignment fidelity.
Meta-GLAR combines global deep representations with local adaptation for improved forecasting accuracy.
problem Joint learning from related time series boosts accuracy but fails for out-of-sample forecasting.
method Meta-GLAR uses a meta-learning approach to adapt RNN representations for each time series.
result Meta-GLAR outperforms state-of-the-art methods in out-of-sample forecasting accuracy.
High-performing equity factor with Sharpe ratio above 13 out-of-sample.
problem Hidden cross-sectional predictability in stock returns.
method Regime-conditional signal activation combining value and short-term reversal signals.
result Annualized returns of 158.6% with 12.0% volatility, strong performance out-of-sample.
A study finds that only a few factors explain corporate bond risk, rendering extensive bond factor literature redundant.
problem The redundancy of extensive bond factor literature in explaining corporate bond risk premia.
method Bayesian Model Averaging Stochastic Discount Factor analysis of 18 quadrillion models.
result A Bayesian Model Averaging SDF explains risk premia better than low-dimensional models, with an out-of-sample Sharpe ratio of 1.5 to 1.8.
Proposes a bond portfolio solution for managing interest rate risk.
problem Managing long-term assets and liabilities under interest rate risk.
method Proposes a bond portfolio solution based on ambiguity-averse preferences, accommodating various constraints and interest rate perturbations.
result Optimal portfolio can be computed as a simple generalized least squares problem, enhancing out-of-sample performance.
Dimensionality reduction is a topic of recent interest. In this paper, we present the classification constrained dimensionality reduction (CCDR) algorithm to account for label information. The algorithm can account for multiple classes as well as the semi-supervised setting. We present an out-of-sample expressions for …
Paper analyzes high-dimensional portfolio risks and finds empirical out-of-sample relative loss is more reliable.
problem Analyzing risks in high-dimensional portfolios using empirical variance.
method Derives asymptotic behavior of out-of-sample variance and relative loss in high-dimensional settings.
result Empirical out-of-sample relative loss is more reliable than variance in high-dimensional portfolios.
Proposes a parametric t-SNE without perplexity tuning.
problem Non-parametric t-SNE's perplexity parameter limits DR quality.
method Multi-scale parametric t-SNE with deep neural network.
result Produces reliable embeddings with competitive neighborhood preservation.
New method for estimating out-of-sample R² from gene expression data.
problem Lack of a well-defined and unbiased estimator for out-of-sample R².
method Explicitly defined out-of-sample R², provided an unbiased estimator, and calculated standard error.
result Demonstrated improved model comparison for gene expression phenotypes.
Proposes a new framework to optimize portfolios with reduced estimation errors.
problem Estimation errors in multiperiod mean-variance portfolio optimization.
method Reference-regulated multiperiod mean-variance (RRMV) framework.
result Improves portfolio stability and out-of-sample Sharpe ratios.
Paper introduces DOO models to outperform SAA out-of-sample.
problem Outperforming SAA in out-of-sample performance.
method Introduces DOO models that consider both worst-case and best-case scenarios.
result DOO models can always outperform SAA out-of-sample.
A new method optimizes diversity and sparsity for index tracking.
problem Accurately replicating a benchmark index with a small number of diverse assets.
method Jointly optimizes diversity and sparsity using a regularizer based on asset similarity.
result The proposed algorithm outperforms existing methods in out-of-sample backtesting.
New test improves tree ensemble pruning for better model performance.
problem Lack of robust theoretical justification for penalty terms in tree ensembles.
method Developed a novel hypothesis test for tree ensemble split quality.
result Significant reduction in out-of-sample loss using the new test.
This study compares three portfolio design approaches for stock selection.
problem Designing a profitable portfolio with precise stock returns and risks.
method Three portfolio design approaches: mean-variance portfolio, hierarchical risk parity, and autoencoder-based portfolio.
result Autoencoder portfolios outperform MVP on annual returns, but MVP is best on risk-adjusted returns.
OTSL improves structure learning accuracy with out-of-sample and resampling strategies.
problem Determining optimal hyperparameters for structure learning algorithms.
method Out-of-sample Tuning for Structure Learning (OTSL) using resampling strategies.
result Improves graphical accuracy of structure learning algorithms.
The paper proposes a new approach to portfolio selection that maximizes diversification and return.
problem Maximizing diversification and return in portfolio selection.
method A bi-objective model that maximizes a diversification measure and portfolio expected return.
result The return-diversification approach outperforms strategies based on diversification or classical risk-return approaches.
Bagging stabilizes linear interpolators, improving their generalization performance.
problem Unstable linear interpolators fail on noisy data.
method Introduced multiplier-bootstrap-based bagged least square estimator.
result Bagging effectively mitigates variance, leading to bounded prediction risk.
Paper proposes a diagnostic tool for evaluating model performance out-of-sample.
problem Evaluating model performance on unseen data.
method Uses a finite calibration dataset to assess future losses.
result Provides guarantees under weak assumptions and quantifies distribution shifts.
Improved global minimum-variance portfolios using cross-validation for high-dimensional covariance estimation.
problem Ill-conditioned sample covariance matrix in high-dimensional data leads to suboptimal portfolios.
method Cross-validation technique to select tuning parameters for efficient covariance matrix estimation methods.
result Data-driven tuning parameters improve out-of-sample performance of global minimum-variance portfolios.
We identify and validate a model for PCR in high dimensions, improving prediction guarantees.
problem Model identification and out-of-sample prediction in high-dimensional error-in-variables settings.
method Analysis of principal component regression (PCR) in fixed design settings, introducing a linear algebraic condition.
result Consistent model identification and improved out-of-sample prediction guarantees.
Study the impact of overfitting on linear predictive models' performance.
problem Overfitting reduces the out-of-sample performance of linear predictive trading strategies.
method Computed in- and out-of-sample means and variances of PnLs to derive replication ratios.
result Replication ratio diminishes for complex strategies with many assets.
New framework for learning policies that converge in out-of-sample regions.
problem Reliable out-of-sample recovery in imitation learning.
method Contractive dynamical systems and recurrent equilibrium networks.
result Policy rollouts converge regardless of perturbations, enabling efficient OOS recovery.
Optimizes decisions without knowing the true distribution using historical data.
problem Optimizing decisions without knowing the true distribution.
method Combines sampling and bisection search algorithms to solve an optimization problem.
result Proves sufficient conditions for local out-of-sample optimality.
Bayesian neural networks show good correlation between out-of-sample performance and Bayesian evidence.
problem Improving the out-of-sample performance of Bayesian neural networks.
method Numerical sampling of Bayesian posterior, ensembling over architectures, analysis of evidence vs. model size.
result Good correlation between out-of-sample performance and Bayesian evidence; ensembling improves performance.
Paper develops a method to predict spatial point processes with guarantees.
problem Predicting the number of events in space with uncertainty.
method Regularized method to learn spatial models with out-of-sample guarantees.
result Method provides valid prediction intervals even when model is misspecified.
Improves test set performance and reduces out-of-sample disappointment for unstable models.
problem Ensuring strong test set performance via cross-validation for unstable models.
method Nested k-fold cross-validation with hyperparameter selection based on a weighted sum of cross-validation metric and model stability measure.
result Improves out-of-sample MSE for sparse ridge regression and CART by 4% and 2% respectively, compared to k-fold cross-validation.
New model predicts stock performance in large equity markets.
problem Predicting stock performance in large equity markets over long time horizons.
method Rank-based volatility stabilized models calibrated to empirical data.
result The model exhibits relative arbitrage and statistically fits empirical features.
Paper presents a new way to analyze machine learning generalization without probabilistic assumptions.
problem Traditional generalization analysis assumes i.i.d. data, which is often unverifiable.
method Uses sensitivity analysis of optimization problems to derive deterministic generalization bounds.
result Obtains generalization bounds that relate in-sample and out-of-sample evaluations through an error term quantifying data similarity.