The paper proves limit theorems for graph embeddings out-of-sample.
problem Proving limit theorems for graph embeddings out-of-sample.
method Least-squares and maximum-likelihood objectives for adjacency and Laplacian spectral embeddings.
result Out-of-sample extensions based on these objectives obey central limit theorems and concentration inequalities.
This paper proposes an out-of-sample extension framework for a global manifold learning algorithm (Isomap) that uses temporal information in out-of-sample points in order to make the embedding more robust to noise and artifacts. Given a set of noise-free training data and its embedding, the proposed framework extends t…
We consider the problem of vertex classification for graphs constructed from the latent position model. It was shown previously that the approach of embedding the graphs into some Euclidean space followed by classification in that space can yields a universally consistent vertex classifier. However, a major technical d…
Many popular dimensionality reduction procedures have out-of-sample extensions, which allow a practitioner to apply a learned embedding to observations not seen in the initial training sample. In this work, we consider the problem of obtaining an out-of-sample extension for the adjacency spectral embedding, a procedure…
Several popular graph embedding techniques for representation learning and dimensionality reduction rely on performing computationally expensive eigendecompositions to derive a nonlinear transformation of the input data space. The resulting eigenvectors encode the embedding coordinates for the training samples only, an…
Two strategies for embedding new data points from proximity data are explored.
problem Embedding new data points using proximity data.
method Two competing strategies: projection and restricted reconstruction.
result Projection and restricted reconstruction can be derived from kernel methods.
Let X=X∪Z be a data set in RD, where X is the training set and Z is the test one. Many unsupervised learning algorithms based on kernel methods have been developed to provide dimensionality reduction (DR) embedding for a given training set $Φ: \mathbf{X} \to \mat…
In machine learning or statistics, it is often desirable to reduce the dimensionality of a sample of data points in a high dimensional space Rd. This paper introduces a dimensionality reduction method where the embedding coordinates are the eigenvectors of a positive semi-definite kernel obtained as the sol…
Non-linear manifold learning enables high-dimensional data analysis, but requires out-of-sample-extension methods to process new data points. In this paper, we propose a manifold learning algorithm based on deep learning to create an encoder, which maps a high-dimensional dataset and its low-dimensional embedding, and …
Dimensionality reduction is a topic of recent interest. In this paper, we present the classification constrained dimensionality reduction (CCDR) algorithm to account for label information. The algorithm can account for multiple classes as well as the semi-supervised setting. We present an out-of-sample expressions for …
The paper uses machine learning to forecast macroeconomic outcomes with high-dimensional data.
problem Forecasting the full conditional distribution of macroeconomic outcomes.
method Systematically integrating three key principles: high-dimensional data with regularization, rigorous out-of-sample validation, and incorporating nonlinearities.
result Regularization via shrinkage is essential to control model complexity, while nonlinearities yield limited improvements in predictive accuracy.
Diffusion maps are a nonlinear manifold learning technique based on harmonic analysis of a diffusion process over the data. Out-of-sample extensions with computational complexity O(N), where N is the number of points comprising the manifold, frustrate applications to online learning applications requiring…
Enhances supervised visualization for unseen data using autoencoders and random forest.
problem Lack of generalization to unseen test sets in supervised dimensionality reduction.
method Combines autoencoder and random forest proximities for out-of-sample extension.
result 40% reduction in training time with 10% of training data, achieving consistent quality.
This research extends the Pareto/NBD model using neural networks for better out-of-sample predictions.
problem The limitations of the Pareto/NBD model in predicting out-of-sample data.
method A neural network-based extension of the Pareto/NBD model.
result The proposed method shows extraordinary predictability on repeat purchases at individual and aggregate levels.
News embeddings improve volatility forecasts.
problem Improving volatility forecasting accuracy.
method Transformed news text into embeddings, evaluated standalone and combined with benchmarks.
result News contains useful predictive information, especially for stock-related content.
New framework for interpretable firm characteristics factors.
problem Creating statistically efficient and economically interpretable factors from firm characteristics.
method Grouping related characteristics and deriving one factor per group, combining economic intuition with data-driven clustering.
result Parsimonious, transparent factors outperform benchmarks in out-of-sample tests.
High-dimensional big data appears in many research fields such as image recognition, biology and collaborative filtering. Often, the exploration of such data by classic algorithms is encountered with difficulties due to `curse of dimensionality' phenomenon. Therefore, dimensionality reduction methods are applied to the…
Proposes a new method for manifold alignment using geometry-regularized twin autoencoders.
problem Traditional MA methods lack out-of-sample extension and real-world applicability.
method Guided representation learning with geometry-regularized twin autoencoders.
result Improves cross-domain generalization and robustness while maintaining alignment fidelity.
Explains SNE, t-SNE, and their variants for manifold learning.
problem Dimensionality reduction and manifold learning.
method Probabilistic approach using Gaussian and Student-t distributions.
result Out-of-sample extension and acceleration methods for t-SNE.
The Joint Optimization of Fidelity and Commensurability (JOFC) manifold matching methodology embeds an omnibus dissimilarity matrix consisting of multiple dissimilarities on the same set of objects. One approach to this embedding optimizes the preservation of fidelity to each individual dissimilarity matrix together wi…
We propose a deep learning approach for discovering kernels tailored to identifying clusters over sample data. Our neural network produces sample embeddings that are motivated by--and are at least as expressive as--spectral clustering. Our training objective, based on the Hilbert Schmidt Information Criterion, can be o…
Proposes a parametric t-SNE without perplexity tuning.
problem Non-parametric t-SNE's perplexity parameter limits DR quality.
method Multi-scale parametric t-SNE with deep neural network.
result Produces reliable embeddings with competitive neighborhood preservation.
The study finds that supply chain information from LLM embeddings improves stock returns predictions.
problem Predicting stock returns using textual information from annual reports.
method Combining LLM embeddings of annual reports with supply chain knowledge graph propagation.
result Network-augmented embeddings significantly predict stock returns with a Sharpe ratio of 0.86 and alpha of 7.27%.
Survey of Laplacian-based methods for data dimensionality reduction and embedding.
problem Efficiently reducing high-dimensional data to lower dimensions while preserving important features and structures.
method Laplacian-based methods including spectral clustering, Laplacian eigenmap, locality preserving projection, graph embedding, and diffusion map.
result Comprehensive overview of various optimization variants and applications of Laplacian-based techniques.
Novel approach to OT using kernel mean embeddings controls overfitting and achieves dimension-free sample complexity.
problem Consistently estimate optimal transport plan from samples.
method Pose OT as learning kernel mean embedding, employ MMD regularization.
result ε-optimal recovery of transport plan and map with dimension-free sample complexity.
CASTLE learns causal DAG to improve model generalization.
problem Improving model generalization to out-of-sample data.
method CASTLE learns causal relationships via adjacency matrix embedded in neural network input layers, reconstructing only causal features.
result CASTLE leads to better out-of-sample predictions compared to other regularizers.
Paper analyzes high-dimensional portfolio risks and finds empirical out-of-sample relative loss is more reliable.
problem Analyzing risks in high-dimensional portfolios using empirical variance.
method Derives asymptotic behavior of out-of-sample variance and relative loss in high-dimensional settings.
result Empirical out-of-sample relative loss is more reliable than variance in high-dimensional portfolios.
Survey of Locally Linear Embedding and its variants.
problem Representing high-dimensional data in a lower-dimensional space while preserving local structure.
method Explains various LLE and variant methods, including kernel LLE, inverse LLE, feature fusion, out-of-sample embedding, incremental LLE, landmark LLE, supervised LLE, robust LLE, fusion with other methods, and weighted LLE.
result Comprehensive overview of LLE and its variants.
This paper reviews MDS, Sammon mapping, and Isomap, explaining their theory and applications.
problem Exploring multidimensional data structures and mappings.
method Explains classical MDS, metric MDS, kernel classical MDS, Sammon mapping, Isomap, and their applications.
result Detailed understanding of MDS, Sammon mapping, and Isomap methods.
The paper predicts responses on out-of-sample nodes using latent positions on unknown curves.
problem Predicting responses on out-of-sample nodes with latent positions on unknown curves.
method Manifold learning and graph embedding technique using latent positions.
result Convergence guarantees for predicting responses on out-of-sample nodes.
New method for estimating out-of-sample R² from gene expression data.
problem Lack of a well-defined and unbiased estimator for out-of-sample R².
method Explicitly defined out-of-sample R², provided an unbiased estimator, and calculated standard error.
result Demonstrated improved model comparison for gene expression phenotypes.
Sep-SpectralNet improves SE for broader applicability and scalability.
problem Three main drawbacks of current SE implementations: generalizability, scalability, and eigenvectors separation.
method Sep-SpectralNet extends SpectralNet with an eigenvector separation post-processing step.
result Sep-SpectralNet achieves consistent SE approximation and generalization, enhancing scalability and applicability.
Paper introduces DOO models to outperform SAA out-of-sample.
problem Outperforming SAA in out-of-sample performance.
method Introduces DOO models that consider both worst-case and best-case scenarios.
result DOO models can always outperform SAA out-of-sample.
MEDAL converts manifold embeddings into models for rigorous validation.
problem Challenges in validating manifold embeddings without held-out validation.
method Develops MEDAL framework that distills embeddings into autoencoder models.
result Enables rigorous validation of manifold embeddings and hyperparameters.
SpecNet2 improves spectral embedding without orthogonalization, achieving better performance and efficiency.
problem Improving spectral embedding methods for better performance and efficiency.
method Optimizes an equivalent objective of the eigen-problem without orthogonalization, allowing separate row and column sampling.
result Local and global convergence of the new objective using batch-based gradient descent is proven, and improved performance and efficiency are demonstrated on simulated and image datasets.
Modern LLMs fail at authorship attribution without fine-tuning, but topic embeddings outperform them.
problem Authorship attribution of the Federalist Papers using modern LLMs.
method Examined popular LLMs, compared word/phrase embeddings, and used Bayesian analysis with topic embeddings.
result Topic embeddings trained on 'function words' outperform default LLM embeddings in authorship attribution.
OTSL improves structure learning accuracy with out-of-sample and resampling strategies.
problem Determining optimal hyperparameters for structure learning algorithms.
method Out-of-sample Tuning for Structure Learning (OTSL) using resampling strategies.
result Improves graphical accuracy of structure learning algorithms.
Dimensionality reduction methods are very common in the field of high dimensional data analysis. Typically, algorithms for dimensionality reduction are computationally expensive. Therefore, their applications for the analysis of massive amounts of data are impractical. For example, repeated computations due to accumula…
Paper proposes a diagnostic tool for evaluating model performance out-of-sample.
problem Evaluating model performance on unseen data.
method Uses a finite calibration dataset to assess future losses.
result Provides guarantees under weak assumptions and quantifies distribution shifts.
Kalshi prediction markets forecast cryptocurrency volatility through monetary policy and inflation signals.
problem Forecasting cryptocurrency volatility using prediction markets.
method Monetary policy and inflation signals from Kalshi prediction markets.
result Signals from Kalshi prediction markets predict cryptocurrency volatility with statistical significance.
End-to-end neural network optimizes portfolios by directly learning allocations from features.
problem Error maximization in two-step portfolio optimization.
method Single feed-forward neural network combining prediction and optimization.
result Model-based end-to-end framework achieves Sharpe ratio of 1.16.
We identify and validate a model for PCR in high dimensions, improving prediction guarantees.
problem Model identification and out-of-sample prediction in high-dimensional error-in-variables settings.
method Analysis of principal component regression (PCR) in fixed design settings, introducing a linear algebraic condition.
result Consistent model identification and improved out-of-sample prediction guarantees.
Study the impact of overfitting on linear predictive models' performance.
problem Overfitting reduces the out-of-sample performance of linear predictive trading strategies.
method Computed in- and out-of-sample means and variances of PnLs to derive replication ratios.
result Replication ratio diminishes for complex strategies with many assets.
New framework for learning policies that converge in out-of-sample regions.
problem Reliable out-of-sample recovery in imitation learning.
method Contractive dynamical systems and recurrent equilibrium networks.
result Policy rollouts converge regardless of perturbations, enabling efficient OOS recovery.
Optimizes decisions without knowing the true distribution using historical data.
problem Optimizing decisions without knowing the true distribution.
method Combines sampling and bisection search algorithms to solve an optimization problem.
result Proves sufficient conditions for local out-of-sample optimality.
Bayesian neural networks show good correlation between out-of-sample performance and Bayesian evidence.
problem Improving the out-of-sample performance of Bayesian neural networks.
method Numerical sampling of Bayesian posterior, ensembling over architectures, analysis of evidence vs. model size.
result Good correlation between out-of-sample performance and Bayesian evidence; ensembling improves performance.
Paper develops a method to predict spatial point processes with guarantees.
problem Predicting the number of events in space with uncertainty.
method Regularized method to learn spatial models with out-of-sample guarantees.
result Method provides valid prediction intervals even when model is misspecified.
Proposes deep hedging for index options using implied volatility surface.
problem Managing risk in index option portfolios with complex dynamics.
method Integrates surface-informed decisions with multiple hedging instruments, accounting for transaction costs and variance risk premium.
result Consistently outperforms traditional hedging strategies across various market conditions.