This paper proposes an out-of-sample extension framework for a global manifold learning algorithm (Isomap) that uses temporal information in out-of-sample points in order to make the embedding more robust to noise and artifacts. Given a set of noise-free training data and its embedding, the proposed framework extends t…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Graph embeddings, a class of dimensionality reduction techniques designed for relational data, have proven useful in exploring and modeling network structure. Most dimensionality reduction methods allow out-of-sample extensions, by which an embedding can be applied to observations not present in the training set. Appli…
Many popular dimensionality reduction procedures have out-of-sample extensions, which allow a practitioner to apply a learned embedding to observations not seen in the initial training sample. In this work, we consider the problem of obtaining an out-of-sample extension for the adjacency spectral embedding, a procedure…
Dimensionality reduction methods are very common in the field of high dimensional data analysis. Typically, algorithms for dimensionality reduction are computationally expensive. Therefore, their applications for the analysis of massive amounts of data are impractical. For example, repeated computations due to accumula…
Let be a data set in , where is the training set and is the test one. Many unsupervised learning algorithms based on kernel methods have been developed to provide dimensionality reduction (DR) embedding for a given training set $Φ: \mathbf{X} \to \mat…
We consider the problem of vertex classification for graphs constructed from the latent position model. It was shown previously that the approach of embedding the graphs into some Euclidean space followed by classification in that space can yields a universally consistent vertex classifier. However, a major technical d…
Diffusion maps are a nonlinear manifold learning technique based on harmonic analysis of a diffusion process over the data. Out-of-sample extensions with computational complexity , where is the number of points comprising the manifold, frustrate applications to online learning applications requiring…
Enhances supervised visualization for unseen data using autoencoders and random forest.
Several popular graph embedding techniques for representation learning and dimensionality reduction rely on performing computationally expensive eigendecompositions to derive a nonlinear transformation of the input data space. The resulting eigenvectors encode the embedding coordinates for the training samples only, an…
Non-linear manifold learning enables high-dimensional data analysis, but requires out-of-sample-extension methods to process new data points. In this paper, we propose a manifold learning algorithm based on deep learning to create an encoder, which maps a high-dimensional dataset and its low-dimensional embedding, and …
In machine learning or statistics, it is often desirable to reduce the dimensionality of a sample of data points in a high dimensional space . This paper introduces a dimensionality reduction method where the embedding coordinates are the eigenvectors of a positive semi-definite kernel obtained as the sol…
RandALO speeds up risk estimation for large datasets.
The paper analyzes LOCV for high-dimensional risk estimation, proving error bounds.
Under the framework of spectral clustering, the key of subspace clustering is building a similarity graph which describes the neighborhood relations among data points. Some recent works build the graph using sparse, low-rank, and -norm-based representation, and have achieved state-of-the-art performance. Howeve…
Paper proposes a method to improve prediction intervals for neural networks.
Manifold learning has been successfully applied to a variety of medical imaging problems. Its use in real-time applications requires fast projection onto the low-dimensional space. To this end, out-of-sample extensions are applied by constructing an interpolation function that maps from the input space to the low-dimen…
New method improves feature selection in tree-based models.
High-dimensional big data appears in many research fields such as image recognition, biology and collaborative filtering. Often, the exploration of such data by classic algorithms is encountered with difficulties due to `curse of dimensionality' phenomenon. Therefore, dimensionality reduction methods are applied to the…
Proposes a new method for manifold alignment using geometry-regularized twin autoencoders.
Meta-GLAR combines global deep representations with local adaptation for improved forecasting accuracy.
High-performing equity factor with Sharpe ratio above 13 out-of-sample.
A study finds that only a few factors explain corporate bond risk, rendering extensive bond factor literature redundant.
Proposes a bond portfolio solution for managing interest rate risk.
Dimensionality reduction is a topic of recent interest. In this paper, we present the classification constrained dimensionality reduction (CCDR) algorithm to account for label information. The algorithm can account for multiple classes as well as the semi-supervised setting. We present an out-of-sample expressions for …
Paper analyzes high-dimensional portfolio risks and finds empirical out-of-sample relative loss is more reliable.
Proposes a parametric t-SNE without perplexity tuning.
New method for estimating out-of-sample R² from gene expression data.
Whether stochastic or parametric, the Pareto/NBD model can only be utilized for an in-sample prediction rather than an out-of-sample prediction. This research thus provides a neural network based extension of the Pareto/NBD model to estimate the out-of-sample parameters, which overrides the estimation burden and the ap…
Paper introduces DOO models to outperform SAA out-of-sample.
Proposes a new framework to optimize portfolios with reduced estimation errors.
The global minimum-variance portfolio is a typical choice for investors because of its simplicity and broad applicability. Although it requires only one input, namely the covariance matrix of asset returns, estimating the optimal solution remains a challenge. In the presence of high-dimensionality in the data, the samp…
This paper proves an abstract theorem addressing in a unified manner two important problems in function approximation: avoiding curse of dimensionality and estimating the degree of approximation for out-of-sample extension in manifold learning. We consider an abstract (shallow) network that includes, for example, neura…
New test improves tree ensemble pruning for better model performance.
OTSL improves structure learning accuracy with out-of-sample and resampling strategies.
This study compares three portfolio design approaches for stock selection.
The paper proposes a new approach to portfolio selection that maximizes diversification and return.
Bagging stabilizes linear interpolators, improving their generalization performance.
Paper proposes a diagnostic tool for evaluating model performance out-of-sample.
We identify and validate a model for PCR in high dimensions, improving prediction guarantees.
Study the impact of overfitting on linear predictive models' performance.
New framework for learning policies that converge in out-of-sample regions.
Optimizes decisions without knowing the true distribution using historical data.
Paper develops a method to predict spatial point processes with guarantees.
Bayesian neural networks show good correlation between out-of-sample performance and Bayesian evidence.
Improves test set performance and reduces out-of-sample disappointment for unstable models.
Paper presents a new way to analyze machine learning generalization without probabilistic assumptions.
New model predicts stock performance in large equity markets.
Ridge regression CV loss may have multiple local optima.