R package psvmSDR simplifies SDR computation for machine learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The principal support vector machines method (Li et al., 2011) is a powerful tool for sufficient dimension reduction that replaces original predictors with their low-dimensional linear combinations without loss of information. However, the computational burden of the principal support vector machines method constrains …
Study identifies and estimates treatment effect heterogeneity within principal stratification subpopulations.
Two new PCA variants improve financial data analysis.
A new method reduces data movement in neural network training.
New algorithm learns principal subspace from random samples.
We consider the classification problem and focus on nonlinear methods for classification on manifolds. For multivariate datasets lying on an embedded nonlinear Riemannian manifold within the higher-dimensional ambient space, we aim to acquire a classification boundary for the classes with labels, using the intrinsic me…
Kernelized PCovR reveals structure-property relations in chemistry and materials.
We present a novel algorithm (Principal Sensitivity Analysis; PSA) to analyze the knowledge of the classifier obtained from supervised machine learning techniques. In particular, we define principal sensitivity map (PSM) as the direction on the input space to which the trained classifier is most sensitive, and use anal…
Study robust estimation of principal components under adversarial perturbations.
Study uses machine learning to estimate effective policies in settings with hidden individual actions.
New PCA method for derivatives problems.
This paper is a tutorial for eigenvalue and generalized eigenvalue problems. We first introduce eigenvalue problem, eigen-decomposition (spectral decomposition), and generalized eigenvalue problem. Then, we mention the optimization problems which yield to the eigenvalue and generalized eigenvalue problems. We also prov…
Survey of SDR methods for high-dimensional regression and embedding.
We propose a fair principal component analysis method that balances reconstruction error and subgroup fairness.
Optimal contracts help principals delegate data collection in decentralized ML.
POTD estimates SDR subspace using optimal transport for binary response.
Capacity control, the bias/variance dilemma, and learning unknown functions from data, are all concerned with identifying effective and consistent fits of unknown geometric loci to random data points. A geometric locus is a curve or surface formed by points, all of which possess some uniform property. A geometric locus…
PCA-Guided Quantile Sampling preserves data structure in large datasets.
Improves model predictability by mixing forecasts and orthogonalizing models.
This paper tackles distributed estimation of the top-L eigenspace in PCA for large data sets.
Principal Component Analysis (PCA) has wide applications in machine learning, text mining and computer vision. Classical PCA based on a Gaussian noise model is fragile to noise of large magnitude. Laplace noise assumption based PCA methods cannot deal with dense noise effectively. In this paper, we propose Cauchy Princ…
ML PCA detects phase transitions in muon spectroscopy data.
NECO detects out-of-distribution data using neural collapse properties.
The study predicts Kronecker coefficients using interpretable machine learning models.
Machine learning techniques have gained prominence for the analysis of resting-state functional Magnetic Resonance Imaging (rs-fMRI) data. Here, we present an overview of various unsupervised and supervised machine learning applications to rs-fMRI. We present a methodical taxonomy of machine learning methods in resting…
The paper addresses causal mediation analysis with post-treatment events, proposing robust estimators and efficient methods.
Machine learning is used to approximate density functionals. For the model problem of the kinetic energy of non-interacting fermions in 1d, mean absolute errors below 1 kcal/mol on test densities similar to the training set are reached with fewer than 100 training densities. A predictor identifies if a test density is …
Tensor completion and robust principal component analysis have been widely used in machine learning while the key problem relies on the minimization of a tensor rank that is very challenging. A common way to tackle this difficulty is to approximate the tensor rank with the norm of singular values based on its …
Is all of machine learning supervised to some degree? The field of machine learning has traditionally been categorized pedagogically into ; where supervised learning has typically referred to learning from labeled data, while unsupervised learning has typically referred to learning …
We investigate coresets - succinct, small summaries of large data sets - so that solutions found on the summary are provably competitive with solution found on the full data set. We provide an overview over the state-of-the-art in coreset construction for machine learning. In Section 2, we present both the intuition be…
We employ unsupervised machine learning techniques to learn latent parameters which best describe states of the two-dimensional Ising model and the three-dimensional XY model. These methods range from principal component analysis to artificial neural network based variational autoencoders. The states are sampled using …
The paper uses news headlines to predict stock prices using embeddings.
Interpretability has become an important issue in the machine learning field, along with the success of layered neural networks in various practical tasks. Since a trained layered neural network consists of a complex nonlinear relationship between large number of parameters, we failed to understand how they could achie…
A hybrid approach detects financial market regime switches using PCA and k-means.
Study shows resampling can drastically alter PCA results.
Develops an ℓ_p theory for PCA and spectral clustering.
Correspondence analysis (CA) is a multivariate statistical tool used to visualize and interpret data dependencies by finding maximally correlated embeddings of pairs of random variables. CA has found applications in fields ranging from epidemiology to social sciences; however, current methods do not scale to large, hig…
Proposes VEESA pipeline for interpreting ML models with functional data.
This paper proposes a brain-inspired approach to quantum machine learning with the goal of circumventing many of the complications of other approaches. The fact that quantum processes are unitary presents both opportunities and challenges. A principal opportunity is that a large number of computations can be carried ou…
A new method clusters rows of a matrix of point processes.
This paper solves tensor robust principal component analysis via scaled gradient descent.
Unified method learns latent spaces from labeled data.
Kernel ridge regression is used to approximate the kinetic energy of non-interacting fermions in a one-dimensional box as a functional of their density. The properties of different kernels and methods of cross-validation are explored, and highly accurate energies are achieved. Accurate {\em constrained optimal densitie…
Synthetic experiments are crucial for assessing causal machine learning methods.
Survey of embedding methods for high-dimensional and network data.
Most of machine learning approaches have stemmed from the application of minimizing the mean squared distance principle, based on the computationally efficient quadratic optimization methods. However, when faced with high-dimensional and noisy data, the quadratic error functionals demonstrated many weaknesses including…
Machine learning improves financial stress testing in Indian markets.