PSMM method optimizes matrix sufficient dimension reduction.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New method corrects missing data bias in dimension reduction.
Paper improves efficiency in matrix computations for Gaussian processes.
Paper introduces a taxonomy of reduction matrices for more efficient graph coarsening.
Efficiently reduces rank of non-negative matrices with quadratic time complexity.
Formula for sectional curvatures on matrix groups.
The paper shows how to recover true node positions from a graph or similarity matrix.
We reframe linear dimensionality reduction as a problem of Bayesian inference on matrix manifolds. This natural paradigm extends the Bayesian framework to dimensionality reduction tasks in higher dimensions with simpler models at greater speeds. Here an orthogonal basis is treated as a single point on a manifold and is…
Efficiently solves large portfolio optimization problems by reducing and sparsifying covariance matrices.
Large textual corpora are often represented by the document-term frequency matrix whose elements are the frequency of terms; however, this matrix has two problems: sparsity and high dimensionality. Four dimension reduction strategies are used to address these problems. Of the four strategies, unsupervised feature trans…
New methods for sketching non-PSD matrices improve regression and optimization tasks.
Sparse random projection (RP) is a popular tool for dimensionality reduction that shows promising performance with low computational complexity. However, in the existing sparse RP matrices, the positions of non-zero entries are usually randomly selected. Although they adopt uniform sampling with replacement, due to lar…
The study proposes a method for risk reduction without relying on risk measurement.
We give a simple explicit formula for turnover reduction when a large number of alphas are traded on the same execution platform and trades are crossed internally. We model turnover reduction via alpha correlations. Then, for a large number of alphas, turnover reduction is related to the largest eigenvalue and the corr…
Unified framework for graph coarsening using node features and graph matrices.
Over the years data has become increasingly higher dimensional, which has prompted an increased need for dimension reduction techniques. This is perhaps especially true for clustering (unsupervised classification) as well as semi-supervised and supervised classification. Although dimension reduction in the area of clus…
Study uses NMF to reduce cancer microarray data dimensions.
We present a very fast algorithm for general matrix factorization of a data matrix for use in the statistical analysis of high-dimensional data via latent factors. Such data are prevalent across many application areas and generate an ever-increasing demand for methods of dimension reduction in order to undertake the st…
sgdGMF efficiently estimates generalized matrix factorization models for single-cell RNA sequencing data.
CCP clusters correlated features and projects them to 1D for efficient dimensionality reduction.
A collaborative convex framework for factoring a data matrix into a non-negative product , with a sparse coefficient matrix , is proposed. We restrict the columns of the dictionary matrix to coincide with certain columns of the data matrix , thereby guaranteeing a physically meaningful dictionary and …
The abstract discusses conditions for hyperkähler manifolds and Kähler reduction.
CDP reduces point cloud dimensions by preserving detour-induced local non-convexity.
The paper defines and classifies Cappell-Shaneson polynomials.
New optimizer MARS-M combines variance reduction with Muon for faster LLM training.
Survey of Laplacian-based methods for data dimensionality reduction and embedding.
A novel approach is put forth that utilizes data similarity, quantified on a graph, to improve upon the reconstruction performance of principal component analysis. The tasks of data dimensionality reduction and reconstruction are formulated as graph filtering operations, that enable the exploitation of data node connec…
We propose dimension reduction methods for sparse, high-dimensional multivariate response regression models. Both the number of responses and that of the predictors may exceed the sample size. Sometimes viewed as complementary, predictor selection and rank reduction are the most popular strategies for obtaining lower-d…
New method solves matrix completion problems to certifiable optimality.
NMF with specific constraints is equivalent to LDA.
Method reduces categorical data to lower dimensions using density matrices.
Improved LDA method for better classification and dimensionality reduction.
In this paper, we examine the problem of approximating a general linear dimensionality reduction (LDR) operator, represented as a matrix with , by a partial circulant matrix with rows related by circular shifts. Partial circulant matrices admit fast implementations via Fourier tra…
We present Matrix Krasulina, an algorithm for online k-PCA, by generalizing the classic Krasulina's method (Krasulina, 1969) from vector to matrix case. We show, both theoretically and empirically, that the algorithm naturally adapts to data low-rankness and converges exponentially fast to the ground-truth principal su…
We propose a unified framework to speed up the existing stochastic matrix factorization (SMF) algorithms via variance reduction. Our framework is general and it subsumes several well-known SMF formulations in the literature. We perform a non-asymptotic convergence analysis of our framework and derive computational and …
In this paper we show that for the purposes of dimensionality reduction certain class of structured random matrices behave similarly to random Gaussian matrices. This class includes several matrices for which matrix-vector multiply can be computed in log-linear time, providing efficient dimensionality reduction of gene…
We consider the concept of Stokes-Dirac structures in boundary control theory proposed by van der Schaft and Maschke. We introduce Poisson reduction in this context and show how Stokes-Dirac structures can be derived through symmetry reduction from a canonical Dirac structure on the unreduced phase space. In this way, …
A new method uses Gram matrix for efficient multivariate functional principal components.
Efficient method estimates intrinsic dimension for big data.
It is difficult to find the optimal sparse solution of a manifold learning based dimensionality reduction algorithm. The lasso or the elastic net penalized manifold learning based dimensionality reduction is not directly a lasso penalized least square problem and thus the least angle regression (LARS) (Efron et al. \ci…
In recent years, data have become increasingly higher dimensional and, therefore, an increased need has arisen for dimension reduction techniques for clustering. Although such techniques are firmly established in the literature for multivariate data, there is a relative paucity in the area of matrix variate, or three-w…
New method estimates Gaussian vector functions more efficiently.
For the problems of low-rank matrix completion, the efficiency of the widely-used nuclear norm technique may be challenged under many circumstances, especially when certain basis coefficients are fixed, for example, the low-rank correlation matrix completion in various fields such as the financial market and the low-ra…
A parsimonious model reduces over-parameterization in skewed matrix variate mixtures.
We obtain the first polynomial-time algorithm for exact tensor completion that improves over the bound implied by reduction to matrix completion. The algorithm recovers an unknown 3-tensor with incoherent, orthogonal components in from randomly observed entries of the tensor…
Dimensional reduction of high dimensional data can be achieved by keeping only the relevant eigenmodes after principal component analysis. However, differentiating relevant eigenmodes from the random noise eigenmodes is problematic. A new method based on the random matrix theory and a statistical goodness-of-fit test i…
PCA (Principal Component Analysis) and its variants areubiquitous techniques for matrix dimension reduction and reduced-dimensionlatent-factor extraction. One significant challenge in using PCA, is thechoice of the number of principal components. The information-theoreticMDL (Minimum Description Length) principle gives…
Linear dimensionality reduction methods are a cornerstone of analyzing high dimensional data, due to their simple geometric interpretations and typically attractive computational properties. These methods capture many data features of interest, such as covariance, dynamical structure, correlation between data sets, inp…