Study extends bounds on sample covariance matrices with general dependence.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New algorithm learns low-rank matrices with linear number of samples.
Estimates covariance matrices with correlations between samples.
We consider the problem of sampling from posterior distributions for Bayesian models where some parameters are restricted to be orthogonal matrices. Such matrices are sometimes used in neural networks models for reasons of regularization and stabilization of training procedures, and also can parameterize matrices of bo…
The paper presents two schemes for sampling matrices from specific distributions on a manifold.
We present a novel technique for learning the mass matrices in samplers obtained from discretized dynamics that preserve some energy function. Existing adaptive samplers use Riemannian preconditioning techniques, where the mass matrices are functions of the parameters being sampled. This leads to significant complexiti…
Scalable method completes ill-conditioned matrices from few samples.
Two algorithms estimate Wasserstein distance matrices from few entries for manifold learning.
Motivated by a sampling problem basic to computational statistical inference, we develop a nearly optimal algorithm for a fundamental problem in spectral graph theory and numerical analysis. Given an SDDM matrix , and a constant , our algorithm gives efficient access to a…
Paper develops new method for detecting latent structure in large symmetric data matrices.
WISDoM uses the Wishart distribution to analyze neurological data like EEG and brain connectivity.
We calculate eigenvector overlaps between intersecting time periods of covariance matrices.
We propose a novel approach for sampling realistic financial correlation matrices. This approach is based on generative adversarial networks. Experiments demonstrate that generative adversarial networks are able to recover most of the known stylized facts about empirical correlation matrices estimated on asset returns.…
Kernel matrices (e.g. Gram or similarity matrices) are essential for many state-of-the-art approaches to classification, clustering, and dimensionality reduction. For large datasets, the cost of forming and factoring such kernel matrices becomes intractable. To address this challenge, we introduce a new adaptive sampli…
We consider the problem of exact recovery of any matrix of rank from a small number of observed entries via the standard nuclear norm minimization framework. Such low-rank matrices have degrees of freedom . We show that any arbitrary low-rank matrices can be recovered exa…
How many samples are sufficient to guarantee that the eigenvectors and eigenvalues of the sample covariance matrix are close to those of the actual covariance matrix? For a wide family of distributions, including distributions with finite second moment and distributions supported in a centered Euclidean ball, we prove …
Improved method for computing Fréchet means on SPD matrices.
It is natural to ask: what kinds of matrices satisfy the Restricted Eigenvalue (RE) condition? In this paper, we associate the RE condition (Bickel-Ritov-Tsybakov 09) with the complexity of a subset of the sphere in , where is the dimensionality of the data, and show that a class of random matrices with indep…
Quantum GBS boosts asset clustering for robust statistical arbitrage portfolios.
The accurate detection of small deviations in given density matrices is important for quantum information processing. Here we propose a new method based on the concept of data mining. We demonstrate that the proposed method can more accurately detect small erroneous deviations in reconstructed density matrices, which c…
Matrix factorization is a simple and effective solution to the recommendation problem. It has been extensively employed in the industry and has attracted much attention from the academia. However, it is unclear what the low-dimensional matrices represent. We show that matrix factorization can actually be seen as simult…
Kernel methods are successful approaches for different machine learning problems. This success is mainly rooted in using feature maps and kernel matrices. Some methods rely on the eigenvalues/eigenvectors of the kernel matrix, while for other methods the spectral information can be used to estimate the excess risk. An …
Three methods for tuning HMC diagonal scale matrices compared.
Study on random matrices in deep neural networks using Gaussian data.
Given two data matrices and , sparse canonical correlation analysis (SCCA) is to seek two sparse canonical vectors and to maximize the correlation between and . However, classical and sparse CCA models consider the contribution of all the samples of data matrices and thus cannot identify an unde…
Leverage score sampling provides an appealing way to perform approximate computations for large matrices. Indeed, it allows to derive faithful approximations with a complexity adapted to the problem at hand. Yet, performing leverage scores sampling is a challenge in its own right requiring further approximations. In th…
Improved bounds for sensitivity sampling reducing the sample complexity for structured matrices.
Paper provides unbiased spectral moment estimates from finite data.
Stochastic gradient descent optimizes Nyström samples for kernel matrix approximation.
Correlation matrices are omnipresent in multivariate data analysis. When the number d of variables is large, the sample estimates of correlation matrices are typically noisy and conceal underlying dependence patterns. We consider the case when the variables can be grouped into K clusters with exchangeable dependence; t…
A new mechanism for differentially private Fréchet mean on SPD matrices.
Study on random matrices in deep neural networks with IID entries.
This paper focuses on the estimation of the sample covariance matrix from low-dimensional random projections of data known as compressive measurements. In particular, we present an unbiased estimator to extract the covariance structure from compressive measurements obtained by a general class of random projection matri…
Twisted Neumann--Zagier matrices for quantum invariants.
Improved statistical computation through efficient matrix sampling.
New method allows generating independent data matrices from summary statistics.
We consider the problem of selecting non-zero entries of a matrix in order to produce a sparse sketch of it, , that minimizes . For large matrices, such that (for example, representing observations over attributes) we give sampling distributions that exhibit four importa…
The paper proves local laws for non-separable sample covariance matrices.
Advances in molecular "omics'" technologies have motivated new methodology for the integration of multiple sources of high-content biomedical data. However, most statistical methods for integrating multiple data matrices only consider data shared vertically (one cohort on multiple platforms) or horizontally (different …
Paper proposes a deep learning method for better covariance matrix forecasting.
Proposes BONMI for integrating noisy matrices from multi-source data.
This paper presents a new method for estimating high dimensional covariance matrices. The method, permuted rank-penalized least-squares (PRLS), is based on a Kronecker product series expansion of the true covariance matrix. Assuming an i.i.d. Gaussian random sample, we establish high dimensional rates of convergence to…
Improved Bayesian regression for large datasets using multilevel Gibbs sampling.
New Bayesian matrix completion method using Stiefel manifolds.
Minimizing the nuclear norm of a matrix has been shown to be very efficient in reconstructing a low-rank sampled matrix. Furthermore, minimizing the sum of nuclear norms of matricizations of a tensor has been shown to be very efficient in recovering a low-Tucker-rank sampled tensor. In this paper, we propose to recover…
Rank-one measurements limit feasible sets for low-rank PSD matrices.
Better signal detection in undersampled data using joint and cross covariances.
Given a real matrix A with n columns, the problem is to approximate the Gram product AA^T by c << n weighted outer products of columns of A. Necessary and sufficient conditions for the exact computation of AA^T (in exact arithmetic) from c >= rank(A) columns depend on the right singular vector matrix of A. For a Monte-…