Proofs high-dimensional spectrum convergence of weighted sample covariance.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Iterative method 'Concent' corrects spectrum bias in covariance matrices.
The salient properties of large empirical covariance and correlation matrices are studied for three datasets of size 54, 55 and 330. The covariance is defined as a simple cross product of the returns, with weights that decay logarithmically slowly. The key general properties of the covariance matrices are the following…
CSTs improve stability in covariance spectrum analysis without training.
WeSpeR speeds up non-linear shrinkage for high-dimensional weighted covariance.
Adaptive Bayesian model for covariate-dependent power spectra analysis.
Standard sparse pseudo-input approximations to the Gaussian process (GP) cannot handle complex functions well. Sparse spectrum alternatives attempt to answer this but are known to over-fit. We suggest the use of variational inference for the sparse spectrum approximation to avoid both issues. We model the covariance fu…
This paper analyzes generalization for linear models with spiked covariance structures.
The exact meaning of the noise spectrum of eigenvalues of the covariance matrix is discussed. In order to better understand the possible phenomena behind the observed noise, the spectrum of eigenvalues of the covariance matrix is studied under a model where most of the true eigenvalues are zero and the parameters are n…
Using Random Matrix Theory one can derive exact relations between the eigenvalue spectrum of the covariance matrix and the eigenvalue spectrum of its estimator (experimentally measured correlation matrix). These relations will be used to analyze a particular case of the correlations in financial series and to show that…
Study improves Hayashi-Yoshida estimator for high-dimensional stock covolatility.
Using Weitzenböck techniques on any compact Riemannian spin manifold we derive a general inequality depending on a real parameter and joining the spectrum of the Dirac operator with terms depending on the Ricci tensor and its first covariant derivatives. The discussion of this inequality yields vanishing theorems for t…
A new debiasing method for high-dimensional regression with applications to PCR.
We consider deep classifying neural networks. We expose a structure in the derivative of the logits with respect to the parameters of the model, which is used to explain the existence of outliers in the spectrum of the Hessian. Previous works decomposed the Hessian into two components, attributing the outliers to one o…
Kernel method is a very powerful tool in machine learning. The trick of kernel has been effectively and extensively applied in many areas of machine learning, such as support vector machine (SVM) and kernel principal component analysis (kernel PCA). Kernel trick is to define a kernel function which relies on the inner-…
Simple bounds for covariance and Gram matrices across various settings.
Robust and reliable covariance estimates play a decisive role in financial and many other applications. An important class of estimators is based on Factor models. Here, we show by extensive Monte Carlo simulations that covariance matrices derived from the statistical Factor Analysis model exhibit a systematic error, w…
We describe a method to determine the eigenvalue density of empirical covariance matrix in the presence of correlations between samples. This is a straightforward generalization of the method developed earlier by the authors for uncorrelated samples. The method allows for exact determination of the experimental spectru…
Power-law spectrum of random feature model is preserved in neural networks.
Study on KRR with power-law data, showing better sample complexity.
CDST improves ensemble prediction by adjusting model weights based on covariates.
We consider the problem of approximating the set of eigenvalues of the covariance matrix of a multivariate distribution (equivalently, the problem of approximating the "population spectrum"), given access to samples drawn from the distribution. The eigenvalues of the covariance of a distribution contain basic informati…
Modeling sequential data has become more and more important in practice. Some applications are autonomous driving, virtual sensors and weather forecasting. To model such systems, so called recurrent models are frequently used. In this paper we introduce several new Deep recurrent Gaussian process (DRGP) models based on…
Study precise sample covariance error for Gaussian centered data.
In dealing with high-dimensional data sets, factor models are often useful for dimension reduction. The estimation of factor models has been actively studied in various fields. In the first part of this paper, we present a new approach to estimate high-dimensional factor models, using the empirical spectral density of …
We propose a novel sparse spectrum approximation of Gaussian process (GP) tailored for Bayesian optimization. Whilst the current sparse spectrum methods provide desired approximations for regression problems, it is observed that this particular form of sparse approximations generates an overconfident GP, i.e. it produc…
New method estimates covariance matrices without restrictive assumptions.
Model improves covariance estimation from shared and distinct datasets.
We provide a unified analysis of the predictive risk of ridge regression and regularized discriminant analysis in a dense random effects model. We work in a high-dimensional asymptotic regime where and , and allow for arbitrary covariance among the features. For both metho…
Paper analyzes ensemble Kalman updates for effective dimension and localization.
The study establishes minimax bounds for estimating operators from noisy samples.
We introduce a covariance matrix estimator that both takes into account the heteroskedasticity of financial returns (by using an exponentially weighted moving average) and reduces the effective dimensionality of the estimation (and hence measurement noise) via techniques borrowed from random matrix theory. We calculate…
Least squares regression shows unexpected double descent in under-parameterized models.
Study on the geometric Dyson Brownian motion of non-square matrix products.
Hybrid ResNet and RMT improve covariance matrix estimation for cryptocurrency portfolios.
Bayesian method uses data spectra to estimate non-sparse high-dimensional models.
In this paper, we study the spectrum and the eigenvectors of radial kernels for mixtures of distributions in . Our approach focuses on high dimensions and relies solely on the concentration properties of the components in the mixture. We give several results describing of the structure of kernel matrices …
This paper shows universality in spectrum behavior for random inner-product kernel matrices in polynomial regime.
Scaling laws in linear regression explain model performance improvements with size and data.
The aim of this work is to build financial crisis indicators based on spectral properties of the dynamics of market data. After choosing an optimal size for a rolling window, the historical market data in this window is seen every trading day as a random matrix from which a covariance and a correlation matrix are obtai…
We provide a method to prepare covariance matrices for quantum datasets.
A neural network improves DOA estimation from a single snapshot.
We investigate if kernel regularization methods can achieve minimax convergence rates over a source condition regularity assumption for the target function. These questions have been considered in past literature, but only under specific assumptions about the decay, typically polynomial, of the spectrum of the the kern…
We obtain a tight distribution-specific characterization of the sample complexity of large-margin classification with L_2 regularization: We introduce the γ-adapted-dimension, which is a simple function of the spectrum of a distribution's covariance matrix, and show distribution-specific upper and lower bounds on the s…
New methods improve cross-correlation analysis of time series data.
Study eigenvalues and eigenvectors in neural networks, focusing on signal propagation.
In this paper we study the behaviour of the continuous spectrum of the Laplacian on a complete Riemannian manifold of bounded curvature under perturbations of the metric. The perturbations that we consider are such that its covariant derivatives up to some order decay with some rate in the geodesic distance from a fixe…
Improved scaling laws in linear regression using data reuse.