Iterative method 'Concent' corrects spectrum bias in covariance matrices.
problem Consistent bias in the spectrum of covariance matrices.
method 'Concent' iterative algorithm.
result Corrects spectrum bias for small and moderate dimensions.
Using Random Matrix Theory one can derive exact relations between the eigenvalue spectrum of the covariance matrix and the eigenvalue spectrum of its estimator (experimentally measured correlation matrix). These relations will be used to analyze a particular case of the correlations in financial series and to show that…
New outlier detection method using graph Laplacian spectrum boosts performance.
problem Detecting outliers in large datasets efficiently.
method Boosted outlier detection based on graph Laplacian spectrum.
result Outperforms existing methods on synthetic datasets.
A new method for unfolding histograms without matrix inversion.
problem Matrix inversion in experimental physics, especially in high-energy particle physics.
method Sampling many distributions, folding them through the response matrix, and choosing the closest one to the data.
result Performs as well as traditional methods in well-defined inverse problems and outperforms them in ill-defined ones.
We use topological methods to prove a semicontinuity property of the Hodge spectra for analytic germs defined on an isolated surface singularity. For this we introduce an analogue of the Seifert matrix (the fractured Seifert matrix), and of the Levine--Tristram signatures associated with it, defined for null-homologous…
Constructs coordinates to diagonalize Toda flow on matrices with simple spectrum.
problem Diagonalizing the Toda flow on matrices with simple spectrum.
method Lie theoretic methods applied to complex semisimple Lie algebras and their real forms.
result Decouples the Toda vector field into simpler components.
Proofs high-dimensional spectrum convergence of weighted sample covariance.
problem High-dimensional spectrum convergence of weighted sample covariance.
method Proposes a new, concise proof with stronger assumptions.
result Spectrum convergence proven for different weight distributions.
Study detects signal in financial stock correlations using phase-ordering kinetics.
problem Detecting meaningful signals in financial stock return correlations.
method Stochastic field theory model to establish a detection threshold.
result Detection of a signal in the largest eigenvalues of the stock return correlation matrix.
Muon replaces matrix gradient with polar factor, optimizing flat spectrum updates
problem Optimization bias in matrix updates
method Using polar factor of gradient
result Muon update maximizes entropy among bounded updates
New method estimates large matrices' spectra from small sub-matrices.
problem Estimating large matrices' spectra when full matrix-vector products are not available.
method Free decompression based on free probability theory.
result Estimates eigenspectrum of impalpable matrices from small sub-matrices.
Analyzes Hessian spectrum for neural networks near optimal learning.
problem Understanding learning dynamics near optimal points in neural networks.
method Characterizes Hessian eigenspectrum for teacher-student problems, using analytical and numerical methods.
result The rank of the Hessian matrix determines effective number of parameters for non-linear networks.
Study reveals an equivalence principle for the spectrum of random inner-product kernel matrices in polynomial scaling.
problem Understanding the spectrum of random kernel matrices in polynomial scaling regimes.
method Investigates random matrices with nonlinear kernel functions applied to inner products of uniformly distributed vectors.
result The spectrum of the random kernel matrix is asymptotically equivalent to a simpler matrix model through free additive convolution.
Study improves Hayashi-Yoshida estimator for high-dimensional stock covolatility.
problem Inconsistent performance of Hayashi-Yoshida estimator in high dimensions.
method Analyzed the limiting spectral distribution of the Hayashi-Yoshida estimator.
result Established the connection between the estimator's spectrum and the true covariance matrix in high dimensions.
Study on neural networks with non-normal interactions reveals unique spectral properties.
problem Understanding episodic memory encoding in the brain.
method Developed a neural network model with non-Hermitian couplings and applied random matrix theory.
result Spectral density of the model is non-uniform and can transition to chaos, providing computational benefits.
Study finds Calabi-Yau models' operator spectra match random matrix theory.
problem Understanding spectra of Calabi-Yau sigma models.
method Numerical methods for Ricci-flat metrics, averaging over complex structure moduli space.
result Spectrum matches Gaussian orthogonal ensemble of random matrix theory.
We use methods of random matrix theory to analyze the cross-correlation matrix C of price changes of the largest 1000 US stocks for the 2-year period 1994-95. We find that the statistics of most of the eigenvalues in the spectrum of C agree with the predictions of random matrix theory, but there are deviations for a fe…
This paper speeds up spectral clustering for large graphs by dilating their eigenspectrum.
problem Slow convergence in spectral clustering due to small eigengaps in graph Laplacians.
method Polynomial approximations to matrix operations that dilate the spectrum without changing eigenvectors.
result Significant acceleration of convergence in spectral clustering.
Muon outperforms GD in associative memory learning by balancing frequency components.
problem Training dynamics and scaling behavior of Muon in associative memory learning.
method Study of Muon in a linear associative memory model with softmax retrieval and hierarchical frequency spectrum over query-answer pairs.
result Muon achieves exponential speedup over GD in noiseless case and superior scaling efficiency in noisy case.
Particles representing tokens cluster in Transformers, influenced by initial tokens and matrix spectrum.
problem Understanding the geometry of learned representations in Transformers.
method Viewing Transformers as particle systems, applying dynamical systems and partial differential equations.
result Particles cluster towards limiting objects, confirming context-awareness and the emergence of leaders.
Kernel method is a very powerful tool in machine learning. The trick of kernel has been effectively and extensively applied in many areas of machine learning, such as support vector machine (SVM) and kernel principal component analysis (kernel PCA). Kernel trick is to define a kernel function which relies on the inner-…
The exact meaning of the noise spectrum of eigenvalues of the covariance matrix is discussed. In order to better understand the possible phenomena behind the observed noise, the spectrum of eigenvalues of the covariance matrix is studied under a model where most of the true eigenvalues are zero and the parameters are n…
Random matrix theory explains how neural networks adapt to data.
problem Understanding how neural networks learn and generalize from data.
method Random matrix analysis of two-layer neural networks.
result Sharp characterization of feature spectrum and generalization error.
Study conic Laplacian on \(\mb P^1\) with explicit model and boundary data.
problem Modeling conic Laplacian on \(\mb P^1\) with specific boundary conditions.
method Fourier decomposition, Legendre equations, gluing map, Friedrichs spectrum, Weyl function.
result Explicit computation of eigenfunctions and \(S\)-matrix.
Pion optimizes LLMs by preserving weight matrix singular values.
problem Training large language models (LLMs) with standard optimizers leads to unstable weight matrices.
method Pion uses orthogonal transformations to update weight matrices, preserving their singular values.
result Pion offers a stable alternative to standard optimizers for LLM pretraining and finetuning.
We propose a new method of learning a sparse nonnegative-definite target matrix. Our primary example of the target matrix is the inverse of a population covariance or correlation matrix. The algorithm first estimates each column of the target matrix by the scaled Lasso and then adjusts the matrix estimator to be symmet…
Mamba struggles with long context lengths, but spectrum scaling improves performance.
problem Mamba's performance degrades with increasing context length.
method Spectrum scaling applied to pre-trained Mamba models to improve long-context generalization.
result Spectrum scaling significantly improves performance in long-context settings.
The paper studies the loss landscape of regularized deep matrix factorization, revealing unique and sharp minimizers.
problem Understanding the loss landscape and minimizers of regularized deep matrix factorization problems.
method Theoretical analysis of ℓ2-regularized deep matrix factorization/deep linear network training problems with squared-error loss. result The unique end-to-end minimizer exists for all target matrices except for a set of Lebesgue measure zero.
This paper analyzes generalization for linear models with spiked covariance structures.
problem Understanding the generalization performance of linear models with spiked covariance structures.
method Derives the generalization error for two simple models with spiked covariances using random matrix theory.
result The eigenvector and eigenvalue corresponding to the spike significantly influence the generalization error.
Derives adjoint formulas for matrix operations and applies them to specific cases.
problem Computing adjoints for matrix operations and specific matrix types.
method Derives adjoint formulas for matrix operations and applies them to specific cases.
result Closed-form expressions for adjoints in specific matrix types.
We analyze kernel matrices in polynomial high-dimensional settings and explain double descent in KRR.
problem Understanding the spectrum of kernel matrices in polynomial high-dimensional settings and its implications for KRR risk.
method Generalized decomposition of kernel matrices into low-rank spike matrix, identity, and Gegenbauer matrix.
result The test error in KRR can exhibit double descent behavior, depending on effective regularization and signal-to-noise ratio.
Singular values of a data in a matrix form provide insights on the structure of the data, the effective dimensionality, and the choice of hyper-parameters on higher-level data analysis tools. However, in many practical applications such as collaborative filtering and network analysis, we only get a partial observation.…
A parameterization that is a modified version of a previous work is proposed for the returns and correlation matrix of financial time series and its properties are studied. This parameterization allows easy introduction of non-stationarity and it shows several of the characteristics of the true, observed realizations, …
A new measure of model complexity based on Fisher Information.
problem Model complexity measurement in statistical models.
method Effective dimension defined by the number of cubes needed to cover the model space.
result The effective dimension is scale-dependent and measures model complexity.
In this paper we describe market in projective geometry language and give definition of a matrix of market rate, which is related to the matrix rate of return and the matrix of judgements in the Analytic Hierarchy Process (AHP). We use these observations to extend the AHP model to projective geometry formalism and gene…
We describe a method to determine the eigenvalue density of empirical covariance matrix in the presence of correlations between samples. This is a straightforward generalization of the method developed earlier by the authors for uncorrelated samples. The method allows for exact determination of the experimental spectru…
We develop the scattering theory of general conformally compact metrics. For low frequencies, the domain of the scattering matrix is shown to be frequency dependent. In particular, generalized eigenfunctions exhibit L^2 decay in directions where the asymptotic curvature is sufficiently negative. The scattering matrix i…
Heavy-tailed regularization improves deep neural network performance.
problem Improving generalization of deep neural networks.
method Introducing Heavy-Tailed Regularization, using differentiable penalty terms and Bayesian statistics.
result Heavy-tailed regularization outperforms conventional regularization techniques.
OMD monitors stock market dynamics through matrix trajectories, revealing crisis patterns and sector rotations.
problem Understanding and predicting stock market dynamics during crises.
method Applying OMD to S&P 500 returns over three crises, analyzing distance matrices and their spectra.
result Market dynamics show coherent changes during crises, with sector-specific patterns and volatility clustering.
We discuss the applications of Random Matrix Theory in the context of financial markets and econometric models, a topic about which a considerable number of papers have been devoted to in the last decade. This mini-review is intended to guide the reader through various theoretical results (the Marcenko-Pastur spectrum …
This paper concerns the problem of matrix completion, which is to estimate a matrix from observations in a small subset of indices. We propose a calibrated spectrum elastic net method with a sum of the nuclear and Frobenius penalties and develop an iterative algorithm to solve the convex minimization problem. The itera…
The paper examines how gradient descent stabilizes low-rank matrix factorization in noisy conditions.
problem Stability of low-rank implicit regularization in perturbed deep matrix factorization.
method Derives spectral conditions for gradient descent to exhibit a low-rank phase in noiseless settings and analyzes perturbed dynamics.
result Gradient descent converges to a low-rank solution under perturbation, with explicit dependence on perturbation size.
A new debiasing method for high-dimensional regression with applications to PCR.
problem Debiasing in high-dimensional statistics with i.i.d. samples and sub-Gaussian covariates.
method Spectrum-Aware Debiasing using rescaled gradient descent with spectral information.
result Achieves debiasing in broader contexts with structured dependencies, heavy tails, and low-rank structures.
We consider deep classifying neural networks. We expose a structure in the derivative of the logits with respect to the parameters of the model, which is used to explain the existence of outliers in the spectrum of the Hessian. Previous works decomposed the Hessian into two components, attributing the outliers to one o…
We introduce a covariance matrix estimator that both takes into account the heteroskedasticity of financial returns (by using an exponentially weighted moving average) and reduces the effective dimensionality of the estimation (and hence measurement noise) via techniques borrowed from random matrix theory. We calculate…
OMD monitors stock market dynamics through matrix trajectories and reveals crisis patterns.
problem Understanding and predicting stock market crises and sector rotations.
method Applying OMD to S\&P 500 returns over three crises, analyzing distance matrices and their spectra.
result Market dynamics show coherent changes during crises, with distinct sector leadership.
New method corrects missing data bias in dimension reduction.
problem Missing data complicates high-dimensional data analysis.
method Developed a bias-corrected Gram matrix for heterogeneous missingness.
result Proposed method improves dimension reduction techniques significantly.
We analyse the structure of the distribution of eigenvalues of the stock market correlation matrix with increasing length of the time series representing the price changes. We use 100 highly-capitalized stocks from the American market and relate result to the corresponding ensemble of Wishart random matrices. It turns …
CSTs improve stability in covariance spectrum analysis without training.
problem Stability and expressiveness in covariance spectrum analysis.
method Sequential application of covariance wavelet filters to input data.
result Stable and expressive hierarchical representations in low-data settings.