This work compresses heavy-tailed weight matrices for tighter generalization bounds.
problem Empirical evidence linking heavy-tailed weight matrices to test set accuracy but lack of formal relationship with generalization bounds.
method Utilized the compression framework to show that heavy-tailed matrices can be compressed, resulting in sparse weight matrices.
result Demonstrated a non-vacuous generalization bound for compressed networks with heavy-tailed weight matrices.
Heavy-tailed regularization improves deep neural network performance.
problem Improving generalization of deep neural networks.
method Introducing Heavy-Tailed Regularization, using differentiable penalty terms and Bayesian statistics.
result Heavy-tailed regularization outperforms conventional regularization techniques.
Study heavy-tailed weights' impact on neural network's spectral distribution.
problem Analyzing spectral distribution of conjugate kernel matrices with heavy-tailed weights.
method Computed limiting eigenvalue distribution through moments, considering heavy-tailed distributions and nonlinear activation functions.
result Heavy-tailed weights induce strong correlations, leading to fundamentally different spectral behavior.
The difficulty of classification affects the weight matrices' heavy tail appearance in deep learning networks.
problem Understanding the spectral properties of weight matrices in deep learning networks.
method Spectral analysis of weight matrices in different modules of DNNs, classification difficulty as a driving factor for heavy tail appearance.
result Higher classification difficulty leads to more frequent appearance of heavy tails in weight matrices spectra.
Study connects covariance cleaning theory to information theory for heavy-tailed distributions.
problem Optimizing covariance matrices for heavy-tailed distributions using information theory.
method Minimizing Frobenius norm and information loss between true and estimated covariance matrices.
result Asymptotic regime of large matrices minimizes information loss for Student's t distributions.
Researchers study heavy-tail properties of SGD using stochastic recurrence equations.
problem Analyzing heavy-tail properties of Stochastic Gradient Descent (SGD).
method Modeling SGD iterations as multivariate affine stochastic recursions and applying the theory of irreducible-proximal (i-p) matrices.
result Extended results of Gürbüzbalaban et al. (2020) by using the theory of i-p matrices.
RMT reveals self-regularization in neural networks, including traditional and heavy-tailed forms.
problem Understanding and quantifying self-regularization in neural networks.
method Application of Random Matrix Theory to analyze weight matrices of various neural network models.
result Identification of 5+1 phases of training in neural networks, including traditional and heavy-tailed self-regularization.
New convergence rates for SGD under heavy-tailed noise with infinite variance.
problem Convergence analysis of SGD under heavy-tailed noise with infinite variance.
method Identifying a condition on the Hessian and providing a convergence rate for the distance to the global optimum.
result SGD can converge to the global optimum under heavy-tailed noise with infinite variance.
Study uses detrended cross-correlation to analyze cryptocurrency market, revealing robust collective modes and distinguishing interdependencies.
problem Nonstationarity, long-range memory, and heavy-tailed fluctuations obscure traditional correlations in complex systems.
method Constructs detrended correlation matrices using multifractal detrended cross-correlation coefficient ρ r ρ_r ρ r to emphasize different fluctuations. result Detrending and fluctuation analysis reveal distinct spectral properties from random case, identifying market and sectoral components.
New models explain heavy-tailed behavior in neural networks.
problem Heavy-tailed spectral densities in neural networks.
method High-temperature Marchenko-Pastur (HTMP) ensemble models.
result Heavy-tailed behavior arises from three factors: data structure, training temperature, and eigenvector entropy.
New clustering method handles heavy-tailed, asymmetric data.
problem Robust clustering of high-dimensional, heavy-tailed data.
method Sparse mixture of generalized hyperbolic distributions with gamma-lasso penalty.
result Improved clustering performance on heavy-tailed data.
Efficiently estimates sparse mean from heavy-tailed data.
problem Robustly estimating sparse mean from heavy-tailed distributions.
method Stability-based approach adapted for heavy-tailed data.
result Optimal sample complexity with logarithmic dependence on dimension.
New theory predicts which large DNNs will have best test accuracy.
problem Predicting which large pre-trained DNNs will have the best test accuracy.
method Heavy-Tailed Self-Regularization (HT-SR) and Universal capacity control metric based on power law exponents.
result Universal capacity control metric correlates well with reported test accuracies of large-scale DNNs.
Deep neural networks implicitly self-regularize, shown by random matrix theory.
problem Understanding and quantifying implicit self-regularization in deep neural networks.
method Random Matrix Theory applied to various DNNs, including pre-trained and self-trained models.
result DNN training implicitly implements self-regularization, observable in empirical spectral densities.
Muon optimizes Transformer training with heavy-tailed data, achieving optimal sample complexity.
problem Theoretical understanding of non-Euclidean optimisation methods for heavy-tailed data in training Transformers.
method Addressing the gap in theoretical understanding, we show Muon achieves optimal sample complexity under heavy-tailed noise.
result Muon finds an ε-stationary point in nuclear norm with optimal sample complexity, absorbing heavy-tailed noise without dimension dependence.
New theory predicts deep neural networks can operate in an extended critical regime without fine-tuning.
problem Understanding the dynamics and computational principles of deep neural networks.
method Combining theories of heavy-tailed random matrices and non-equilibrium statistical physics.
result Deep neural networks can operate in an extended critical regime without fine-tuning parameters.
Study learns linear system dynamics from noisy bilinear data.
problem Learning linear dynamics from bilinear observations with process and measurement noise.
method Regression with Kronecker product design, data-dependent and independent error bounds.
result Upper bounds on statistical error rates and sample complexity for learning dynamics matrices.
Study on estimating rank-one tensors in noisy data with heavy tails.
problem Estimating rank-one spiked tensors in the presence of heavy tailed errors.
method Analysis of spectral norm of random tensors with iid entries.
result Signal strength requirements for optimal estimation are similar for heavy tailed and Gaussian noise, but vanish for noise with finite fourth moment.
New estimator for covariance of heavy-tailed data with affine-invariant bound.
problem Estimating covariance of heavy-tailed multivariate distributions.
method Affine-invariant bound for covariance estimation with fourth-order moment requirement.
result Proposed estimator S ^ \widehat{\mathbf{S}} S has an affine-invariant bound of ( 1 − ε ) S ≼ S ^ ≼ ( 1 + ε ) S (1-\varepsilon) \mathbf{S} \preccurlyeq \widehat{\mathbf{S}} \preccurlyeq (1+\varepsilon) \mathbf{S} ( 1 − ε ) S ≼ S ≼ ( 1 + ε ) S in high probability. Study LASSO for high-dimensional VAR models with weakly dependent innovations.
problem Understanding sparse regularization in high-dimensional VAR models with weakly dependent innovations.
method LASSO estimation for weakly sparse VAR models with heavy tailed innovations, under L 1 L^1 L 1 mixingale condition. result Oracle properties of LASSO estimation in high-dimensional VAR models with weakly dependent innovations.
New inequalities for matrix supermartingales converge under various conditions.
problem Convergence and maximal inequalities of supermartingales in positive semidefinite matrices.
method Developed new concentration inequalities for matrix supermartingales.
result New inequalities for matrix supermartingales under different tail conditions.
Investigates spectral properties of neural networks, showing invariance under certain conditions.
problem Understanding the spectral evolution and invariance in linear-width neural networks.
method Empirical and theoretical analysis of spectra of weight matrices in high-dimensional settings.
result Spectra of weight matrices are invariant under certain training conditions, with implications for feature learning.
The paper analyzes heavy-tailed multivariate distributions in non-stationary systems using random matrix theory.
problem Risk assessment for rare events in complex, non-stationary systems.
method Generalized scalar product between correlation matrices, model for non-stationary fluctuations.
result Formulae for multivariate distributions with reduced parameters, facilitating applications.
AlphaPruning optimizes LLM pruning using HT-SR theory for better performance.
problem Improving pruning of large language models to reduce size without sacrificing performance.
method AlphaPruning uses HT-SR theory to allocate layerwise sparsity ratios more theoretically.
result AlphaPruning prunes LLaMA-7B to 80% sparsity with reasonable perplexity.
We consider a Gaussian process formulation of the multiple kernel learning problem. The goal is to select the convex combination of kernel matrices that best explains the data and by doing so improve the generalisation on unseen data. Sparsity in the kernel weights is obtained by adopting a hierarchical Bayesian approa…
Paper proposes a new algorithm for graph learning with covariance constraints.
problem Graphical models and factor analysis not jointly leveraged in graph learning processes.
method Penalized maximum likelihood estimation of an elliptical distribution with Riemannian optimization.
result Effectiveness of the proposed approach demonstrated on real-world data sets.
The paper provides exact multivariate amplitude distributions for non-stationary Gaussian or algebraic fluctuations.
problem Capturing the statistical properties of fluctuating correlations in non-stationary systems.
method Developed a random matrix model to average multivariate amplitude distributions from short time scales to large time scales.
result Explicit multivariate distributions for non-stationary correlation systems are provided, capturing the degree of non-stationarity.
We review recent progress in modeling credit risk for correlated assets. We start from the Merton model which default events and losses are derived from the asset values at maturity. To estimate the time development of the asset values, the stock prices are used whose correlations have a strong impact on the loss distr…
Bayesian VAR and Elliptical Black-Litterman models improve portfolio optimization during regime changes and heavy-tailed returns.
problem Portfolio optimization under market regime changes and heavy-tailed returns.
method BAVAR-BLED algorithm combining BAVAR and Black-Litterman models with Elliptical Distributions.
result Significant outperformance of state-of-the-art methods in Sharpe, Sortino ratios, and total returns.
Modeling spatial extremes with non-Gaussian fields using SAR models and CNNs.
problem Challenges in modeling spatial data with heavy-tailed distributions and missing cells.
method Spatial autoregressive models with Generalized Extreme Value innovations, combined with CNN for fast parameter estimation.
result Effective modeling of spatial extremes in non-Gaussian fields, demonstrated on precipitation data.
New random feature maps for Laplacian and related kernels.
problem Challenges in approximating the Laplacian kernel and its generalizations.
method Developed random feature maps for Laplacian and related kernels, providing efficient sampling schemes.
result Demonstrated the efficacy of these random feature maps on real datasets.
We define a family of probability distributions for random count matrices with a potentially unbounded number of rows and columns. The three distributions we consider are derived from the gamma-Poisson, gamma-negative binomial, and beta-negative binomial processes. Because the models lead to closed-form Gibbs sampling …
Unified framework detects overfitting in crash classification models.
problem Evaluation metrics fail to detect overfitting in crash classification models.
method Random Matrix Theory and Heavy-Tailed Self-Regularization framework applied to various model types.
result Power-law exponent α reliably distinguishes well-regularized from overfit models.
Study improves ERM for heavy-tailed data with dependent inputs.
problem Empirical Risk Minimization with dependent and heavy-tailed data.
method Extending risk bounds for ERM with heavy-tailed, dependent data.
result Established risk bounds for ERM with dependent and heavy-tailed data.
Non-negative matrix factorization (NMF) approximates a non-negative matrix X X X by a product of two non-negative low-rank factor matrices W W W and H H H . NMF and its extensions minimize either the Kullback-Leibler divergence or the Euclidean distance between X X X and W T H W^T H W T H to model the Poisson noise or the Gaussian noise.…
Efficiently estimates sparse linear regression with heavy-tailed and outlier-contaminated data.
problem Estimating sparse linear regression coefficients with heavy-tailed and outlier-contaminated data.
method Efficient computation of estimators with sharp error bounds.
result Sharp error bounds for efficient estimators.
Is AdamW effective under heavy-tailed noise?
problem Stochastic gradient noise in LLM pretraining is typically heavy-tailed.
method Formulate as an open problem, prove a positive weighted-metric benchmark, and give a corridor lower-bound mechanism.
result No rigorous convergence theory for AdamW established in heavy-tailed regime.
Self-regulating annealing improves sampling from heavy-tailed datasets.
problem Sampling from heavy-tailed distributions using diffusion models.
method Proposed an SDE-based sampler with a state-dependent diffusion coefficient.
result State dependence induces a self-regulating annealing mechanism.
New diffusion models capture heavy-tailed distributions better.
problem Diffusion models struggle with rare or extreme events in heavy-tailed distributions.
method Repurposed diffusion framework using multivariate Student-t distributions, tailored perturbation kernel, and γ γ γ -divergence. result Our models generate rare and extreme events more effectively than standard diffusion models.
New concentration inequalities for tensors with heavy-tailed coefficients.
problem Developing bounds for Euclidean functions of tensors with sub-Weibull distributions.
method Extending concentration inequalities to sub-Weibull random tensors, using new inequalities for heavy-tailed random variables and martingale analysis.
result Established a phase transition between sub-gaussian and heavy-tailed regimes for Euclidean functions of tensors.
New PAC-Bayes bounds for heavy-tailed losses using supermartingales.
problem Extending PAC-Bayes bounds to heavy-tailed losses.
method Using supermartingales and bounded variance assumption.
result PAC-Bayes generalization bounds for heavy-tailed losses.
New bounds for heavy-tailed SDEs without info-theory terms.
problem Understanding generalization of heavy-tailed stochastic optimization.
method Fractional Fokker-Planck equation to estimate entropy flows.
result High-probability bounds with better dimension dependence.
Study on error probability for classification of heavy-tailed renewal processes.
problem Error probability in classification of heavy-tailed renewal processes.
method Asymptotic expressions for Bhattacharyya bound on misclassification error probabilities.
result Obtained asymptotic expressions for misclassification error probabilities.
TTF improves performance of normalizing flows for heavy-tailed distributions.
problem Improving performance of normalizing flows for heavy-tailed distributions.
method Uses a Gaussian base distribution and a final transformation layer to produce heavy tails.
result Experimental results show TTF outperforms current methods, especially in high-dimensional or heavy-tailed scenarios.
New sampling method for heavy-tailed distributions using Langevin Algorithm.
problem Sampling from heavy-tailed distributions efficiently.
method Transformed Unadjusted Langevin Algorithm on specific transformations.
result Polynomial-order oracle complexities for certain heavy-tailed densities.
Study tail behavior of sum of heavy-tailed risks with copulas.
problem Analyzing the tail behavior of sums of heavy-tailed risks with dependence modeled by copulas.
method Modeling dependence with copulas and analyzing tail asymptotics of sums of heavy-tailed risks.
result Obtained asymptotic expansions for Value-at-Risk of aggregate risk.
Survey on mean estimation and regression for heavy-tailed data.
problem Estimating mean and regression functions in heavy-tailed distributions.
method Sub-Gaussian mean estimators, median-of-means, trimmed mean, Catoni's estimator.
result Detailed proofs for estimators in heavy-tailed settings.
Heavy-tailed distributions emerge in SGD's parameter evolution.
problem Understanding heavy-tailed distributions in SGD parameter evolution.
method Continuous diffusion approximation of SGD (homogenized SGD) analysis.
result Explicit upper and lower bounds on tail-index of homogenized SGD.