Work on SGDm under heavy-tailed noise, revealing its generalization properties.
problem Understanding generalization of SGDm under heavy-tailed noise.
method Analysis of continuous-time limit (SDE) and discrete-time SGDm, establishing generalization bounds.
result SGDm can have worse generalization in the presence of heavy-tailed noise for quadratic loss functions.
Econometric framework integrates heavy-tailed distributions with behavioral probability weighting for better asset pricing.
problem Underestimation of Value-at-Risk by traditional models in asset pricing.
method Developed an econometric framework combining heavy-tailed Student's t distributions with behavioral probability weighting. result Student's t specifications outperform Gaussian models in 88.4% of cases, reducing underestimation of Value-at-Risk by 16.5 percentage points. Researchers study heavy-tail properties of SGD using stochastic recurrence equations.
problem Analyzing heavy-tail properties of Stochastic Gradient Descent (SGD).
method Modeling SGD iterations as multivariate affine stochastic recursions and applying the theory of irreducible-proximal (i-p) matrices.
result Extended results of Gürbüzbalaban et al. (2020) by using the theory of i-p matrices.
Study tail behavior of sum of heavy-tailed risks with copulas.
problem Analyzing the tail behavior of sums of heavy-tailed risks with dependence modeled by copulas.
method Modeling dependence with copulas and analyzing tail asymptotics of sums of heavy-tailed risks.
result Obtained asymptotic expansions for Value-at-Risk of aggregate risk.
Paper develops heavy-tailed embeddings for better text classification and augmentation.
problem Improving text classification, especially for extreme values.
method Develops heavy-tailed embeddings using multivariate extreme value theory and introduces a scale-invariant classifier.
result The classifier outperforms baselines and generates meaningful augmented text.
New bounds for heavy-tailed SDEs without info-theory terms.
problem Understanding generalization of heavy-tailed stochastic optimization.
method Fractional Fokker-Planck equation to estimate entropy flows.
result High-probability bounds with better dimension dependence.
New bounds for SGD generalize without mutual information terms.
problem Generalizing SGD's learning dynamics for heavy-tailed distributions.
method Introducing a geometric decoupling term and bounding it computably.
result Proved generalization bounds without mutual information terms.
The paper examines how heavy-tailed risks behave under Gaussian copula models.
problem Understanding tail risk probabilities with heavy-tailed marginal risks and Gaussian dependence.
method Modeling heavy-tailed risks using regular variation and analyzing tail probabilities under Gaussian copula.
result The rate of decay of tail set probabilities varies with the type of tail sets and Gaussian correlation matrix.
We propose a new heavy-tailed distribution --- Gaussian-Chain (GC) distribution, which is inspirited by the hierarchical structures prevailing in social organizations. We determine the mean, variance and kurtosis of the Gaussian-Chain distribution to show its heavy-tailed property, and compute the tail distribution tab…
Investigates spectral properties of neural networks, showing invariance under certain conditions.
problem Understanding the spectral evolution and invariance in linear-width neural networks.
method Empirical and theoretical analysis of spectra of weight matrices in high-dimensional settings.
result Spectra of weight matrices are invariant under certain training conditions, with implications for feature learning.
Study reveals heavy-tailed behavior in training ReLU gates.
problem Understanding heavy-tailed distribution in stochastic deep learning.
method Experimental study of heavy-tail index for S.G.D. and a variant.
result Two algorithms exhibit similar heavy-tail behavior on ReLU data.
We consider strictly stationary heavy tailed time series whose finite-dimensional exponent measures are concentrated on axes, and hence their extremal properties cannot be tackled using classical multivariate regular variation that is suitable for time series with extremal dependence. We recover relevant information ab…
Efficiently estimates sparse mean from heavy-tailed data.
problem Robustly estimating sparse mean from heavy-tailed distributions.
method Stability-based approach adapted for heavy-tailed data.
result Optimal sample complexity with logarithmic dependence on dimension.
We examine the performance of six estimators of the power-law cross-correlations -- the detrended cross-correlation analysis, the detrending moving-average cross-correlation analysis, the height cross-correlation analysis, the averaged periodogram estimator, the cross-periodogram estimator and the local cross-Whittle e…
Copula-based normalizing flows improve flexibility and stability for heavy-tailed data.
problem Limited expressive power of vanilla normalizing flows.
method Generalize base distribution to copula for more accurate representation of target distribution.
result Copula-based normalizing flows improve flexibility, stability, and effectiveness for heavy-tailed data.
Paper examines the structure of stochastic gradients in deep learning.
problem Exploring the structure and heavy tails of stochastic gradients in deep learning.
method Conducted formal statistical tests on stochastic gradients and gradient noise.
result Stochastic gradients and gradient noise do not exhibit power-law heavy tails, but their covariance spectra do.
Study asymptotic properties of generalized shortfall risk measures for heavy-tailed risks.
problem Understanding risk measures for heavy-tailed risks.
method Derive asymptotic expansions for generalized shortfall risk measures.
result Unified theory for risk measures including distortion and utility-based measures.
In this paper, we show how the sampling properties of the Hurst exponent methods of estimation change with the presence of heavy tails. We run extensive Monte Carlo simulations to find out how rescaled range analysis (R/S), multifractal detrended fluctuation analysis (MF-DFA), detrending moving average (DMA) and genera…
Studied how heavy-tailed behavior affects SGD's generalization in quadratic optimization.
problem Link between heavy-tailed behavior and generalization in SGD.
method Used heavy-tailed stochastic differential equation and proved stability bounds.
result Stability of SGD depends on the loss function's tail behavior.
Study examines robust regression in high dimensions with heavy-tailed data.
problem Analyzing robust regression in high-dimensional settings with heavy-tailed data.
method Sharp asymptotic characterisation of M-estimators and ridge regression in elliptical distributions.
result Ridge regression is optimal and universal for finite second moments but can decay faster without them.
This work analyzes CVaR under heavy-tailed data, providing generalization and robustness bounds.
problem Understanding CVaR's behavior under heavy-tailed data and rare high-impact losses.
method Learning-theoretic analysis of CVaR-based empirical risk minimization.
result Sharp, high-probability generalization and excess risk bounds under minimal moment assumptions.
Value-at-Risk can be superadditive for sufficiently heavy-tailed losses.
problem Value-at-Risk (VaR) subadditivity failure
method Random vector perspective
result Universal Value-at-Risk superadditivity (UVS)
Study uses detrended cross-correlation to analyze cryptocurrency market, revealing robust collective modes and distinguishing interdependencies.
problem Nonstationarity, long-range memory, and heavy-tailed fluctuations obscure traditional correlations in complex systems.
method Constructs detrended correlation matrices using multifractal detrended cross-correlation coefficient ρr to emphasize different fluctuations. result Detrending and fluctuation analysis reveal distinct spectral properties from random case, identifying market and sectoral components.
This study benchmarks likelihood-free inference methods for models with heavy-tailed or discrete data.
problem Comparing likelihood-free inference methods for models with structural features like heavy-tails or discreteness.
method Four approaches: MLE, NBE, EOT, and AW-NBE are evaluated using simulations.
result The choice of evaluation tools is crucial for models with extremes and discrete data.
New robust estimator improves variable selection and coefficient estimation in linear regression with heavy-tailed errors and outliers.
problem Heavy-tailed errors and anomalous predictors in high-dimensional regression.
method Adaptive PENSE estimator for robust variable selection and estimation.
result Adaptive PENSE estimator provides reliable results even under very heavy-tailed errors and aberrant predictors.
Proposes a method to model financial returns with extreme shocks using flexible tail transformations.
problem Capturing extreme shocks in financial return data.
method Introduces a transformation layer in normalizing flows to model heavy-tailed distributions.
result Trained models can generate synthetic sets of extreme returns.
New theory explains why normalization is preferred in SGD under heavy-tailed noise.
problem Understanding why normalization is preferred in stochastic gradient descent (SGD) under heavy-tailed noise.
method Developed a worst-case complexity theory for stochastically preconditioned SGD and its variants.
result Normalization guarantees convergence at optimal rates, while clipping may fail in the worst case.
DE-SGD shows heavy-tailed behavior in decentralized settings.
problem Heavy-tailed behavior in decentralized SGD.
method Analyzes the emergence of heavy-tails in DE-SGD, considering both quadratic and twice continuously differentiable strongly convex loss functions.
result DE-SGD exhibits heavier tails than centralized SGD, and tail behavior depends on network parameters.
TSLiNGAM improves causal discovery in heavy-tailed data.
problem Identifying causal relationships in data with heavy tails.
method Combines DAGs with structural causal models, leveraging non-Gaussian noise.
result Significantly better performance on heavy-tailed and skewed data.
New MCMC methods map high-dimensional problems to spheres for better mixing.
problem Mixing issues in high-dimensional distributions, especially heavy-tailed ones.
method Stereographic Markov Chain Monte Carlo (MCMC) methods that map high-dimensional problems to spheres.
result Uniformly ergodic samplers for various distributions, including heavy-tailed ones, with faster convergence in higher dimensions.
New method approximates CVaR with less data for heavy-tailed risks.
problem Lack of data for accurate CVaR approximation in heavy-tailed distributions.
method Importance sampling based extrapolation for heavy-tailed distributions.
result Statistically consistent approximations with reduced data requirements.
The paper studies quantile contributions and their relationship with order statistics in heavy-tailed distributions.
problem Challenges of classical statistical models in heavy-tailed distributions.
method Theoretical study of quantile contribution statistic and its relationship with order statistics. Derivation of closed-form expression for joint CDF of order statistics and quantile contributions.
result Established asymptotic normality of quantile contributions and characterized their limiting distribution.
Introduces Polar Depth for analyzing multivariate heavy-tailed data extremes.
problem Analyzing the behavior of extremes from multivariate heavy-tailed distributions.
method Introduces Polar Depth, a novel statistical depth function expressed in polar coordinates.
result The polar depth of the largest observations converges to the polar depth of the limiting distribution as the threshold increases.
Study improves ERM for heavy-tailed data with dependent inputs.
problem Empirical Risk Minimization with dependent and heavy-tailed data.
method Extending risk bounds for ERM with heavy-tailed, dependent data.
result Established risk bounds for ERM with dependent and heavy-tailed data.
Network embedding aims to learn the low-dimensional representations of vertexes in a network, while structure and inherent properties of the network is preserved. Existing network embedding works primarily focus on preserving the microscopic structure, such as the first- and second-order proximity of vertexes, while th…
Extended univariate Range Value-at-Risk to multivariate settings.
problem Inability of traditional risk measures for heavy-tail distributions and infinite tail expectations.
method Multivariate definitions of robust truncated tail expectations, robustness and properties derived, closed-form expressions and special cases discussed.
result Empirical estimators accuracy examined through numerical and graphical examples.
Efficiently estimates sparse linear regression with heavy-tailed and outlier-contaminated data.
problem Estimating sparse linear regression coefficients with heavy-tailed and outlier-contaminated data.
method Efficient computation of estimators with sharp error bounds.
result Sharp error bounds for efficient estimators.
Study on price fluctuations in NFT market, showing heavy-tailed distributions and long-range memory.
problem Characterizing price fluctuations in NFT market.
method Analysis of capitalization, floor price, transactions, inter-transaction times, and volume value of NFTs.
result NFT market exhibits heavy-tailed probability distribution functions, well described by stretched exponentials, with long-range memory.
Is AdamW effective under heavy-tailed noise?
problem Stochastic gradient noise in LLM pretraining is typically heavy-tailed.
method Formulate as an open problem, prove a positive weighted-metric benchmark, and give a corridor lower-bound mechanism.
result No rigorous convergence theory for AdamW established in heavy-tailed regime.
Self-regulating annealing improves sampling from heavy-tailed datasets.
problem Sampling from heavy-tailed distributions using diffusion models.
method Proposed an SDE-based sampler with a state-dependent diffusion coefficient.
result State dependence induces a self-regulating annealing mechanism.
New algorithm robustly optimizes data streams with heavy-tailed or infinite variance samples.
problem Optimizing data streams with heavy-tailed or infinite variance samples.
method Gradient quantile clipping for SGD, leveraging Markov chain connections.
result Algorithm converges to a concentrated distribution with high probability bounds.
New diffusion models capture heavy-tailed distributions better.
problem Diffusion models struggle with rare or extreme events in heavy-tailed distributions.
method Repurposed diffusion framework using multivariate Student-t distributions, tailored perturbation kernel, and γ-divergence. result Our models generate rare and extreme events more effectively than standard diffusion models.
Stochastic gradient descent with momentum (SGDm) is one of the most popular optimization algorithms in deep learning. While there is a rich theory of SGDm for convex problems, the theory is considerably less developed in the context of deep learning where the problem is non-convex and the gradient noise might exhibit a…
New study reveals how heavy-tailed SGD dynamics lead to compressible neural networks.
problem Understanding why large neural networks can be compressed effectively.
method Linking SGD dynamics to compressibility properties of neural networks.
result Large step-size/batch-size ratios and overparametrization lead to heavy-tailed SGD dynamics, making networks compressible.
New concentration inequalities for tensors with heavy-tailed coefficients.
problem Developing bounds for Euclidean functions of tensors with sub-Weibull distributions.
method Extending concentration inequalities to sub-Weibull random tensors, using new inequalities for heavy-tailed random variables and martingale analysis.
result Established a phase transition between sub-gaussian and heavy-tailed regimes for Euclidean functions of tensors.
New models explain heavy-tailed behavior in neural networks.
problem Heavy-tailed spectral densities in neural networks.
method High-temperature Marchenko-Pastur (HTMP) ensemble models.
result Heavy-tailed behavior arises from three factors: data structure, training temperature, and eigenvector entropy.
New PAC-Bayes bounds for heavy-tailed losses using supermartingales.
problem Extending PAC-Bayes bounds to heavy-tailed losses.
method Using supermartingales and bounded variance assumption.
result PAC-Bayes generalization bounds for heavy-tailed losses.
Proves generalization bounds for SGD using Feller processes and Hausdorff dimension.
problem Characterizing generalization properties of SGD in deep learning.
method Proves generalization bounds for SGD under Feller process approximation, linking generalization error to the Hausdorff dimension of trajectories.
result Generalization error controlled by the Hausdorff dimension of trajectories, which is linked to the tail behavior of the driving process.