Study shows how anisotropic data affects learning dynamics in phase retrieval.
problem Understanding learning dynamics in phase retrieval with anisotropic Gaussian inputs.
method Developed a tractable reduction to reveal a three-phase trajectory and derived scaling laws.
result Found that anisotropy leads to a three-phase trajectory: fast escape, slow convergence, and spectral-tail learning.
Study examines risk premium convergence rates in risk sharing contracts.
problem Analyzing risk premium convergence rates in risk sharing contracts.
method Examines the limiting behavior of risk premium associated with Pareto optimal risk sharing contracts under general law-invariant risk measures.
result Risk premium convergence rate is typically n1/2, not n. Study convergence of simulated annealing in continuous and discrete settings.
problem Analyzing convergence rate of simulated annealing methods.
method Apply Eyring-Kramers law to prove polynomial decay of tail probabilities.
result Explicit rate of convergence for continuous and discrete simulated annealing.
Study on eigenvalue distribution of correlated time series deforming the semi-circle law.
problem Eigenvalue distribution of correlated time series differs from the semi-circle law.
method Analysis of Wigner random matrix with temporal correlation.
result Eigenvalue distribution converges to a deformed semi-circle law with longer tail and higher peak.
In this paper, we prove pointwise convergence of heat kernels for mGH-convergent sequences of RCD∗(K,N)-spaces. We obtain as a corollary results on the short-time behavior of the heat kernel in RCD∗(K,N)-spaces. We use then these results to initiate the study of Weyl's law in the RCD setting
Paper formulates EnKF as optimal transport problem for unique control law.
problem Unique control law for EnKF algorithms.
method Formulated as optimal transportation problem, derived explicit control law.
result Mean squared error converges to zero with finite particles.
We derive scaling laws for optimizing neural networks in hardware.
problem Optimizing the large parameter space of neural networks in hardware.
method Analytical derivation of scaling laws for Coordinate Descent optimization.
result Convergence is exponential and scales linearly with the number of neurons.
New bounds on neural network convergence using information theory.
problem Quantifying convergence rates of neural networks to Gaussian distributions.
method Entropic inequalities and Gaussian approximations.
result Improved convergence rates in various distances for neural networks.
The study uses the Merton model to estimate PD and finds a phase transition affecting convergence speed.
problem Estimating the probability of default (PD) using limited historical data.
method Adopted the Merton model and analyzed phase transitions in default correlation.
result PD estimation converges slowly when temporal correlation decays by power law less than one.
Study of Gaussian random fields on manifolds, focusing on their differential topology.
problem Understanding the differential topology of Gaussian random fields on manifolds.
method Systematic study using Gaussian measures and weak Whitney topology, focusing on convergence in law and transversality.
result The convergence in law of Gaussian random fields is related to the convergence of their covariance structures, with important technical tools like the Thom transversality theorem.
Develops mixed quantization for graph vector bundles.
problem Solving asymptotic spectral problems on graph vector bundles.
method Mixed quantization technique for graph vector bundles.
result Applications to various spectral problems.
LLMs learn peaked distributions slowly due to power-law losses.
problem Slow convergence of loss in training large language models.
method Systematic analysis of toy models and empirical evaluation of LLMs.
result Power-law time scaling with an exponent of 1/3 for learning peaked distributions.
New scaling laws explain deep learning performance growth.
problem Understanding neural network performance growth.
method Analyzed entire training dynamics of various architectures.
result Identified two dynamical scaling laws.
Study shows how heavy-tailed Hawkes processes can model rough volatility in financial markets.
problem Modeling rough volatility in financial markets with heavy-tailed Hawkes processes.
method Established weak convergence of Hawkes process with power-law kernel, derived scaling limit for financial market model.
result Price-volatility process converges weakly to a rough Heston model after rescaling.
This work analyzes neural scaling laws using power-law data spectra and derives analytical expressions for generalization error.
problem Understanding how neural network performance scales with key factors like data size and model complexity.
method Statistical mechanics techniques applied to one-pass stochastic gradient descent in a student-teacher framework.
result Derivation of analytical expressions for generalization error under power-law data spectra and identification of conditions for power-law scaling.
We construct non-symmetric diffusion processes associated with Dirichlet forms consisting of uniformly elliptic forms and derivation operators with killing terms on RCD spaces by aid of non-smooth differential structures introduced by Gigli '16. After constructing diffusions, we investigate conservativeness and the wea…
A universal learner achieves best rates for all distributions.
problem Improving learning algorithm rates under various settings.
method Simple extension of Levin's universal search.
result Achieves best-possible rates for all distributions.
Scaling laws for neural language models reveal optimal model size and compute allocation.
problem Understanding the optimal model size and compute allocation for neural language models.
method Empirical analysis of scaling laws for cross-entropy loss across model size, dataset size, and compute.
result Simple equations govern the dependence of overfitting and training speed on model/dataset size and model size, respectively.
Space exploration technology advances exponentially, consistent with Moore's and Wright's laws.
problem Predicting the advancement of space exploration technology.
method Analysis of Moore's and Wright's laws applied to space exploration technology.
result Spacecraft technology advances exponentially, consistent with Moore's and Wright's laws.
Paper approximates risk measures using SGD with Langevin dynamics.
problem Approximating arbitrary law invariant risk measures.
method Stochastic Gradient Langevin Dynamics (SGD-Langevin) for general risk measures.
result Non-asymptotic convergence rates of the approximation algorithm.
The paper investigates heavy-tailed behavior in offline SGD, showing it approximates power-law tails.
problem Understanding heavy-tailed behavior in offline (multi-pass) SGD with finite data.
method Proves nonasymptotic Wasserstein convergence bounds for offline SGD to online SGD.
result Offline SGD exhibits approximate power-law tails as the number of data points increases.
Optimizes sample reweighting to match laws under covariate shift using Wasserstein distance.
problem Matching laws of samples with different distributions under covariate shift.
method Minimizes Wasserstein distance between empirical measures of samples using Nearest Neighbors weights.
result Consistent reweighting leads to asymptotic convergence of empirical measures.
Random neural networks with ReLU activations are non-Gaussian processes.
problem Understanding the behavior of neural networks with random initialization and rectified linear units.
method Proving these networks are non-Gaussian processes and deriving their properties.
result These networks can converge to non-Gaussian processes under certain conditions.
Study nonlocal minimal surfaces for minimal surfaces in 3D manifolds.
problem Existence and regularity of minimal surfaces in 3D manifolds.
method Min-max methods, fractional perimeters, uniform estimates.
result Uniform estimates for min-max s-minimal surfaces in 3-manifolds, convergence to smooth minimal surfaces. We introduce a new statistical tool (the TP-statistic and TE-statistic) designed specifically to compare the behavior of the sample tail of distributions with power-law and exponential tails as a function of the lower threshold u. One important property of these statistics is that they converge to zero for power laws o…
VAV method optimizes learning rate for faster, stable SGD convergence.
problem Optimizing learning rate for efficient and stable machine learning models.
method Energy-based self-adaptive learning rate with auxiliary variable r. result VAV method achieves faster convergence and superior stability with larger learning rates.
Study laws of large numbers in online classification, determining optimal regret bounds.
problem Understanding how sequential sampling affects online learning and classification.
method Characterized online learnable classes and determined optimal regret bounds using Littlestone's dimension.
result Optimal regret bounds in online learning are determined, resolving open questions.
Model predicts neural network performance scaling laws across various factors.
problem Understanding the performance of neural networks across different training factors.
method Random feature model trained with gradient descent, analyzing compute-optimal scaling laws.
result Predicts asymmetric compute-optimal scaling rule and behavior of training and test loss gap.
The paper analyzes convergence rates for stochastic approximation and reinforcement learning.
problem Establishing almost sure convergence rates for stochastic approximation and reinforcement learning under Markovian noise.
method A novel Lyapunov drift construction that applies a Poisson-equation based correction for Markovian noise to the Moreau-envelope smoothing for contractive mappings.
result Almost sure convergence rates for specific learning rates are derived, with rates arbitrarily close to o(n1−2η) and o(n−1). Study approximates Plateau's laws using the Allen-Cahn equation.
problem Approximating Plateau's laws with the Allen-Cahn equation.
method Minimizing the Allen-Cahn energy under volume and spanning constraints.
result Energy minimizing solutions approximate Plateau-type singularities.
Paper finds optimal mini-batch size for SGD to speed up learning.
problem Optimizing mini-batch size for faster SGD convergence.
method Empirical inverse law and theoretical bound on mini-batch SGD training.
result An accurate model for predicting training time and identifying implications for algorithm and hardware.
SGD with constant stepsize converges to a non-Gaussian limit near flat minima.
problem Behavior of SGD near flat minima with convex objectives.
method Analyzes SGD with Markovian noise and contractive driving chain.
result Invariant law concentrates on scale α1/m and converges weakly to a non-Gaussian stationary distribution. Study on SGD dynamics and scaling laws for training quadratic neural networks in high dimensions.
problem Optimizing and understanding the training dynamics of quadratic neural networks in high-dimensional settings.
method Sharp analysis of SGD dynamics, combining matrix Riccati differential equations and matrix monotonicity arguments.
result Derivation of scaling laws for prediction risk, highlighting power-law dependencies on optimization time, sample size, and model width.
This paper addresses the statistical properties of time series driven by rational bubbles a la Blanchard and Watson (1982), corresponding to multiplicative maps, whose study has recently be revived recently in physics as a mechanism of intermittent dynamics generating power law distributions. Using insights on the beha…
Paper learns interaction laws from multiple trajectories of heterogeneous systems.
problem Estimating unknown interaction laws from multiple trajectories of heterogeneous systems.
method Nonparametric learning of interaction kernels based on pairwise distances, with convergence guarantees in L2 space. result Estimators converge at optimal min-max rate for 1-dimensional nonparametric regression.
The paper proves local laws for non-separable sample covariance matrices.
problem Analyzing non-separable sample covariance matrices with dependent or nonlinearly transformed data.
method Tensor network framework for analyzing fluctuation averaging in the presence of higher-order cumulant structure.
result Optimal averaged local law and full anisotropic local law for non-separable sample covariance matrices.
Self-balancing sampler improves sampling efficiency and unpredictability.
problem Efficient and unpredictable sampling in various applications.
method Adaptive biasing of sampling probabilities to achieve faster convergence and unpredictability.
result Self-balancing sampler converges at O(n−1) rate, outperforming IID sampling. Inspired by constructions in complex geometry we introduce a thermodynamic framework for Monge-Ampère equations on real tori. We show convergence in law of the associated point processes and explain connections to complex Monge-Ampère equations and optimal transport.
A new optimization algorithm improves convergence in unconstrained problems.
problem Unconstrained optimization problems.
method Element-wise relaxed scalar auxiliary variable (E-RSAV) algorithm.
result Improved convergence and alignment of modified and original energy.
The study examines the behavior of Gaussian processes' minimums and overshoots.
problem Understanding the behavior of Gaussian processes' minimums and overshoots.
method Analyzing conditional distributions and subsequential limits of minimizers.
result The scaled overshoot converges to an exponential random variable with mean σ_*^2.
Optimal trading strategy derived for nonlinear price impact models.
problem Optimal trading with nonlinear price impact induced by alpha signals.
method Variational approach, nonlinear Fredholm equation, iterative scheme.
result Existence and uniqueness of optimal trading strategy under monotonicity condition.
The paper analyzes Bayesian neural networks trained with VI, proving a law of large numbers for different schemes.
problem Training Bayesian neural networks with variational inference.
method Analyzes three training schemes: exact estimation, Bayes by Backprop, and Minimal VI.
result All training schemes converge to the same mean-field limit.
The study introduces a high-dimensional tail index model for viral post analysis.
problem Empirical observation of power-law distributions in viral posts.
method High-dimensional tail index regression model, regularized estimator, debiasing for inference.
result Consistency and asymptotic normality of debiased estimator.
Gradient flow method solves for optimal transport starting distributions.
problem Finding the optimal starting distribution for a martingale in optimal transport.
method Following the gradient flow of the Bass functional's L2-lift.
result Gradient flow converges to a minimizer of the Bass functional.
Study shows Direct Feedback Alignment fails to offer more efficient scaling than backpropagation.
problem Understanding and optimizing training methods for neural networks.
method Use of scaling laws to compare Direct Feedback Alignment (DFA) and backpropagation.
result DFA fails to offer more efficient scaling than backpropagation.
Polynomial networks converge to Gaussian processes at a rate of O(n^(-1/2)).
problem Understanding the convergence rate of polynomial networks to Gaussian processes.
method Examined one-hidden-layer neural networks with random weights, focusing on polynomial activations and their convergence rate in the 2-Wasserstein metric.
result The rate of convergence for polynomial networks to Gaussian processes is $O(n^{-rac{1}{2}})$.
The paper proves almost sure convergence of MCES algorithm for a specific class of MDPs.
problem Establishing convergence for the Monte Carlo Exploring Starts (MCES) algorithm in reinforcement learning.
method Introduced a novel inductive approach based on the strong law of large numbers.
result Almost sure convergence for Optimal Policy Feed-Forward MDPs.
Study on length distribution of random multicurves on large genus surfaces converging to Poisson-Dirichlet distribution.
problem Length statistics of random multicurves on large genus hyperbolic surfaces.
method Analytical proof of convergence to Poisson-Dirichlet distribution as genus tends to infinity.
result Mean lengths of the three longest components converge to specific percentages of total length as genus increases.