We found that factors decay over time, with momentum fitting best.
problem Understanding how factors decay over time and their impact on performance.
method Derived a hyperbolic decay model for factors, tested against linear and exponential alternatives.
result Momentum exhibits hyperbolic decay, outperforming linear and exponential models.
SAD-DPSGD improves model performance on imbalanced medical datasets like HAM10000.
problem Data leakage and imbalanced distribution in medical image classification datasets.
method SAD-DPSGD uses a linear decaying mechanism for noise and clipping thresholds to enhance performance.
result SAD-DPSGD outperforms Auto-DPSGD on HAM10000, improving accuracy by 2.15%.
Weight decay is one of the standard tricks in the neural network toolbox, but the reasons for its regularization effect are poorly understood, and recent results have cast doubt on the traditional interpretation in terms of L2 regularization. Literal weight decay has been shown to outperform L2 regularization for…
In this paper, we give a description for steady Ricci solitons with a linear decay of sectional curvature. In particular, we classify all 3-dimensional steady Ricci solitons and 4-dimensional κ-noncollpased steady Ricci solitons with nonnegative sectional curvature under the linear curvature decay.
The study shows that the visible range from a point on harmonic manifolds follows an exponential distribution.
problem Understanding the visible range from a point on harmonic manifolds.
method Analyzing Poisson Boolean models on harmonic manifolds, focusing on the geometric mechanism of tube volumes around geodesic segments.
result The visible range from a point on harmonic manifolds follows an exponential distribution.
Last SGD iterate bounds for overparameterized linear regression.
problem Analyzing the last iterate risk bounds of SGD with decaying stepsize for overparameterized linear regression.
method Problem-dependent analysis of last iterate risk bounds of SGD with geometrically decaying stepsize.
result Proved nearly matching upper and lower bounds on the excess risk for last iterate SGD with geometrically decaying stepsize.
Study on scalar curvature decay on non-compact manifolds linked at infinity.
problem Understanding scalar curvature decay on non-compact manifolds with topological linking at infinity.
method Analyzing polynomial decay, developing obstruction theory, using μ--bubble exhaustions, and index theory. result Topological linking at infinity forces polynomial decay of scalar curvature on manifolds of weakly bounded geometry.
In this article we study the dependence degree of the traded volume of the Dow Jones 30 constituent equities by using a nonextensive generalised form of the Kullback-Leibler information measure. Our results show a slow decay of the dependence degree as a function of the lag. This feature is compatible with the existenc…
In this paper, we prove the linear stability to gravitational and electromagnetic perturbations of the Reissner-Nordström family of charged black holes with small charge. Solutions to the linearized Einstein-Maxwell equations around a Reissner-Nordström solution arising from regular initial data remain globally bounded…
WildCat efficiently compresses neural network attention mechanisms.
problem Expensive quadratic runtime of attention mechanisms in neural networks.
method Uses a weighted coreset with a fast subsampling algorithm to approximate attention with near-linear time complexity.
result Approximates exact attention with super-polynomial error decay and near-linear runtime.
Unified framework for analyzing gradient flows of measures with exponential decay of entropy.
problem Analyzing exponential decay of entropy functionals in gradient flows of measures.
method Characterization of global exponential decay behaviors using Hellinger-Kantorovich geometry, shape-mass decomposition, and Polyak-Łojasiewicz-type inequalities.
result Unified theoretical framework for gradient flows with complete analysis of exponential decay behaviors.
We prove that any noncompact κ-noncollapsed steady Ricci soliton with nonnegative curvature operator must be rotationally symmetric if it has a linear curvature decay.
A random walk wn on a separable, geodesic hyperbolic metric space X converges to the boundary ∂X with probability one when the step distribution supports two independent loxodromics. In particular, the random walk makes positive linear progress. Progress is known to be linear with exponential decay when …
Analyzes minima of deep linear networks with weight decay.
problem Understanding the loss landscape of deep neural networks.
method Analytical solutions for global minima with weight decay and stochastic neurons.
result The origin is a special point with qualitatively different minima in networks with more than 1 hidden layer.
SignSGD outperforms SGD in linear regression with optimal scaling laws under PLRF model.
problem Improving linear regression performance with signSGD under power-law random features.
method Analysis of signSGD risk under PLRF model, comparison with SGD, identification of unique effects.
result SignSGD can have a steeper compute-optimal slope than SGD in noisy regimes, especially with WSD schedule.
Develops a new deep learning formulation using Mori-Zwanzig formalism.
problem Improves deep learning by introducing a new concept of memory.
method Uses Mori-Zwanzig formalism to propagate quantities of interest through neural networks.
result Rigorously transforms deep networks into shallow ones using decay property of memory operator.
Study on curvature decay in steady Ricci solitons, proving dichotomy.
problem Curvature decay in steady Ricci solitons.
method Established a dichotomy for curvature decay in specific types of solitons.
result Proved a dichotomy on curvature decay for certain steady Ricci solitons.
This paper reveals periodic behavior in neural network training with BN and weight decay.
problem Understanding the dynamics of neural network training with BN and weight decay.
method Rigorous investigation of empirical and theoretical mechanisms.
result Periodic behavior in training is a generalization of previously opposing perspectives.
Weight decay stabilizes training dynamics by slowing progressive sharpening.
problem Understanding how weight decay affects training stability in deep learning models.
method Analyzing weight decay effects at the Edge of Stability, developing a mathematical framework.
result Weight decay dampens oscillations and stabilizes sharpness in CNNs, causing a phase transition in MLPs.
We prove in this paper the linear stability of the celebrated Schwarzschild family of black holes in general relativity: Solutions to the linearisation of the Einstein vacuum equations around a Schwarzschild metric arising from regular initial data remain globally bounded on the black hole exterior and in fact decay to…
We analyze a simple prefiltered variation of the least squares estimator for the problem of estimation with biased, semi-parametric noise, an error model studied more broadly in causal statistics and active learning. We prove an oracle inequality which demonstrates that this procedure provably mitigates the variance in…
Active data collection improves convergence rates in operator learning.
problem Improving convergence rates in operator learning with linear target and stochastic input.
method Active data collection strategies with mean-zero stochastic process and continuous covariance kernels.
result Achieves arbitrarily fast error convergence rates with eigenvalue decay of covariance kernels.
Framework models supervised learning in non-stationary data.
problem Non-stationary data in supervised learning.
method Statistical physics methods applied to LVQ and neural networks.
result LVQ and ReLU have different sensitivity to concept drift.
Gradient methods work well on overparameterized diagonal linear networks.
problem Understanding why gradient-based methods work well in overparameterized models.
method Study of Deep Diagonal Linear Networks with gradient flow analysis.
result Gradient flow on layer parameters induces a mirror-flow dynamic in the effective parameter space, leading to explicit convergence guarantees.
Analyzes why neural networks generalize beyond training data.
problem Understanding why neural networks generalize beyond training data.
method Examined loss landscapes of neural networks to identify the mismatch between training and test losses.
result Identified the 'LU mechanism' explaining grokking in various tasks.
Optimal learning rates decay to zero in easy tasks and maintain a warmup phase in hard tasks.
problem Optimizing learning rates under functional scaling laws for model training.
method Deriving optimal learning-rate schedules based on exponents s and β. result Sharp phase transition between easy and hard tasks, with different decay behaviors.
In this paper, we study the theory of linearized gravity and prove the linear stability of Schwarzschild black holes as solutions of the vacuum Einstein equations. In particular, we prove that solutions to the linearized vacuum Einstein equations centered at a Schwarzschild metric, with suitably regular initial data, r…
Gradient flow with weight decay shows grokking effect in deep learning.
problem Understanding the grokking effect in deep learning.
method Analyzing gradient flow dynamics with weight decay.
result Weight decay causes slow norm reduction, explaining grokking.
Several important applications, such as streaming PCA and semidefinite programming, involve a large-scale positive-semidefinite (psd) matrix that is presented as a sequence of linear updates. Because of storage limitations, it may only be possible to retain a sketch of the psd matrix. This paper develops a new algorith…
The paper establishes curvature estimates for solitons in higher dimensions.
problem Curvature estimates for steady and expanding solitons in higher dimensions.
method Curvature estimates using gradient Ricci solitons and integral estimates.
result Curvature operator decays at specific rates for different cases of solitons.
Paper tackles causal inference with partially labeled data, introducing robust methods.
problem Challenges in causal inference due to partially labeled datasets and potential bias.
method Decaying missing-at-random framework and BRSS estimator for doubly robust causal inference.
result Established asymptotic normality of BRSS estimator under decaying labeling propensity scores.
We review our recent work on linear stability for scalar perturbations of Kerr spacetimes, that is to say, boundedness and decay properties for solutions of the scalar wave equation \Box_gψ = 0 on Kerr exterior backgrounds. We begin with the very slowly rotating case |a| \ll M, where first boundedness and then decay ha…
Stochastic (sub)gradient methods require step size schedule tuning to perform well in practice. Classical tuning strategies decay the step size polynomially and lead to optimal sublinear rates on (strongly) convex problems. An alternative schedule, popular in nonconvex optimization, is called \emph{geometric step decay…
New model explains volatility after extreme stock market events.
problem Understanding volatility dynamics after extreme stock market events.
method Proposed a new dynamical model using high frequency minute data.
result Volatility after extreme events follows a stretched exponential decay initially and a power law decay later.
Survey on stability of Minkowski spacetime in relativity.
problem Nonlinear stability of Minkowski spacetime in general relativity.
method Decay assumptions, geometric foliations, energy identities, and gauge choices.
result Understanding of decay, dispersion, and geometry-analysis interplay.
This paper contains the second part of a two-part series on the stability and instability of extreme Reissner-Nordstrom spacetimes for linear scalar perturbations. We continue our study of solutions to the linear wave equation on a suitable globally hyperbolic subset of such a spacetime, arising from regular initial da…
We study the problem of stability and instability of extreme Reissner-Nordstrom spacetimes for linear scalar perturbations. Specifically, we consider solutions to the linear wave equation on a suitable globally hyperbolic subset of such a spacetime, arising from regular initial data prescribed on a Cauchy hypersurface …
Researchers found solutions to a complex equation on spheres, overcoming a key difficulty.
problem Finding solutions to a specific equation on spheres with a background metric.
method Constructed a smooth metric invariant under antipodal map, used a noncompact family of solutions, and addressed the loss of ellipticity.
result Provided solutions to the σ2-Yamabe equation for n=27 and beyond, overcoming a main difficulty. Gradient descent converges linearly in finite-width networks with positive NTK and compatible conditions.
problem Local convergence of gradient descent in finite-width networks.
method Positive Neural Tangent Kernel (NTK), local Polyak-Łojasiewicz inequality, fixed-step containment in Locally Quasi-Convex Region (LQCR).
result Linear convergence achieved under specific conditions.
Study proves global existence and decay for complex wave equations.
problem Global existence and decay for quasilinear wave equations with weak-null condition.
method Novel decoupling of higher order energy estimates, focusing on tangential components.
result Established global existence and decay for solutions with small data.
We study the problem of what causes prices to change. We define the mechanical impact of a trading order as the change in future prices in the absence of any future changes in decision making, and its it informational impact as the remainder of the total impact once mechanical impact is removed. We introduce a method o…
AdamNX improves Adam's stability by adjusting its learning rate.
problem Adam's tendency to converge to non-flat minima in large-scale models.
method Proposes a novel exponential decay mechanism for Adam's second-order moment estimate.
result AdamNX outperforms Adam and its variants in stability and performance.
Two models predict similar high-frequency price dynamics but differ in low-frequency impact strength.
problem Understanding the relationship between market prices and fundamental information.
method Comparing a microfounded linear model with a data-driven model at high and low frequencies.
result Both models predict similar high-frequency price dynamics but differ in low-frequency impact strength.
We consider a model for linear transient price impact for multiple assets that takes cross-asset impact into account. Our main goal is to single out properties that need to be imposed on the decay kernel so that the model admits well-behaved optimal trade execution strategies. We first show that the existence of such s…
It is generally accepted that many time series of practical interest exhibit strong dependence, i.e., long memory. For such series, the sample autocorrelations decay slowly and log-log periodogram plots indicate a straight-line relationship. This necessitates a class of models for describing such behavior. A popular cl…
We develop heat kernel and Green's function estimates for manifolds with positive bottom spectrum. The results are then used to establish existence and sharp estimates of the solution to the Poisson equation on such manifolds with Ricci curvature bounded below. As an application, we show that the curvature of a steady …
Study reveals dynamics of neural networks with normalization, weight decay, and SGD.
problem Understanding the equilibrium condition in Spherical Motion Dynamics (SMD).
method Investigates SMD by exploring the cause of equilibrium condition, introducing assumptions, proposing angular update, and verifying theoretical results.
result Proves weight norm and angular update can converge at linear rate under given assumptions.
Wide neural networks with weight decay exhibit neural collapse.
problem Proving neural collapse in wide neural networks trained with weight decay.
method Generic guarantees on neural collapse for wide networks with weight decay, proving low training error and balancedness, and bounded conditioning.
result First proof of neural collapse in end-to-end training of wide neural networks with weight decay.