Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

95191286381 · Jun 202019922001200920172026
48 results for gradient decay

The paper studies steady solitons with curvature decay and proves their smoothness.

problem Analyzing the properties of steady solitons with curvature decay.
method Bootstrap regularity in harmonic coordinates using the soliton equation.
result Steady gradient Ricci solitons are asymptotically cylindrical under certain curvature decay conditions.

In this paper we introduce a novel method of gradient normalization and decay with respect to depth. Our method leverages the simple concept of normalizing all gradients in a deep neural network, and then decaying said gradients with respect to their depth in the network. Our proposed normalization and decay techniques…

2017-12-10abs ↗pdf ↗

In this paper we study potential function of gradient steady Ricci solitons. We prove that infimum of potential function decays linearly; in particular, potential function of rectifiable gradient steady Ricci solitons decays linearly. As a consequence, we show that a gradient steady Ricci soliton with bounded potential…

2011-02-15abs ↗pdf ↗

Unified framework for analyzing gradient flows of measures with exponential decay of entropy.

problem Analyzing exponential decay of entropy functionals in gradient flows of measures.
method Characterization of global exponential decay behaviors using Hellinger-Kantorovich geometry, shape-mass decomposition, and Polyak-Łojasiewicz-type inequalities.
result Unified theoretical framework for gradient flows with complete analysis of exponential decay behaviors.

This paper analyzes convergence of large-scale Transformers with weight decay.

problem Understanding optimization guarantees in large-scale Transformer training.
method Construct mean-field limit, show gradient flow convergence to PDE, demonstrate global minimum consistency.
result Gradient flow reaches global minimum in large-scale Transformers with small weight decay.

Riemannian stochastic gradient descent converges faster with increasing batch size.

problem Improving convergence rate of Riemannian stochastic gradient descent.
method Theoretical analysis and numerical investigation of increasing batch size effects.
result Riemannian stochastic gradient descent converges faster with increasing batch size.

Constructs expanding gradient Ricci solitons with unique properties.

problem Creating expanding gradient Ricci solitons with specific characteristics.
method Combining previous work with localized maximum principle.
result Constructs various examples of expanding gradient Ricci solitons with positive curvature and exotic curvature decay.

Adam's generalization performance is improved by batch size and weight decay in neural networks.

problem Understanding how batch size and weight decay affect Adam's generalization in neural networks.
method Theoretical analysis of two-layer over-parameterized CNNs on image data.
result Adam's mini-batch variants can achieve near-zero test error, unlike full-batch Adam.

This paper investigates the effectiveness of decoupled weight decay at the start of training.

problem The traditional approach to weight decay is not effective throughout training.
method The authors investigate decoupled weight decay, applying it only at the start of training.
result Applying weight decay only at the start of training stabilizes network weights and improves performance.

Regularization in the optimization of deep neural networks is often critical to avoid undesirable over-fitting leading to better generalization of model. One of the most popular regularization algorithms is to impose L-2 penalty on the model parameters resulting in the decay of parameters, called weight-decay, and the …

2019-07-21abs ↗pdf ↗

Step decay schedules improve convergence in non-convex optimization.

problem Improving convergence in non-convex optimization problems.
method Analyzing convergence rates of step decay schedules in non-convex, convex, and strongly convex problems.
result Step decay schedules achieve O(lnT/T)\mathcal{O}(\ln T/\sqrt{T}) convergence rates in various optimization scenarios.

Gradient descent outperforms ridge regression under certain covariance matrix decay conditions.

problem Comparing the performance of gradient descent and ridge regression in linear models.
method Investigated gradient descent and ridge regression for linear regression with random isotropic ground truth.
result Gradient descent outperforms ridge regression under specific covariance matrix decay conditions.

We study stability of non-compact gradient Kaehler-Ricci flow solitons with positive holomorphic bisectional curvature. Our main result is that any compactly supported perturbation and appropriately decaying perturbations of the Kaehler potential of the soliton will converge to the original soliton under Kaehler-Ricci …

2003-07-22abs ↗pdf ↗

New findings on gradient expanding Ricci solitons with finite scalar curvature ratio.

problem Understanding the behavior of gradient expanding Ricci solitons with finite scalar curvature ratio.
method Analyzing complete gradient expanding Ricci solitons with nonnegative Ricci curvature.
result Riemann curvature tensor must have at least sub-quadratic decay for finite asymptotic scalar curvature ratio.

In this paper, we study the gradient Ricci soliton equation on a complete Riemannian manifold. We show that under a natural decay condition on the Ricci curvature, the Ricci soliton is Ricci-flat and ALE.

2004-11-19abs ↗pdf ↗

SAD-DPSGD improves model performance on imbalanced medical datasets like HAM10000.

problem Data leakage and imbalanced distribution in medical image classification datasets.
method SAD-DPSGD uses a linear decaying mechanism for noise and clipping thresholds to enhance performance.
result SAD-DPSGD outperforms Auto-DPSGD on HAM10000, improving accuracy by 2.15%.

We show that the only complete shrinking gradient Ricci solitons with vanishing Weyl tensor are quotients of the standard ones. This gives a new proof of the Hamilton-Ivey-Perel'man classification of 3-dimensional shrinking gradient solitons. We also prove a classification for expanding gradient Ricci solitons with con…

2007-12-08abs ↗pdf ↗

Last SGD iterate bounds for overparameterized linear regression.

problem Analyzing the last iterate risk bounds of SGD with decaying stepsize for overparameterized linear regression.
method Problem-dependent analysis of last iterate risk bounds of SGD with geometrically decaying stepsize.
result Proved nearly matching upper and lower bounds on the excess risk for last iterate SGD with geometrically decaying stepsize.

Weight decay stabilizes training dynamics by slowing progressive sharpening.

problem Understanding how weight decay affects training stability in deep learning models.
method Analyzing weight decay effects at the Edge of Stability, developing a mathematical framework.
result Weight decay dampens oscillations and stabilizes sharpness in CNNs, causing a phase transition in MLPs.

The paper establishes curvature estimates for solitons in higher dimensions.

problem Curvature estimates for steady and expanding solitons in higher dimensions.
method Curvature estimates using gradient Ricci solitons and integral estimates.
result Curvature operator decays at specific rates for different cases of solitons.

SGD and weight decay encourage neural networks to learn low-rank weight matrices.

problem The bias of SGD towards low-rank weight matrices in neural networks.
method The study investigates the effect of SGD and weight decay on the rank of weight matrices in neural networks, both theoretically and empirically.
result Training with SGD and weight decay induces a bias towards rank minimization in weight matrices, which becomes more pronounced with smaller batch sizes and stronger weight decay.

We develop heat kernel and Green's function estimates for manifolds with positive bottom spectrum. The results are then used to establish existence and sharp estimates of the solution to the Poisson equation on such manifolds with Ricci curvature bounded below. As an application, we show that the curvature of a steady …

2017-01-11abs ↗pdf ↗

Uniform diffusion approximation for SGD in non-convex settings.

problem Finite-time diffusion approximation for SGD.
method Establishing uniform-in-time diffusion approximation with strong convexity and mild conditions.
result Uniform-in-time diffusion approximation of SGD without convexity of each loss function.

Paper tackles dynamic pricing in a geometrically decaying environment, achieving better occupancy with lower rates.

problem Minimizing expected loss in a dynamically changing environment with decisions dependent on the data distribution.
method Introduces algorithms for information and loss function settings, using repeated decision deployment to allow mixing of the environment.
result Iteration complexity matches first and zero order stochastic gradient methods up to logarithmic factors.

In this paper, we study the minimax optimization problem in the smooth and strongly convex-strongly concave setting when we have access to noisy estimates of gradients. In particular, we first analyze the stochastic Gradient Descent Ascent (GDA) method with constant stepsize, and show that it converges to a neighborhoo…

2020-02-13abs ↗pdf ↗

We study minimal graphic functions on complete Riemannian manifolds $\Si$ with non-negative Ricci curvature, Euclidean volume growth and quadratic curvature decay. We derive global bounds for the gradients for minimal graphic functions of linear growth only on one side. Then we can obtain a Liouville type theorem with …

2013-10-08abs ↗pdf ↗

The paper analyzes reg-SGD for convex problems, proving convergence and quantifying the rate of convergence.

problem Minimizing convex, L-smooth functions in a Hilbert space.
method Regularized stochastic gradient descent with decaying regularization.
result Strong convergence to the minimum-norm solution without boundedness assumptions.

New method trains neural networks with threshold activation functions efficiently.

problem Training neural networks with threshold activation functions is challenging due to zero gradients.
method We study weight decay regularized training problems of deep neural networks with threshold activations, showing they can be formulated as convex optimization problems.
result Regularized deep threshold network training problems can be formulated as standard convex optimization problems, paralleling the LASSO method.