Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

92183275366 · May 202619922001200920172026
48 results for geometric step decay

Geometric step decay schedules improve stochastic algorithms' convergence on sharp nonconvex problems.

problem Convergence of stochastic algorithms on sharp nonconvex problems.
method Geometric step decay schedule applied to stochastic algorithms.
result Geometric step decay schedules lead to local linear convergence rates for sharp nonconvex problems.

Improved learning rate schedule for least squares regression.

problem Achieving optimal convergence rates for least squares regression.
method Step Decay schedule with geometrically decaying learning rates.
result Final iterate behavior with Step Decay schedules is off the minimax rate by only log factors.

Step decay schedules improve convergence in non-convex optimization.

problem Improving convergence in non-convex optimization problems.
method Analyzing convergence rates of step decay schedules in non-convex, convex, and strongly convex problems.
result Step decay schedules achieve O(lnT/T)\mathcal{O}(\ln T/\sqrt{T}) convergence rates in various optimization scenarios.

Study of 4D Ricci solitons with symmetry, finding precise geometric asymptotics.

problem Classifying 4D gradient steady Ricci solitons and understanding their geometric properties.
method Analysis of 4D gradient steady Ricci solitons with O(3)-symmetry under a weak curvature decay condition.
result Find precise geometric asymptotics similar to 3D compact κ-solutions.

Last SGD iterate bounds for overparameterized linear regression.

problem Analyzing the last iterate risk bounds of SGD with decaying stepsize for overparameterized linear regression.
method Problem-dependent analysis of last iterate risk bounds of SGD with geometrically decaying stepsize.
result Proved nearly matching upper and lower bounds on the excess risk for last iterate SGD with geometrically decaying stepsize.

A random walk wnw_n on a separable, geodesic hyperbolic metric space XX converges to the boundary X\partial X with probability one when the step distribution supports two independent loxodromics. In particular, the random walk makes positive linear progress. Progress is known to be linear with exponential decay when …

2017-10-14abs ↗pdf ↗

This is the second in a series of three papers in which we initiate the study of very rough solutions to the initial value problem for the Einstein vacuum equations expressed relative to wave coordinates. By very rough we mean solutions which cannot be constructed by the classical techniques of energy estimates and Sob…

2001-09-23abs ↗pdf ↗

Paper develops an online learning algorithm for functional data models.

problem Recovering slope functions or predictors in functional data models.
method Online regularized learning algorithm in reproducing kernel Hilbert spaces with polynomially decaying step-size.
result Established fast convergence rates for estimation error without capacity assumption.

Demon improves neural network training with a decaying momentum approach.

problem Improving neural network training efficiency and robustness.
method Proposes a decaying momentum ( extsc{Demon}) rule for neural network optimization.
result Demon achieves the highest number of Top-1 and Top-3 finishes across various settings and architectures.

Paper tackles dynamic pricing in a geometrically decaying environment, achieving better occupancy with lower rates.

problem Minimizing expected loss in a dynamically changing environment with decisions dependent on the data distribution.
method Introduces algorithms for information and loss function settings, using repeated decision deployment to allow mixing of the environment.
result Iteration complexity matches first and zero order stochastic gradient methods up to logarithmic factors.

One knows that the large time heat decay exponent on a nilpotent group is given by half the growing rate of the volume of its large balls. This work deals with the similar problem of trying to interpret geometrically the heat decay on (one) forms. We will show how it is (partially) related to the depth of the relations…

2001-12-06abs ↗pdf ↗

New insights into how to inspect and learn from multi-stage processes and AI reasoning.

problem Understanding how to attribute outcomes to early stages in multi-stage operations and AI reasoning.
method Information-theoretic analysis and mathematical proofs of four key results.
result Uniform checkpoint spacing is minimax-optimal for inspection design under homogeneous signal attenuation.

We present the results of computer experiments suggesting that the probability that a random multiword in a free group is virtually geometric decays to zero exponentially quickly in the length of the multiword. We then prove this fact.

2014-07-29abs ↗pdf ↗

Gradient descent converges linearly in finite-width networks with positive NTK and compatible conditions.

problem Local convergence of gradient descent in finite-width networks.
method Positive Neural Tangent Kernel (NTK), local Polyak-Łojasiewicz inequality, fixed-step containment in Locally Quasi-Convex Region (LQCR).
result Linear convergence achieved under specific conditions.

New theory sharpens Q-learning with LDTZ rate, proving it's best of both worlds.

problem Improving Q-learning's theoretical and practical performance.
method Developed a sharp non-asymptotic error bound and central limit theory for Q-learning with PD2Z-ν schedule.
result Q-learning with LDTZ schedule achieves rapid decay and asymptotic convergence guarantees.

Regularizers change the geometric properties of loss functions in neural networks.

problem Understanding how different regularizers affect the geometric properties of loss functions in neural networks.
method Examined several regularizers, including weight decay, to determine if the regularized loss function becomes Morse.
result For certain regularizers, the regularized loss function becomes Morse, indicating a change in geometric properties.

SignSGD outperforms SGD in linear regression with optimal scaling laws under PLRF model.

problem Improving linear regression performance with signSGD under power-law random features.
method Analysis of signSGD risk under PLRF model, comparison with SGD, identification of unique effects.
result SignSGD can have a steeper compute-optimal slope than SGD in noisy regimes, especially with WSD schedule.

The paper studies Teichmüller TQFT for hyperbolic knots, proving exponential decay of partition functions.

problem Analyzing Teichmüller TQFT for hyperbolic knots with generalized FAMED triangulations.
method Introducing generalized FAMED property, proving exponential decay of partition functions in semi-classical limit.
result Partition functions decay exponentially with the volume of knot complements, and the 1-loop invariant emerges.

The energy in a square membrane ΩΩ subject to constant viscous damping on a subset ωΩω\subset Ω decays exponentially in time as soon as ωω satisfies a geometrical condition known as the "Bardos-Lebeau-Rauch" condition. The rate τ(ω)τ(ω) of this decay satisfies τ(ω)=2min(μ(ω),g(ω))τ(ω)= 2 \min(-μ(ω), g(ω)) (see Lebeau [Math. Phys. Stud. …

2007-06-01abs ↗pdf ↗

Study geometric step options with jumps, deriving pricing equations and characterizations.

problem Pricing geometric step options in markets with jumps.
method Symmetry and parity relations, partial integro-differential equations, ordinary integro-differential equations.
result Derive semi-analytical pricing results for geometric step options.

Model proposes neural network for continuous time dynamics with inductive biases.

problem Training neural networks for small datasets with nonlinear dynamics.
method Inductive biases on decay rates and frequencies using Koopman operator theory.
result Higher forecasting performance with single short training sequence.

Improved analysis for fair federated learning reduces dependence on noise floor.

problem Asymptotic stationarity in group fair federated learning with reduced noise floor dependence.
method DS FedProxGrad framework with inexact local proximal solutions and fairness regularization.
result Algorithm converges asymptotically to stationarity without dependence on a noise floor.

We introduce a novel algorithm that computes the kk-sparse principal component of a positive semidefinite matrix AA. Our algorithm is combinatorial and operates by examining a discrete set of special vectors lying in a low-dimensional eigen-subspace of AA. We obtain provable approximation guarantees that depend on t…

2013-03-03abs ↗pdf ↗

New method uses entropy dissipation to prove isoperimetric inequalities.

problem Proving isoperimetric inequalities in geometric settings.
method Information-theoretic approach based on entropy dissipation under heat flow.
result New proof of Euclidean isoperimetric inequality with sharp constant.

Geometric focusing affects dispersive estimates for Schrödinger and wave equations.

problem Long-time decay rate in dispersive estimates for Schrödinger and wave equations on non-trapping asymptotically conic manifolds and exact metric cones.
method Classifying the long-time decay rate in dispersive estimates for the Schrödinger and wave equations on non-trapping asymptotically conic manifolds and exact metric cones in terms of the intensity of geometric focusing.
result Each multiplicity of conjugate points within distance π on Y = ∂X0 leads to a |t|1/2-loss in the long-time decay order and a half-order shift in the regularity index in the dispersive estimate for the Schrödinger equation.

Stagewise training strategy is widely used for learning neural networks, which runs a stochastic algorithm (e.g., SGD) starting with a relatively large step size (aka learning rate) and geometrically decreasing the step size after a number of iterations. It has been observed that the stagewise SGD has much faster conve…

2018-12-10abs ↗pdf ↗

The paper analyzes Teukolsky equations on Kerr backgrounds, proving boundedness and decay of solutions.

problem Analyzing boundedness and decay of solutions to Teukolsky equations on Kerr backgrounds.
method Frequency space analysis of transformed Teukolsky equations on Kerr backgrounds.
result Fixed frequency solutions remain bounded and decay in time for subextremal Kerr backgrounds.

This paper is motivated by the non-linear stability problem for the expanding region of Kerr de Sitter cosmologies in the context of Einstein's equations with positive cosmological constant. We show that under dynamically realistic assumptions the conformal Weyl curvature of the spacetime decays towards future null inf…

2016-10-13abs ↗pdf ↗

Temporal Difference Learning analysis under non-i.i.d. data and nonlinear approximation.

problem Finite-sample behavior of TD(0) under non-i.i.d. data and nonlinear approximation.
method High-probability, finite-sample analysis of vanilla TD(0) on polynomially mixing Markov data, assuming Holder continuity and bounded generalized gradients.
result Bounds on the convergence rate of TD(0) with high probability, matching known i.i.d. rates and holding even with nonstationary initialization.

In this paper, we study the online learning algorithm without explicit regularization terms. This algorithm is essentially a stochastic gradient descent scheme in a reproducing kernel Hilbert space (RKHS). The polynomially decaying step size in each iteration can play a role of regularization to ensure the generalizati…

2017-10-10abs ↗pdf ↗

AdamNX improves Adam's stability by adjusting its learning rate.

problem Adam's tendency to converge to non-flat minima in large-scale models.
method Proposes a novel exponential decay mechanism for Adam's second-order moment estimate.
result AdamNX outperforms Adam and its variants in stability and performance.

WSD schedule improves model training efficiency by adapting learning rates dynamically.

problem Fixed compute budgets limit training efficiency of language models.
method Introduces a WSD schedule that uses a constant learning rate followed by a rapid decay phase.
result WSD schedule generates a non-traditional loss curve with stable and decay phases.

Gradient descent outperforms ridge regression under certain covariance matrix decay conditions.

problem Comparing the performance of gradient descent and ridge regression in linear models.
method Investigated gradient descent and ridge regression for linear regression with random isotropic ground truth.
result Gradient descent outperforms ridge regression under specific covariance matrix decay conditions.