Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

3876114152 · May 202619922001200920172026
48 results for Adams theorem

This paper establishes inequalities on quaternionic hyperbolic spaces and the Cayley hyperbolic plane.

problem Establishing higher order Poincaré-Sobolev and Hardy-Sobolev-Maz'ya inequalities on quaternionic hyperbolic spaces and the Cayley hyperbolic plane.
method Developing factorization theorems and introducing Geller's operators, combining with Helgason-Fourier analysis and kernel estimates.
result Established higher order Poincaré-Sobolev and Hardy-Sobolev-Maz'ya inequalities on quaternionic hyperbolic spaces and the Cayley hyperbolic plane.

Sharp inequalities on Siegel domains and complex hyperbolic spaces established.

problem Establishing inequalities on complex hyperbolic spaces and Siegel domains.
method Helgason-Fourier analysis, Kunze-Stein phenomenon, factorization theorem.
result Sharp Hardy-Adams and Adams type inequalities on Sobolev spaces of any positive fractional order on complex hyperbolic spaces.

Sharp inequalities for radial functions on hyperbolic spaces without boundary conditions.

problem Establishing inequalities for radial functions on hyperbolic spaces without zero boundary conditions.
method Novel approach considering both bounded and unbounded domains, focusing on weighted Sobolev and Adams-Trudinger-Moser embeddings.
result Theorems 1.2, 1.3, and 1.4 for weighted Sobolev embedding theorems, and Theorems 1.5 and 1.6 for Adams-Trudinger-Moser type embedding theorems.

Polyak-Ruppert CLT for SA-Adam with momentum and non-convergent adaptive preconditioning

problem Adaptive optimizers combining momentum and non-convergent preconditioning
method Proving positive drift stability and a non-autonomous Polyak-Ruppert CLT for SA-Adam
result The iterate-marginal covariance is exactly the plain stochastic gradient descent (SGD) sandwich

Proves a conjecture for a specific group using spectral sequences and homology.

problem Proves the Gromov-Lawson-Rosenberg Conjecture for the group Z/4xZ/4.
method Used the Adams spectral sequence and detection theorems to compute connective real k-homology.
result Determines differentials of the Adams spectral sequence and studies the cap structure of relevant sub-hopf algebras.

A curvature model (V,A) is a real vector space V which is equipped with a "curvature operator" A(x,y)z that A has the same symmetries as an affine curvature operator; A(x,y)z=-A(y,x)z and A(x,y)z+A(y,z)x+A(z,x)y=0. Such a model is called projective affine Osserman if the spectrum of the Jacobi operator J(y):x->A(x,y)y,…

2014-03-07abs ↗pdf ↗

We describe pairs (p,n) such that n-dimensional affine space is fibered by pairwise skew p-dimensional affine subspaces. The problem is closely related with the theorem of Adams on vector fields on spheres and the Hurwitz-Radon theory of composition of quadratic forms.

2012-03-16abs ↗pdf ↗

Adam's generalization performance is improved by batch size and weight decay in neural networks.

problem Understanding how batch size and weight decay affect Adam's generalization in neural networks.
method Theoretical analysis of two-layer over-parameterized CNNs on image data.
result Adam's mini-batch variants can achieve near-zero test error, unlike full-batch Adam.

AdamS uses momentum as a denominator to optimize LLMs efficiently.

problem Optimizing large language models (LLMs) with efficient and effective methods.
method AdamS introduces a novel denominator based on the root of the weighted sum of squares of momentum and current gradient.
result AdamS achieves superior optimization performance with minimal memory and compute requirements.

Simpler, parameter-free AdaGrad and Adam variants with convergence guarantees.

problem Inefficiencies in ad-hoc learning rate tuning for optimization algorithms.
method Developed AdaGrad++ and Adam++ without predefined learning rates and proved their convergence.
result AdaGrad++ and Adam++ achieve comparable convergence rates to AdaGrad and Adam respectively.

A new memory-efficient Adam variant reduces second moments when feasible.

problem Memory constraints in training machine learning models.
method Signal-to-Noise Ratio (SNR) analysis to identify dimensions where second moments can be replaced by means.
result Memory-efficient Adam variant (SlimAdam) matches performance and stability of Adam while saving up to 98% of second moments.

AdaX improves Adam by exponentially accumulating past gradients, leading to better performance in machine learning tasks.

problem Adam's fast convergence can lead to local minimums in non-convex problems.
method AdaX exponentially accumulates past gradients to adaptively tune the learning rate.
result AdaX outperforms Adam in various machine learning tasks, including computer vision and natural language processing.

Adaptive optimization algorithms, such as Adam and RMSprop, have shown better optimization performance than stochastic gradient descent (SGD) in some scenarios. However, recent studies show that they often lead to worse generalization performance than SGD, especially for training deep neural networks (DNNs). In this wo…

2017-09-13abs ↗pdf ↗

Improved Adam for time series forecasting with distributional drift.

problem Non-stationary data challenges Adam's effectiveness.
method Proposed TS_Adam, removing Adam's second-order bias correction.
result TS_Adam achieves 12.8% reduction in MSE and 5.7% in MAE on ETT datasets.

DualAdam improves generalization of Adam by integrating its update mechanisms.

problem Adam's tendency to converge to sharp minima leading to suboptimal generalization.
method DualAdam combines Adam and inverse Adam's update mechanisms to enhance generalization.
result DualAdam outperforms Adam and state-of-the-art variants in generalization performance.

Novel Adam-family method with decoupled weight decay for training neural networks.

problem Training nonsmooth neural networks with weight decay.
method Proposes a novel Adam-family method with decoupled weight decay, establishing convergence properties and demonstrating superior performance.
result Asymptotically approximates SGD and enhances generalization performance.

Paper studies Adam's convergence under relaxed assumptions, proving a rate of O(poly(log T)/sqrt(T)).

problem Understanding Adam's convergence in non-convex, stochastic optimization with unbounded gradients and noise.
method Introduced a comprehensive noise model and used it to prove Adam's convergence rate.
result Adam finds a stationary point with a rate of O(poly(log T)/sqrt(T)) in high probability.

Adam optimization algorithm can have non-zero average regret under certain conditions.

problem Non-zero average regret in Adam optimization algorithm.
method Used a three-periodic sequence of linear functions on [-1,1] with slopes c, -1, -1, and analyzed Adam variants.
result Adam optimization algorithm can have non-zero average regret under certain conditions.

This work analyzes Adam's preconditioning effect on quadratic functions and quantifies its impact on condition number.

problem Understanding and quantifying the preconditioning effect of Adam to alleviate ill-conditioning in gradient descent.
method Detailed analysis of Adam's preconditioning effect for quadratic functions, including empirical evidence.
result Adam can mitigate the condition number but at a dimension-dependent cost, with specific bounds for different types of Hessians.

GRU models with Adam optimizer outperform other combinations in stock market forecasting.

problem Comparing optimization techniques for time series forecasting in LSTM and GRU networks.
method Examined Adam and Nesterov Accelerated Gradient (NAG) on LSTM and GRU models for stock market forecasting.
result GRU models with Adam optimizer produced the lowest RMSE and outperformed other combinations.

Adam converges with high probability under unconstrained non-convex smooth stochastic optimizations.

problem Theoretical limitations of Adam's convergence under unconstrained non-convex smooth stochastic optimizations.
method Deep analysis of Adam's convergence rate under affine variance noise, without bounded gradient assumptions.
result Adam converges to the stationary point with a high probability rate of $\mathcal{O}\left({ m poly}(\log T)/\sqrt{T} ight)$.

Adam can lead to worse test errors than GD in deep learning, especially with over-parameterized networks.

problem Adam's inferior generalization performance compared to GD in deep learning optimization.
method Theoretical analysis of Adam and gradient descent in over-parameterized two-layer convolutional neural networks.
result Adam and GD can converge to different global solutions with different generalization errors in nonconvex optimization landscapes.

The adaptive optimizer for training neural networks has continually evolved to overcome the limitations of the previously proposed adaptive methods. Recent studies have found the rare counterexamples that Adam cannot converge to the optimal point. Those counterexamples reveal the distortion of Adam due to a small secon…

2019-11-01abs ↗pdf ↗

Adam's bias shifts from full-batch to max-margin of different norms for separable data.

problem Understanding Adam's implicit bias in the incremental batch setting.
method Analyzing incremental Adam on linearly separable data, constructing datasets, and using a proxy algorithm.
result Incremental Adam can converge to different max-margin classifiers depending on the dataset and batching scheme.

This research explains why SGD generalizes better than ADAM in deep learning.

problem Understanding the generalization gap between SGD and ADAM in deep learning.
method Analyzing local convergence behaviors through Levy-driven stochastic differential equations (SDEs).
result SGD is more locally unstable and better escapes from sharp minima to flatter ones, leading to better generalization.

We discuss an universal bordism invariant obtained from the Atiyah-Patodi-Singer eta-invariant from the analytic and homotopy theoretic point of view. Classical invariants like the Adams e-invariant, ρρ-invariants and StringString-bordism invariants are derived as special cases. The main results are a secondary index theo…

2011-03-22abs ↗pdf ↗

Adam and RMSProp are two of the most influential adaptive stochastic algorithms for training deep neural networks, which have been pointed out to be divergent even in the convex setting via a few simple counterexamples. Many attempts, such as decreasing an adaptive learning rate, adopting a big batch size, incorporatin…

2018-11-23abs ↗pdf ↗

The paper analyzes Adam and SGD in nonstationary optimization, revealing tradeoffs between noise and drift.

problem Analyzing Adam and SGD in nonstationary optimization problems.
method Theoretical analysis of Adam and SGD under non-stationary stochastic objectives, separating two regimes.
result Characterizes the tradeoff between noise and drift in Adam and SGD, revealing when adaptive step-sizing is beneficial or harmful.