Improved analysis shows momentum in SGD reduces batch size needs for non-convex objectives.
problem Reducing batch size requirements for SGD in non-convex optimization.
method Normalized SGD with momentum, adaptive method for small gradient variance.
result Normalized SGD with momentum achieves ε-critical points in O(1/ε3.5) iterations. The paper investigates how target normalization and momentum affect dying ReLUs in neural networks.
problem Understanding and mitigating the dying ReLU problem in neural networks.
method Empirical analysis and theoretical modeling of a discrete-time linear autonomous system.
result Target variance plays a crucial role in the dying ReLU phenomenon, and momentum exacerbates this issue.
Inverts rank m symmetric tensor fields using line integrals.
problem Recovering symmetric tensor fields from line integrals.
method Computes normal operator and presents inversion formula.
result Recovering rank m tensor fields from data (Nm0f,…,Nmmf). New methods solve optimization problems with heavy-tailed noise, improving upon existing complexity bounds.
problem Optimization problems with heavy-tailed noise and weakly average smoothness.
method Normalized stochastic first-order methods with Polyak, multi-extrapolated, and recursive momentum.
result First-order oracle complexity results for finding approximate stochastic stationary points under heavy-tailed noise.
SGDM accelerates faster than SGD with large batch sizes and permits broader learning rates.
problem Understanding the role of momentum in SGDM and its convergence rates.
method Analysis of SGDM convergence rates under strongly convex settings, including finite-sample rates and asymptotic normality of the averaged estimator.
result SGDM converges faster than SGD with large batch sizes and permits broader learning rates.
Theory of symplectic reduction in infinite dimensions developed.
problem Challenges in symplectic reduction in infinite dimensions.
method Normal form of momentum map for infinite-dimensional equivariant maps.
result Theory of singular symplectic reduction in infinite dimensions.
We give a generalization of toric symplectic geometry to Poisson manifolds which are symplectic away from a collection of hypersurfaces forming a normal crossing configuration. We introduce the tropical momentum map, which takes values in a generalization of affine space called a log affine manifold. Using this momentu…
In this paper we prove rigidity theorems for Poisson Lie group actions on Poisson manifolds. In particular, we prove that close infinitesimal momentum maps associated to Poisson Lie group actions are equivalent using a normal form theorem for SCI spaces. When the Poisson structure of the acted manifold is integrable, t…
Intriguing empirical evidence exists that deep learning can work well with exoticschedules for varying the learning rate. This paper suggests that the phenomenon may be due to Batch Normalization or BN, which is ubiquitous and provides benefits in optimization and generalization across all standard architectures. The f…
SNGM improves large-batch training accuracy.
problem Improving generalization in large-batch training.
method Stochastic Normalized Gradient Descent with Momentum.
result SNGM achieves better test accuracy than MSGD and other large-batch methods.
AuON is a linear-time optimizer that improves upon Muon's performance without approximate orthogonal matrices.
problem High memory and computational costs of orthogonal momentum updates.
method AuON uses normalized nonlinear scaling and a 'emergency brake' to handle exploding attention logits.
result AuON achieves strong performance without approximate orthogonal matrices, preserving structural alignment and reconditioning.
Normality equations describe Newtonian dynamical systems admitting normal shift of hypersurfaces. These equations were first derived in Euclidean geometry. Then very soon they were rederived in Riemannian and in Finslerian geometry. Recently I have found that normality equations can be derived in geometry given by clas…
AdamS uses momentum as a denominator to optimize LLMs efficiently.
problem Optimizing large language models (LLMs) with efficient and effective methods.
method AdamS introduces a novel denominator based on the root of the weighted sum of squares of momentum and current gradient.
result AdamS achieves superior optimization performance with minimal memory and compute requirements.
This work addresses the instability in asynchronous data parallel optimization. It does so by introducing a novel distributed optimizer which is able to efficiently optimize a centralized model under communication constraints. The optimizer achieves this by pushing a normalized sequence of first-order gradients to a pa…
AdamP optimizes momentum-based optimizers for scale-invariant weights, improving model performance.
problem Premature decay of effective step sizes in momentum-based optimizers for scale-invariant weights.
method Proposes SGDP and AdamP to eliminate the radial component at each optimizer step, preserving convergence properties.
result Uniform gains across multiple benchmarks, improving model performance.
A new method uses trivialized momentum to generate data on Lie groups.
problem Generating data on Lie groups with high fidelity and efficiency.
method Introducing an auxiliary momentum variable that stays in a fixed vector space, and using a manifold preserving integrator.
result Achieves state-of-the-art performance on protein and RNA torsion angle generation and high-dimensional Lie groups.
Simplified optimization for structured matrices in deep learning.
problem Computational challenges in Riemannian submanifold optimization for structured symmetric positive-definite matrices.
method Proposed a generalized Riemannian normal coordinates that dynamically orthonormalizes the metric and converts the problem into an unconstrained Euclidean space problem.
result Simplified existing approaches for structured covariances and developed matrix-inverse-free 2nd-order optimizers for deep learning with low precision.
A fast method for training linear classifiers maximizes margins.
problem Training linear classifiers with maximum margins.
method Momentum-based gradient method derived from convex dual with Nesterov acceleration.
result Exponentially faster convergence rate compared to standard methods.
MuonEq improves training of matrix-valued parameters by rebalancing momentum before orthogonalization.
problem Training matrix-valued parameters with orthogonalized-update optimizers like Muon.
method MuonEq introduces three lightweight pre-orthogonalization equilibration schemes: two-sided row/column normalization (RC), row normalization (R), and column normalization (C).
result Row/column normalization acts as a zeroth-order surrogate for whitening and improves the geometry seen by orthogonalization.
A local normal form theorem for smooth equivariant maps between Fréchet manifolds is established. Moreover, an elliptic version of this theorem is obtained. The proof these normal form results is inspired by the Lyapunov-Schmidt reduction for dynamical systems and by the Kuranishi method for moduli spaces, and uses a s…
This article concerns cotangent-lifted Lie group actions; our goal is to find local and ``semi-global'' normal forms for these and associated structures. Our main result is a constructive cotangent bundle slice theorem that extends the Hamiltonian slice theorem of Marle, Guillemin and Sternberg. The result applies to a…
FedNNNN improves FL by adjusting model update vector norms.
problem Slow convergence and low prediction accuracy in FL.
method Norm-Normalized Neural Network Aggregation (FedNNNN) with momentum control.
result Up to 5.4% accuracy improvement on multiple datasets.
The paper shows how overreactions in stock prices can be predicted and used for trading.
problem Predicting and monetizing overreactions in stock prices as momentum signals.
method High-frequency data from Twitter, machine learning models (XGBoost, Random Forests, Deep Neural Networks, Bidirectional LSTMs), and SHAP for explainability.
result Machine learning models significantly outperform traditional overreaction rules at ultra short horizons.
New formula extracts full local information from ray transform data.
problem Determining symmetric tensor fields from ray transform data.
method Deriving explicit formula for Saint Venant operator.
result Explicit formula for extracting full local information.
Arguably, the two most popular accelerated or momentum-based optimization methods in machine learning are Nesterov's accelerated gradient and Polyaks's heavy ball, both corresponding to different discretizations of a particular second order differential equation with friction. Such connections with continuous-time dyna…
The study finds that factor momentum is significant only at short lags compared to stock momentum.
problem Investigating the relationship between factor momentum and stock momentum.
method Replicated earlier findings and conducted a spanning test controlling for stock momentum and factor exposure.
result Factor momentum is significant only at short lags after controlling for stock momentum and factor exposure.
Symplectic reduction by abelian subgroups coincides under specific conditions.
problem Conditions for symplectic reduction by abelian subgroups to match.
method Reduction-by-stages framework, nilpotent Lie groups, generic momentum values.
result Symplectomorphic reduced spaces under mild conditions.
Muon dynamics study uses spectral Wasserstein flow for optimization stability.
problem Optimizing deep learning models with gradient normalization.
method Introduces Spectral Wasserstein distances for matrix flows, proving equivalence with Benamou--Brenier formulation.
result Gradient-flow interpretation of mean-field normalized training dynamics.
We test the price momentum effect in the Korean stock markets under the momentum universe shrinkage to subuniverses of the KOSPI 200. Performance of the momentum strategy is not homogeneous with respect to change of the momentum universe. It is found that some submarkets generate the higher momentum returns than other …
Introduces homotopy momentum sections on multisymplectic manifolds.
problem No specific problem stated; focuses on introducing a new concept.
method Introduces a new concept of homotopy momentum sections on multisymplectic manifolds.
result Shows that a gauged nonlinear sigma model with Wess-Zumino term has homotopy momentum section structure.
Customer momentum is a positive relationship between a firm's returns and past returns of its customers.
problem Understanding the relationship between a firm's returns and its customers' past returns.
method Examined customer momentum using a long-short equally-weighted decile portfolio and Fama-French factor models.
result Customer momentum generates significant monthly returns and is statistically significant.
This paper re-evaluates hyperparameters for fine-tuning pre-trained models.
problem Current hyperparameter settings for fine-tuning are often ad-hoc and fixed.
method Empirical evaluation of learning rate, batch size, and momentum for fine-tuning.
result Optimal hyperparameters are not only dataset-dependent but also sensitive to domain similarity.
New insights into SGD and SGD-M in high dimensions.
problem Understanding and comparing SGD and SGD-M in high-dimensional settings.
method Developed high-dimensional scaling limits for SGD-M and online SGD, examining their dynamics and performance.
result SGD-M amplifies high-dimensional effects, potentially degrading performance compared to online SGD.
Standard optimizers perform as well as LARS and LAMB at large batch sizes.
problem Comparing optimizers for neural network training at large batch sizes.
method Used standard optimizers like Nesterov momentum and Adam to match or exceed LARS and LAMB results.
result Standard optimizers can match or exceed LARS and LAMB at large batch sizes.
This paper examines momentum spillover across multiple asset classes using only pricing data.
problem Challenges in studying momentum spillover across diverse asset classes due to lack of common characteristics.
method Utilised a linear and interpretable graph learning model to reveal momentum spillover network.
result Network momentum strategy yields a Sharpe ratio of 1.5 and an annual return of 22%.
The paper analyzes how hyperparameters affect SGD with momentum's convergence rate.
problem The role of hyperparameters in SGD with momentum's convergence rate.
method Theoretical analysis using a hyperparameters-dependent stochastic differential equation (hp-dependent SDE).
result The optimal linear rate of convergence depends on both the learning rate and the momentum coefficient.
This paper presents generalized momentum mappings for covariant Hamiltonian field theories. The new momentum mappings arise from a generalization of symplectic geometry to LVY, the bundle of vertically adapted linear frames over the bundle of field configurations Y. Specifically, the generalized field momentum obs…
We give a detailed discussion about existence and uniqueness of Lu's momentum map. More precisely, we introduce the infinitesimal momentum map, and we study its properties. This allows us to describe the theory of reconstruction of the momentum map from the infinitesimal one. We provide the conditions for the uniquenes…
Momentum ResNets improve ResNets' memory efficiency.
problem Memory inefficiency in deep residual neural networks (ResNets).
method Adding a momentum term to the forward rule of ResNets to make them invertible.
result Momentum ResNets can learn any linear mapping up to a multiplicative factor, improving memory efficiency.
We introduce various quantitative and mathematical definitions for price momentum of financial instruments. The price momentum is quantified with velocity and mass concepts originated from the momentum in physics. By using the physical momentum of price as a selection criterion, the weekly contrarian strategies are imp…
In previous work with M.C. Fernandes, we found a Lie algebroid symmetry for the Einstein evolution equations of general relativity. The present work was motivated by the effort to explain the coisotropic structure of the constraint subset for the initial value problem by extending the notion of hamiltonian structure fr…
New algorithm Momentum-QNG improves optimization of quantum circuits.
problem Optimizing variational quantum circuits to avoid local minima.
method Applied Langevin dynamics to QNG, introducing momentum term.
result Momentum-QNG outperforms basic QNG and other optimizers.
We demonstrate the possibility of what we call sparse learning: accelerated training of deep neural networks that maintain sparse weights throughout training while achieving dense performance levels. We accomplish this by developing sparse momentum, an algorithm which uses exponentially smoothed gradients (momentum) to…
Momentum speeds up evolutionary processes in machine learning.
problem Accelerating convergence in evolutionary dynamics.
method Combining momentum from machine learning with evolutionary dynamics using information divergences as Lyapunov functions.
result Momentum accelerates convergence of evolutionary dynamics, including the replicator equation and Euclidean gradient descent.
We propose NovoGrad, an adaptive stochastic gradient descent method with layer-wise gradient normalization and decoupled weight decay. In our experiments on neural networks for image classification, speech recognition, machine translation, and language modeling, it performs on par or better than well tuned SGD with mom…
Contact manifolds' momentum polytopes are convex.
problem Understanding the structure of contact manifolds.
method Using isomorphism to toric varieties.
result Momentum polytopes of contact manifolds are convex.
Unified model learns from both time-series and cross-sectional momentum features.
problem Separate time-series and cross-sectional momentum strategies do not consider concurrent relationships.
method Spatio-Temporal Momentum strategies using neural networks to combine both types of momentum.
result Simple neural network with single fully connected layer generates trading signals for all assets.
SMG combines shuffling and momentum for non-convex optimization.
problem Non-convex finite-sum optimization problems.
method Shuffling Gradient-based method with momentum.
result Established state-of-the-art convergence rates for SMG.