The paper analyzes how hyperparameters affect SGD with momentum's convergence rate.
problem The role of hyperparameters in SGD with momentum's convergence rate.
method Theoretical analysis using a hyperparameters-dependent stochastic differential equation (hp-dependent SDE).
result The optimal linear rate of convergence depends on both the learning rate and the momentum coefficient.
A new family of momentum coefficients improves the convergence rate of accelerated algorithms.
problem Improving the convergence rate of accelerated gradient methods for strongly convex functions.
method Introducing a family of controllable momentum coefficients for forward-backward accelerated methods.
result Established a controllable $O\left(1/k^{2α}
ight)$ convergence rate for the NAG-α method. Formulae for mass and angular momentum transformations under BMS transformations derived from curvature and metric.
problem Deriving transformation formulae for mass and angular momentum under BMS transformations.
method Two approaches: from curvature tensor and metric coefficients.
result Exact expressions for Drey-Streubel angular momentum of a general section.
In this article, we consider the limit of quasi-local conserved quantities [31,9] at the infinity of an asymptotically hyperbolic initial data set in general relativity. These give notions of total energy-momentum, angular momentum, and center of mass. Our assumption on the asymptotics is less stringent than any previo…
Momentum is a simple and widely used trick which allows gradient-based optimizers to pick up speed along low curvature directions. Its performance depends crucially on a damping coefficient β. Large β values can potentially deliver much larger speedups, but are prone to oscillations and instability; hence one typic…
CoolMomentum combines momentum and Simulated Annealing for deep learning optimization.
problem Global optimization of non-convex functions in deep learning.
method Discretized Langevin dynamics with Simulated Annealing.
result CoolMomentum achieves high accuracy on Resnet-20 on Cifar-10 and Efficientnet-B0 on Imagenet.
It is common practice to decay the learning rate. Here we show one can usually obtain the same learning curve on both training and test sets by instead increasing the batch size during training. This procedure is successful for stochastic gradient descent (SGD), SGD with momentum, Nesterov momentum, and Adam. It reache…
The study finds that factor momentum is significant only at short lags compared to stock momentum.
problem Investigating the relationship between factor momentum and stock momentum.
method Replicated earlier findings and conducted a spanning test controlling for stock momentum and factor exposure.
result Factor momentum is significant only at short lags after controlling for stock momentum and factor exposure.
In stochastic gradient descent, especially for neural network training, there are currently dominating first order methods: not modeling local distance to minimum. This information required for optimal step size is provided by second order methods, however, they have many difficulties, starting with full Hessian having…
We test the price momentum effect in the Korean stock markets under the momentum universe shrinkage to subuniverses of the KOSPI 200. Performance of the momentum strategy is not homogeneous with respect to change of the momentum universe. It is found that some submarkets generate the higher momentum returns than other …
In the current paper the Lagrangian of a classical, relativistic point particle is obtained whose conjugate momentum satisfies the dispersion relation of a quantum wave packet that is subject to Lorentz violation based on a particular coefficient of the nonminimal Standard-Model Extension (SME). The properties of this …
Introduces homotopy momentum sections on multisymplectic manifolds.
problem No specific problem stated; focuses on introducing a new concept.
method Introduces a new concept of homotopy momentum sections on multisymplectic manifolds.
result Shows that a gauged nonlinear sigma model with Wess-Zumino term has homotopy momentum section structure.
Customer momentum is a positive relationship between a firm's returns and past returns of its customers.
problem Understanding the relationship between a firm's returns and its customers' past returns.
method Examined customer momentum using a long-short equally-weighted decile portfolio and Fama-French factor models.
result Customer momentum generates significant monthly returns and is statistically significant.
This paper revisits the classic iterative proportional scaling (IPS) from a modern optimization perspective. In contrast to the criticisms made in the literature, we show that based on a coordinate descent characterization, IPS can be slightly modified to deliver coefficient estimates, and from a majorization-minimizat…
This paper examines momentum spillover across multiple asset classes using only pricing data.
problem Challenges in studying momentum spillover across diverse asset classes due to lack of common characteristics.
method Utilised a linear and interpretable graph learning model to reveal momentum spillover network.
result Network momentum strategy yields a Sharpe ratio of 1.5 and an annual return of 22%.
New formula extracts full local information from ray transform data.
problem Determining symmetric tensor fields from ray transform data.
method Deriving explicit formula for Saint Venant operator.
result Explicit formula for extracting full local information.
This paper presents generalized momentum mappings for covariant Hamiltonian field theories. The new momentum mappings arise from a generalization of symplectic geometry to LVY, the bundle of vertically adapted linear frames over the bundle of field configurations Y. Specifically, the generalized field momentum obs…
We give a detailed discussion about existence and uniqueness of Lu's momentum map. More precisely, we introduce the infinitesimal momentum map, and we study its properties. This allows us to describe the theory of reconstruction of the momentum map from the infinitesimal one. We provide the conditions for the uniquenes…
Momentum ResNets improve ResNets' memory efficiency.
problem Memory inefficiency in deep residual neural networks (ResNets).
method Adding a momentum term to the forward rule of ResNets to make them invertible.
result Momentum ResNets can learn any linear mapping up to a multiplicative factor, improving memory efficiency.
The paper analyzes how momentum affects convergence in stochastic gradient methods.
problem Lack of clear understanding of momentum's impact on convergence and performance.
method Unified analysis of several popular algorithms using the QHM formulation.
result Provides practical guidelines for setting learning rate and momentum parameters.
We introduce various quantitative and mathematical definitions for price momentum of financial instruments. The price momentum is quantified with velocity and mass concepts originated from the momentum in physics. By using the physical momentum of price as a selection criterion, the weekly contrarian strategies are imp…
Adapting momentum from optimization to reinforcement learning.
problem Improving the convergence and stability of reinforcement learning algorithms.
method Introducing Momentum Value Iteration (MoVI) by incorporating an average of consecutive state-action value functions, inspired by the concept of momentum in optimization.
result MoVI improves the convergence and stability of reinforcement learning algorithms, as demonstrated by experiments on Atari games.
New algorithm Momentum-QNG improves optimization of quantum circuits.
problem Optimizing variational quantum circuits to avoid local minima.
method Applied Langevin dynamics to QNG, introducing momentum term.
result Momentum-QNG outperforms basic QNG and other optimizers.
We demonstrate the possibility of what we call sparse learning: accelerated training of deep neural networks that maintain sparse weights throughout training while achieving dense performance levels. We accomplish this by developing sparse momentum, an algorithm which uses exponentially smoothed gradients (momentum) to…
Momentum speeds up evolutionary processes in machine learning.
problem Accelerating convergence in evolutionary dynamics.
method Combining momentum from machine learning with evolutionary dynamics using information divergences as Lyapunov functions.
result Momentum accelerates convergence of evolutionary dynamics, including the replicator equation and Euclidean gradient descent.
Contact manifolds' momentum polytopes are convex.
problem Understanding the structure of contact manifolds.
method Using isomorphism to toric varieties.
result Momentum polytopes of contact manifolds are convex.
Unified model learns from both time-series and cross-sectional momentum features.
problem Separate time-series and cross-sectional momentum strategies do not consider concurrent relationships.
method Spatio-Temporal Momentum strategies using neural networks to combine both types of momentum.
result Simple neural network with single fully connected layer generates trading signals for all assets.
SMG combines shuffling and momentum for non-convex optimization.
problem Non-convex finite-sum optimization problems.
method Shuffling Gradient-based method with momentum.
result Established state-of-the-art convergence rates for SMG.
New method shows stochastic momentum can converge quickly on optimization problems.
problem Improving convergence of stochastic optimization methods.
method Stochastic heavy ball momentum with minibatching.
result Stochastic heavy ball momentum retains fast linear rate on quadratic problems.
The paper analyzes dynamics of momentum in high dimensions with sparse updates.
problem Theoretical analysis of momentum dynamics in high-dimensional sparse settings.
method Theoretical analysis of two models: least squares with sparse inputs and logistic regression with a rare class.
result Characterization of high-dimensional limits of momentum dynamics and phase structure.
Optimization algorithms with momentum, e.g., (ADAM), have been widely used for building deep learning models due to the faster convergence rates compared with stochastic gradient descent (SGD). Momentum helps accelerate SGD in the relevant directions in parameter updating, which can minify the oscillations of parameter…
One has not any conventional energy-momentum conservation law in Lagrangian field theory, but relations involving different stress-energy-momentum tensors associated with different connections. It is not obvious how to choose the true energy-momentum tensor. This problem is solved in the framework of the multimomentum …
The paper extends a theorem about momentum maps to singular symplectic spaces.
problem Extending a theorem about momentum maps to singular symplectic spaces.
method Using integral affine stratification and equivariant locally trivial fibrations, the paper extends the linear variation theorem to singular values of the momentum map.
result Cohomology classes of symplectic forms on reduced spaces vary linearly within strata.
This paper analyzes momentum Q-learning with finite-sample guarantees.
problem Improving Q-learning performance with momentum schemes.
method Proposes MomentumQ algorithm integrating Nesterov and Polyak's momentum schemes, analyzes convergence for function approximations.
result Establishes finite-sample convergence rates for MomentumQ, demonstrating better performance than vanilla Q-learning.
Momentum improves deep learning generalization by stabilizing noise and learning features.
problem Improving generalization in deep learning models.
method Empirical and theoretical analysis of gradient descent with momentum (GD+M) vs. gradient descent (GD) in binary classification tasks.
result GD+M outperforms GD in generalization, especially in datasets with shared features and varying margins.
Generalizes momentum map to Courant algebroid for constrained mechanics.
problem Generalizing momentum map to new geometric structures.
method Generalized momentum section on Lie algebroid to Courant algebroid, constructed cohomological formulations.
result Identified momentum section in constrained Hamiltonian mechanics with Courant algebroid symmetry.
Momentum SGD fails to track nonstationary optima due to drift amplification.
problem Tracking nonstationary optima in stochastic optimization.
method Theoretical analysis of SGD and momentum variants under strong convexity and smoothness.
result Momentum incurs a drift-amplification penalty that diverges as the momentum parameter approaches 1, leading to systematic lag.
This paper concentrates on the time series momentum or contrarian effects in the Chinese stock market. We evaluate the performance of the time series momentum strategy applied to major stock indices in mainland China and explore the relation between the performance of time series momentum strategies and some firm-speci…
A new accelerated method with simpler momentum update rules.
problem Optimizing parameters in machine learning models.
method Proposes a novel accelerated stochastic gradient method with simpler momentum update rules.
result The method outperforms Sgdm and Adam in practical problems.
New method uses momentum to converge in DC optimization with small batches.
problem Lack of convergence properties for stochastic difference-of-convex optimization with small batch sizes.
method Introduces momentum to enable convergence under standard assumptions for any batch size.
result Proves convergence of the algorithm under smoothness and bounded variance assumptions.
Dynamic econometric models improve trading signals in momentum strategies.
problem Static momentum strategies are inefficient; dynamic models enhance accuracy.
method Dynamic binary classifier model to learn time-varying momentum importance.
result Dynamic classifier outperforms traditional naive time series momentum strategy.
Stochastic proximal point algorithm with momentum converges faster and is more stable than standard methods.
problem Improving convergence and stability of stochastic optimization methods.
method Developed and analyzed the convergence and stability of the stochastic proximal point algorithm with momentum (SPPAM).
result SPPAM converges faster and is more stable than standard stochastic proximal point algorithm (SPPA) and stochastic gradient descent with momentum (SGDM).
Negative momentum accelerates convergence in minimax games but at a suboptimal rate.
problem The convergence rate of negative momentum in minimax games is suboptimal.
method Extending variational inequality formulation, connecting momentum method with Chebyshev polynomials.
result Negative momentum accelerates convergence locally but at a suboptimal rate.
In this paper we study several classes of stochastic optimization algorithms enriched with heavy ball momentum. Among the methods studied are: stochastic gradient descent, stochastic Newton, stochastic proximal point and stochastic dual subspace ascent. This is the first time momentum variants of several of these metho…
Lion optimizer performs well in training AI models with memory efficiency.
problem Lion optimizer's theoretical basis is unclear.
method Continuous-time and discrete-time analysis of Lion updates with a new Lyapunov function.
result Lion is a novel and principled approach for constrained optimization.
Two major momentum-based techniques that have achieved tremendous success in optimization are Polyak's heavy ball method and Nesterov's accelerated gradient. A crucial step in all momentum-based methods is the choice of the momentum parameter m which is always suggested to be set to less than 1. Although the choice…
DeepUnifiedMom uses deep learning to create better momentum portfolios.
problem Lack of unified momentum portfolios across different time frames.
method Multi-task learning with multi-gate mixture of experts.
result DeepUnifiedMom outperforms benchmark models in diverse asset classes.
This study uses continuous-time analysis to understand how momentum affects the optimisation of diagonal linear networks.
problem The effect of momentum on the optimisation trajectory of gradient descent.
method Leveraging a continuous-time approach to analyze momentum gradient descent with step size γ and momentum parameter β.
result Small values of λ help recover sparse solutions in overparametrised regression settings.