Momentum is a popular technique to accelerate the convergence in practical training, and its impact on convergence guarantee has been well-studied for first-order algorithms. However, such a successful acceleration technique has not yet been proposed for second-order algorithms in nonconvex optimization.In this paper, …
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper develops momentum schemes with variance reduction for non-convex composition optimization.
A new Bayesian modeling method is proposed by combining the maximization of the marginal likelihood with a momentum-space renormalization group transformation for Gaussian graphical models. Moreover, we present a scheme for computint the statistical averages of hyperparameters and mean square errors in our proposed met…
The study analyzes momentum-based optimization algorithms from dynamical systems perspective.
This paper analyzes momentum Q-learning with finite-sample guarantees.
Tuning hyperparameters of learning algorithms is hard because gradients are usually unavailable. We compute exact gradients of cross-validation performance with respect to all hyperparameters by chaining derivatives backwards through the entire training procedure. These gradients allow us to optimize thousands of hyper…
Integrating adaptive learning rate and momentum techniques into SGD leads to a large class of efficiently accelerated adaptive stochastic algorithms, such as AdaGrad, RMSProp, Adam, AccAdaGrad, \textit{etc}. In spite of their effectiveness in practice, there is still a large gap in their theories of convergences, espec…
This paper analyzes two Lie group momentum optimization algorithms and their convergence rates.
New HMC method uses asymmetrical momentum distributions and improves performance.
SRSGD improves DNN training speed and accuracy.
FedMCC learns from distributed data to cluster and extract features.
New methods reduce constraint violations to certainty in stochastic optimization.
MSGD outperforms SGD in overparametrized settings with faster convergence rates.
SARAH and SPIDER are two recently developed stochastic variance-reduced algorithms, and SPIDER has been shown to achieve a near-optimal first-order oracle complexity in smooth nonconvex optimization. However, SPIDER uses an accuracy-dependent stepsize that slows down the convergence in practice, and cannot handle objec…
Regularized nonlinear acceleration (RNA) estimates the minimum of a function by post-processing iterates from an algorithm such as the gradient method. It can be seen as a regularized version of Anderson acceleration, a classical acceleration scheme from numerical analysis. The new scheme provably improves the rate of …
A new algorithm improves convergence rates for convex optimization problems.
Modern distributed training of machine learning models suffers from high communication overhead for synchronizing stochastic gradients and model parameters. In this paper, to reduce the communication complexity, we propose \emph{double quantization}, a general scheme for quantizing both model parameters and gradients. …
Distributed asynchronous SGD has become widely used for deep learning in large-scale systems, but remains notorious for its instability when increasing the number of workers. In this work, we study the dynamics of distributed asynchronous SGD under the lens of Lagrangian mechanics. Using this description, we introduce …
The study finds that factor momentum is significant only at short lags compared to stock momentum.
In this paper we explore acceleration techniques for large scale nonconvex optimization problems with special focuses on deep neural networks. The extrapolation scheme is a classical approach for accelerating stochastic gradient descent for convex optimization, but it does not work well for nonconvex optimization typic…
We test the price momentum effect in the Korean stock markets under the momentum universe shrinkage to subuniverses of the KOSPI 200. Performance of the momentum strategy is not homogeneous with respect to change of the momentum universe. It is found that some submarkets generate the higher momentum returns than other …
Introduces homotopy momentum sections on multisymplectic manifolds.
Customer momentum is a positive relationship between a firm's returns and past returns of its customers.
This paper examines momentum spillover across multiple asset classes using only pricing data.
The paper analyzes how hyperparameters affect SGD with momentum's convergence rate.
We propose a stepsize adaptation scheme for stochastic gradient descent. It operates directly with the loss function and rescales the gradient in order to make fixed predicted progress on the loss. We demonstrate its capabilities by conclusively improving the performance of Adam and Momentum optimizers. The enhanced op…
This paper presents generalized momentum mappings for covariant Hamiltonian field theories. The new momentum mappings arise from a generalization of symplectic geometry to , the bundle of vertically adapted linear frames over the bundle of field configurations . Specifically, the generalized field momentum obs…
The vast majority of successful deep neural networks are trained using variants of stochastic gradient descent (SGD) algorithms. Recent attempts to improve SGD can be broadly categorized into two approaches: (1) adaptive learning rate schemes, such as AdaGrad and Adam, and (2) accelerated schemes, such as heavy-ball an…
We give a detailed discussion about existence and uniqueness of Lu's momentum map. More precisely, we introduce the infinitesimal momentum map, and we study its properties. This allows us to describe the theory of reconstruction of the momentum map from the infinitesimal one. We provide the conditions for the uniquenes…
Momentum ResNets improve ResNets' memory efficiency.
The paper analyzes how momentum affects convergence in stochastic gradient methods.
We introduce various quantitative and mathematical definitions for price momentum of financial instruments. The price momentum is quantified with velocity and mass concepts originated from the momentum in physics. By using the physical momentum of price as a selection criterion, the weekly contrarian strategies are imp…
This paper studies accelerations in Q-learning algorithms. We propose an accelerated target update scheme by incorporating the historical iterates of Q functions. The idea is conceptually inspired by the momentum-based accelerated methods in the optimization theory. Conditions under which the proposed accelerated algor…
Adapting momentum from optimization to reinforcement learning.
New algorithm Momentum-QNG improves optimization of quantum circuits.
We demonstrate the possibility of what we call sparse learning: accelerated training of deep neural networks that maintain sparse weights throughout training while achieving dense performance levels. We accomplish this by developing sparse momentum, an algorithm which uses exponentially smoothed gradients (momentum) to…
Momentum speeds up evolutionary processes in machine learning.
Optimizes web page freshness with limited crawling frequencies.
Contact manifolds' momentum polytopes are convex.
Unified model learns from both time-series and cross-sectional momentum features.
SMG combines shuffling and momentum for non-convex optimization.
New method shows stochastic momentum can converge quickly on optimization problems.
The paper analyzes dynamics of momentum in high dimensions with sparse updates.
Optimization algorithms with momentum, e.g., (ADAM), have been widely used for building deep learning models due to the faster convergence rates compared with stochastic gradient descent (SGD). Momentum helps accelerate SGD in the relevant directions in parameter updating, which can minify the oscillations of parameter…
One has not any conventional energy-momentum conservation law in Lagrangian field theory, but relations involving different stress-energy-momentum tensors associated with different connections. It is not obvious how to choose the true energy-momentum tensor. This problem is solved in the framework of the multimomentum …
The paper extends a theorem about momentum maps to singular symplectic spaces.
Momentum improves deep learning generalization by stabilizing noise and learning features.
Generalizes momentum map to Courant algebroid for constrained mechanics.