New method shows stochastic momentum can converge quickly on optimization problems.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Heavy Ball method speeds up finding global optima in non-convex problems.
This paper analyzes two Lie group momentum optimization algorithms and their convergence rates.
In this work we establish the first linear convergence result for the stochastic heavy ball method. The method performs SGD steps with a fixed stepsize, amended by a heavy ball momentum term. In the analysis, we focus on minimizing the expected loss and not on finite-sum minimization, which is typically a much harder p…
Arguably, the two most popular accelerated or momentum-based optimization methods in machine learning are Nesterov's accelerated gradient and Polyaks's heavy ball, both corresponding to different discretizations of a particular second order differential equation with friction. Such connections with continuous-time dyna…
Proposes momentum methods for Lie groups, improving on classical algorithms.
Study accelerates optimization methods in non-convex problems, but doesn't improve the algorithm's performance.
A new Bayesian filtering method speeds up stochastic Newton optimization.
In this paper we study several classes of stochastic optimization algorithms enriched with heavy ball momentum. Among the methods studied are: stochastic gradient descent, stochastic Newton, stochastic proximal point and stochastic dual subspace ascent. This is the first time momentum variants of several of these metho…
Fine-grained analysis of gradient descent with momentum provides modified loss equations.
Recently, {\it stochastic momentum} methods have been widely adopted in training deep neural networks. However, their convergence analysis is still underexplored at the moment, in particular for non-convex optimization. This paper fills the gap between practice and theory by developing a basic convergence analysis of t…
The paper analyzes dynamics of momentum in high dimensions with sparse updates.
The paper analyzes how momentum affects convergence in stochastic gradient methods.
Paper analyzes SHB method for neural networks, proving stability, connectivity, and global convergence.
We take a Hamiltonian-based perspective to generalize Nesterov's accelerated gradient descent and Polyak's heavy ball method to a broad class of momentum methods in the setting of (possibly) constrained minimization in Euclidean and non-Euclidean normed vector spaces. Our perspective leads to a generic and unifying non…
Two major momentum-based techniques that have achieved tremendous success in optimization are Polyak's heavy ball method and Nesterov's accelerated gradient. A crucial step in all momentum-based methods is the choice of the momentum parameter which is always suggested to be set to less than . Although the choice…
Improved SHB method for faster convergence on strongly-convex quadratics.
A novel decentralized deep learning algorithm using gradient-based optimization.
Momentum SGD fails to track nonstationary optima due to drift amplification.
Analysis of momentum methods on quadratic models, showing SGD's superiority.
While momentum-based accelerated variants of stochastic gradient descent (SGD) are widely used when training machine learning models, there is little theoretical understanding on the generalization error of such methods. In this work, we first show that there exists a convex loss function for which the stability gap fo…
Integrating adaptive learning rate and momentum techniques into SGD leads to a large class of efficiently accelerated adaptive stochastic algorithms, such as AdaGrad, RMSProp, Adam, AccAdaGrad, \textit{etc}. In spite of their effectiveness in practice, there is still a large gap in their theories of convergences, espec…
New methods solve optimization problems with heavy-tailed noise, improving upon existing complexity bounds.
Optimizes web page freshness with limited crawling frequencies.
Nonconvex optimization algorithms with random initialization have attracted increasing attention recently. It has been showed that many first-order methods always avoid saddle points with random starting points. In this paper, we answer a question: can the nonconvex heavy-ball algorithms with random initialization avoi…
Stochastic momentum methods have been widely adopted in training deep neural networks. However, their theoretical analysis of convergence of the training objective and the generalization error for prediction is still under-explored. This paper aims to bridge the gap between practice and theory by analyzing the stochast…
New insights into training machine learning models with momentum.
The paper analyzes convergence rates for SGD and SHB methods.
Rod flow models Adam's behavior at the edge of stability.
New method improves optimization and DP in FL.
Unified algorithm for stochastic optimization with time-varying momentum converges under general conditions.
Paper extends Noether's Theorem to nonholonomic systems, proving conserved momentum.
This paper deals with a natural stochastic optimization procedure derived from the so-called Heavy-ball method differential equation, which was introduced by Polyak in the 1960s with his seminal contribution [Pol64]. The Heavy-ball method is a second-order dynamics that was investigated to minimize convex functions f .…
Work on SGDm under heavy-tailed noise, revealing its generalization properties.
Standard optimizers perform as well as LARS and LAMB at large batch sizes.
Gradient descent-based optimization methods underpin the parameter training of neural networks, and hence comprise a significant component in the impressive test results found in a number of applications. Introducing stochasticity is key to their success in practical problems, and there is some understanding of the rol…
Momentum methods such as Polyak's heavy ball (HB) method, Nesterov's accelerated gradient (AG) as well as accelerated projected gradient (APG) method have been commonly used in machine learning practice, but their performance is quite sensitive to noise in the gradients. We study these methods under a first-order stoch…
Paper studies stochastic optimization methods with momentum, proving convergence and avoiding traps.
Stochastic momentum methods trade compute efficiency for serial runtime.
Study accelerates gradient methods in machine learning, revealing risk and stability connections.
A method for estimating the median of gradients in stochastic optimization.
In this paper, we revisit the convergence of the Heavy-ball method, and present improved convergence complexity results in the convex setting. We provide the first non-ergodic O(1/k) rate result of the Heavy-ball algorithm with constant step size for coercive objective functions. For objective functions satisfying a re…
The choice of how to retain information about past gradients dramatically affects the convergence properties of state-of-the-art stochastic optimization methods, such as Heavy-ball, Nesterov's momentum, RMSprop and Adam. Building on this observation, we use stochastic differential equations (SDEs) to explicitly study t…
New step-size methods improve SHB convergence for stochastic optimization.
The vast majority of successful deep neural networks are trained using variants of stochastic gradient descent (SGD) algorithms. Recent attempts to improve SGD can be broadly categorized into two approaches: (1) adaptive learning rate schemes, such as AdaGrad and Adam, and (2) accelerated schemes, such as heavy-ball an…
New method targets deep learning optima with heavy-tailed noise.
Derives new optimization methods using variational integrators.
Momentum based stochastic gradient methods such as heavy ball (HB) and Nesterov's accelerated gradient descent (NAG) method are widely used in practice for training deep networks and other supervised learning models, as they often provide significant improvements over stochastic gradient descent (SGD). Rigorously speak…