New methods using natural gradient for structured optimization.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Enhances gradient estimates for Hermitian Monge-Ampère equations.
New algorithms solve complex minimax problems efficiently.
In this paper, we study stochastic non-convex optimization with non-convex random functions. Recent studies on non-convex optimization revolve around establishing second-order convergence, i.e., converging to a nearly second-order optimal stationary points. However, existing results on stochastic non-convex optimizatio…
Natural gradient descent, which preconditions a gradient descent update with the Fisher information matrix of the underlying statistical model, is a way to capture partial second-order information. Several highly visible works have advocated an approximation known as the empirical Fisher, drawing connections between ap…
Improved SVRG method using BB techniques for faster convergence.
In this paper, we provide an overview of first-order and second-order variants of the gradient descent method that are commonly used in machine learning. We propose a general framework in which 6 of these variants can be interpreted as different instances of the same approach. They are the vanilla gradient descent, the…
Deep neural networks are usually trained with stochastic gradient descent (SGD), which minimizes objective function using very rough approximations of gradient, only averaging to the real gradient. Standard approaches like momentum or ADAM only consider a single direction, and do not try to model distance from extremum…
Variance reduction techniques like SVRG provide simple and fast algorithms for optimizing a convex finite-sum objective. For nonconvex objectives, these techniques can also find a first-order stationary point (with small gradient). However, in nonconvex optimization it is often crucial to find a second-order stationary…
Second-order methods improve differential privacy in convex optimization.
We analyze stochastic gradient algorithms for optimizing nonconvex problems. In particular, our goal is to find local minima (second-order stationary points) instead of just finding first-order stationary points which may be some bad unstable saddle points. We show that a simple perturbed version of stochastic recursiv…
Gradient descent and its variants are widely used in machine learning. However, oracle access of gradient may not be available in many applications, limiting the direct use of gradient descent. This paper proposes a method of estimating gradient to perform gradient descent, that converges to a stationary point for gene…
New SGD algorithm finds critical points faster with second-order corrections.
Improved robustness in optimization methods using second-order information.
SLEDGE algorithm reduces gradient computation errors in optimization.
A new algorithm for solving constrained convex optimization problems efficiently.
Paper improves compressed SGD to reach second-order stationary points.
Trust region and cubic regularization methods have demonstrated good performance in small scale non-convex optimization, showing the ability to escape from saddle points. Each iteration of these methods involves computation of gradient, Hessian and function value in order to obtain the search direction and adjust the r…
Paper derives estimates for complex Hessian equations on Hermitian manifolds.
New algorithm TURTLE outperforms MAML and meta-learner LSTM.
The study provides interior estimates for -flows and translators in .
This paper concerns local gradient estimates to solutions of general conformally invariant fully nonlinear elliptic equations of second order.
PDHAMS improves sampling for discrete distributions with quadratic potential functions.
Improved algorithm finds second-order stationary points in non-convex optimization.
Second-order optimizers retain residual information after data deletion, affecting machine unlearning.
Simple gradient descent algorithm escapes saddle points efficiently.
Finite-sum optimization problems are ubiquitous in machine learning, and are commonly solved using first-order methods which rely on gradient computations. Recently, there has been growing interest in \emph{second-order} methods, which rely on both gradients and Hessians. In principle, second-order methods can require …
Single-site Markov Chain Monte Carlo (MCMC) is a variant of MCMC in which a single coordinate in the state space is modified in each step. Structured relational models are a good candidate for this style of inference. In the single-site context, second order methods become feasible because the typical cubic costs assoc…
Recent years have seen increased interest in performance guarantees of gradient descent algorithms for non-convex optimization. A number of works have uncovered that gradient noise plays a critical role in the ability of gradient descent recursions to efficiently escape saddle-points and reach second-order stationary p…
EvoGrad improves efficiency in meta-learning and hyperparameter optimization.
Optimistic method adapted for faster convex-concave min-max problems.
Develops first and second-order pseudo-mirror descent methods for nonnegative function estimation.
New method speeds up deep learning optimization.
Enhances SMC² with Hessian info for more efficient posterior approximation.
Second-order guarantees for federated learning algorithms.
New method learns population dynamics from snapshots, outperforming existing models.
PWGF escapes saddle points in nonconvex optimization.
We consider the case of derivative-free algorithms for non-convex optimization, also known as zero order algorithms, that use only function evaluations rather than gradients. For a wide variety of gradient approximators based on finite differences, we establish asymptotic convergence to second order stationary points u…
Online learning with limited information feedback (bandit) tries to solve the problem where an online learner receives partial feedback information from the environment in the course of learning. Under this setting, Flaxman et al.[8] extended Zinkevich's classical Online Gradient Descent (OGD) algorithm [29] by proposi…
A new optimizer for deep learning improves accuracy and reduces training time.
Incorporating second order curvature information in gradient based methods have shown to improve convergence drastically despite its computational intensity. In this paper, we propose a stochastic (online) quasi-Newton method with Nesterov's accelerated gradient in both its full and limited memory forms for solving lar…
New method improves online covariance estimation for SGD.
First-order stochastic methods are the state-of-the-art in large-scale machine learning optimization owing to efficient per-iteration complexity. Second-order methods, while able to provide faster convergence, have been much less explored due to the high cost of computing the second-order information. In this paper we …
New method provides tighter robustness guarantees for adversarial attacks.
The paper analyzes second-order guarantees for optimization in various architectures.
Improved computational complexity in statistical models using second-order information.
Stein variational gradient descent (SVGD) was recently proposed as a general purpose nonparametric variational inference algorithm [Liu & Wang, NIPS 2016]: it minimizes the Kullback-Leibler divergence between the target distribution and its approximation by implementing a form of functional gradient descent on a reprod…
Estimates for complex equations on manifolds derived from a conjecture.