We describe stochastic Newton and stochastic quasi-Newton approaches to efficiently solve large linear least-squares problems where the very large data sets present a significant computational burden (e.g., the size may exceed computer memory or data are collected in real-time). In our proposed framework, stochasticity…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We introduce a framework for Newton's flows in probability space with information metrics, named information Newton's flows. Here two information metrics are considered, including both the Fisher-Rao metric and the Wasserstein-2 metric. A known fact is that overdamped Langevin dynamics correspond to Wasserstein gradien…
Newton's method solves variational problems on manifolds.
Muon with Newton-Schulz converges to the same stationary point as SVD-polar, up to a constant factor.
RNN operators solve Newton's equations with large timesteps for molecular dynamics.
We generalize Newton-type methods for minimizing smooth functions to handle a sum of two convex functions: a smooth function and a nonsmooth function with a simple proximal mapping. We show that the resulting proximal Newton-type methods inherit the desirable convergence behavior of Newton-type methods for minimizing s…
New algorithm improves convergence of gradient boosting trees.
Boosting algorithms are frequently used in applied data science and in research. To date, the distinction between boosting with either gradient descent or second-order Newton updates is often not made in both applied and methodological research, and it is thus implicitly assumed that the difference is irrelevant. The g…
Newton's method tackles nonlinear mappings into vector bundles with connections and retractions.
Study uses Newton polytopes to distinguish Lagrangian fillings of Legendrian submanifolds.
Unified approach to Bayesian inference with guarantees on covariance matrices.
Paper proposes an online covariance estimator for sketched Newton methods.
New Q-Newton's method avoids saddle points and converges quadratically.
A new optimization method improves deep learning accuracy without hyper-parameter tuning.
This thesis disentangles Gauss-Newton and variational approximations in Bayesian deep learning.
SVRN accelerates Newton methods by reducing variance and improving performance.
In this article we present a natural generalization of Newton's Second Law valid in field theory, i.e., when the parameterized curves are replaced by parameterized submanifolds of higher dimension. For it we introduce what we have called the geodesic -vector field, analogous to the ordinary geodesic field and which …
The second order method as Newton Step is a suitable technique in Online Learning to guarantee regret bound. The large data is a challenge in Newton method to store second order matrices as hessian. In this paper, we have proposed an modified online Newton step that store first and second order matrices of dimension m …
Deep learning involves a difficult non-convex optimization problem, which is often solved by stochastic gradient (SG) methods. While SG is usually effective, it may not be robust in some situations. Recently, Newton methods have been investigated as an alternative optimization technique, but nearly all existing studies…
Study of measured laminations on surfaces using Newton polytopes and Poisson brackets.
Approximate Newton methods are a standard optimization tool which aim to maintain the benefits of Newton's method, such as a fast rate of convergence, whilst alleviating its drawbacks, such as computationally expensive calculation or estimation of the inverse Hessian. In this work we investigate approximate Newton meth…
We present two new remarkably simple stochastic second-order methods for minimizing the average of a very large number of sufficiently smooth and strongly convex functions. The first is a stochastic variant of Newton's method (SN), and the second is a stochastic variant of cubically regularized Newton's method (SCN). W…
Newton-LESS sparsifies Gaussian sketching for faster optimization.
New quasi-Newton method guarantees global superlinear convergence.
Paper develops a robust PP distributed quasi-Newton estimation for Byzantine machines.
Four decades after their invention, quasi-Newton methods are still state of the art in unconstrained numerical optimization. Although not usually interpreted thus, these are learning algorithms that fit a local quadratic approximation to the objective function. We show that many, including the most popular, quasi-Newto…
A new quasi-Newton method uses cubic regularization to avoid saddle points in deep learning.
Extends Newton's minimal resistance problem to Riemannian surfaces.
A new Bayesian filtering method speeds up stochastic Newton optimization.
The Gauss-Newton method is analyzed for neural networks using Riemannian optimization techniques.
Extends Newton's minimal resistance problem to Lorentz-Minkowski space.
ISAAC Newton uses input-based curvature for efficient training.
Improved root-finding method for smooth functions.
A new method for machine learning updates reduces complexity and improves robustness.
Paper proposes an efficient online Newton method with Nesterov's acceleration for streaming data.
A new method solves distributed optimization problems over networks.
In [19], a general, inexact, efficient proximal quasi-Newton algorithm for composite optimization problems has been proposed and a sublinear global convergence rate has been established. In this paper, we analyze the convergence properties of this method, both in the exact and inexact setting, in the case when the obje…
A new method for faster optimization on statistical manifolds.
In this work, we present a globalized stochastic semismooth Newton method for solving stochastic optimization problems involving smooth nonconvex and nonsmooth convex terms in the objective function. We assume that only noisy gradient and Hessian information of the smooth part of the objective function is available via…
Recently algorithms incorporating second order curvature information have become popular in training neural networks. The Nesterov's Accelerated Quasi-Newton (NAQ) method has shown to effectively accelerate the BFGS quasi-Newton method by incorporating the momentum term and Nesterov's accelerated gradient vector. A sto…
Paper proposes a new method to efficiently incorporate curvature information in stochastic optimization.
For distributed computing environment, we consider the empirical risk minimization problem and propose a distributed and communication-efficient Newton-type optimization method. At every iteration, each worker locally finds an Approximate NewTon (ANT) direction, which is sent to the main driver. The main driver, then, …
A new method avoids saddle points in Newton's method.
A new hybrid Newton algorithm improves convergence in logistic regression.
In this work, we discuss graph like image of curves under moment maps and their relation with the Newton polygon of the curve, which has applications to Lagrangian torus fibration of Calabi-Yau manifolds.
Recent studies incorporate Nesterov's accelerated gradient method for the acceleration of gradient based training. The Nesterov's Accelerated Quasi-Newton (NAQ) method has shown to drastically improve the convergence speed compared to the conventional quasi-Newton method. This paper implements NAQ for non-convex optimi…
We develop a non-relativistic twistor theory, in which Newton--Cartan structures of Newtonian gravity correspond to complex three-manifolds with a four-parameter family of rational curves with normal bundle . We show that the Newton--Cartan space-times are unstable under the general K…
A general class of Newton algorithms on Graßmann and Lagrange-Graßmann manifolds is introduced, that depends on an arbitrary pair of local coordinates. Local quadratic convergence of the algorithm is shown under a suitable condition on the choice of coordinate systems. Our result extends and unifies previous convergenc…