Study shows how steepest descent algorithms' geometric margin increases during training.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We develop a progressive training approach for neural networks which adaptively grows the network structure by splitting existing neurons to multiple off-springs. By leveraging a functional steepest descent idea, we derive a simple criterion for deciding the best subset of neurons to split and a splitting gradient for …
This work studies the implicit bias of mini-batch SGD in classification.
Let S be a complete surface of constant curvature K = + 1 or -1, i.e. the sphere S^2 or the Lobachevskij plane L^2, and D a bounded convex subset of S. If S = S^2, assume also diameter (D) < pi/2. It is proved that the length of any steepest descent curve of a quasi-convex function in D is less than or equal to the per…
New method escapes local optima in neural architecture optimization.
This letter proposes a sparse diffusion steepest-descent algorithm for one bit compressed sensing in wireless sensor networks. The approach exploits the diffusion strategy from distributed learning in the one bit compressed sensing framework. To estimate a common sparse vector cooperatively from only the sign of measur…
Study shows momentum-based optimizers like Muon and MomentumGD bias towards KKT points in smooth homogeneous models.
New method accelerates steepest descent for convex optimization.
Improved greedy 2-coordinate updates for optimization problems with constraints.
Two popular examples of first-order optimization methods over linear spaces are coordinate descent and matching pursuit algorithms, with their randomized variants. While the former targets the optimization by moving along coordinates, the latter considers a generalized notion of directions. Exploiting the connection be…
GANs excel at learning high dimensional distributions, but they can update generator parameters in directions that do not correspond to the steepest descent direction of the objective. Prominent examples of problematic update directions include those used in both Goodfellow's original GAN and the WGAN-GP. To formally d…
Method approximates Riemannian barycenter on manifolds.
The paper optimizes regret using covariance between costs and decisions.
Straight lines are a basin of attraction for the elastic flow at least to level 1.9615π.
In optimization, the negative gradient of a function denotes the direction of steepest descent. Furthermore, traveling in any direction orthogonal to the gradient maintains the value of the function. In this work, we show that these orthogonal directions that are ignored by gradient descent can be critical in equilibri…
Adaptive step-size improves optimization in complex geometries.
Logistic regression is one of the most popular methods in binary classification, wherein estimation of model parameters is carried out by solving the maximum likelihood (ML) optimization problem, and the ML estimator is defined to be the optimal solution of this problem. It is well known that the ML estimator exists wh…
Study uses outer metrics for PDE-constrained shape optimization over diffeomorphism group.
We study the implicit bias of generic optimization methods, such as mirror descent, natural gradient descent, and steepest descent with respect to different potentials and norms, when optimizing underdetermined linear regression or separable linear classification problems. We explore the question of whether the specifi…
Study identifies stable configurations of intertwined threads with repulsive interactions.
SONIA optimizes machine learning problems with a novel algorithm.
A quantum generalization of Natural Gradient Descent is presented as part of a general-purpose optimization framework for variational quantum circuits. The optimization dynamics is interpreted as moving in the steepest descent direction with respect to the Quantum Information Geometry, corresponding to the real part of…
New method optimizes SDE models using continuous-time gradient descent.
We present here a new model and algorithm which performs an efficient Natural gradient descent for Multilayer Perceptrons. Natural gradient descent was originally proposed from a point of view of information geometry, and it performs the steepest descent updates on manifolds in a Riemannian space. In particular, we ext…
Paper accelerates Bayesian few-shot classification using mirror descent.
Scalable and robust TR decomposition for large-scale data with missing entries and outliers.
New optimizers control network width scaling, improving stability and transfer across different model sizes.
This paper presents a stochastic behavior analysis of a kernel-based stochastic restricted-gradient descent method. The restricted gradient gives a steepest ascent direction within the so-called dictionary subspace. The analysis provides the transient and steady state performance in the mean squared error criterion. It…
A new algorithm for training generative models using Sinkhorn divergence.
New algorithm optimizes nonlinear SDEs online with convergence guarantees.
New RL algorithms learn policies competitive with best in class without assuming optimal policy.
AdamW optimizes a constrained loss with norm constraint.
Bayesian inference problems require sampling or approximating high-dimensional probability distributions. The focus of this paper is on the recently introduced Stein variational gradient descent methodology, a class of algorithms that rely on iterated steepest descent steps with respect to a reproducing kernel Hilbert …
We revisit the choice of SGD for training deep neural networks by reconsidering the appropriate geometry in which to optimize the weights. We argue for a geometry invariant to rescaling of weights that does not affect the output of the network, and suggest Path-SGD, which is an approximate steepest descent method with …
This paper presents the application of a newly developed nature-inspired metaheuristic optimization method, namely the Adaptive Wind Driven Optimization (AWDO), to the training of feedforward artificial neural networks (NN) and presents a discussion into the future research of AWDO implementation in Deep Learning (DL).…
The affine Grassmannian is a noncompact smooth manifold that parameterizes all affine subspaces of a fixed dimension. It is a natural generalization of Euclidean space, points being zero-dimensional affine subspaces. We will realize the affine Grassmannian as a matrix manifold and extend Riemannian optimization algorit…
Characterizes corridors in loss surfaces for gradient-based optimization.
Gradient flow expands curves to round shapes.
Sign-based optimization methods have become popular in machine learning due to their favorable communication cost in distributed optimization and their surprisingly good performance in neural network training. Furthermore, they are closely connected to so-called adaptive gradient methods like Adam. Recent works on sign…
Real flag manifolds are the isotropy orbits of noncompact symmetric spaces . Any such manifold enjoys two very peculiar geometric properties: It carries a transitive action of the (noncompact) Lie group , and it is embedded in euclidean space as a taut submanifold. The aim of the paper is to link these two …
This letter presents an improved version of diffusion least mean ppower (LMP) algorithm for distributed estimation. Instead of sum of mean square errors, a weighted sum of mean square error is defined as the cost function for global and local cost functions of a network of sensors. The weight coefficients are updated b…
Information geometry applies concepts in differential geometry to probability and statistics and is especially useful for parameter estimation in exponential families where parameters are known to lie on a Riemannian manifold. Connections between the geometric properties of the induced manifold and statistical properti…
New Muon and Momo variants improve neural network optimization robustness.
New optimization method combines gradient clipping and non-Euclidean smoothness.
New data-driven Cartan connection tracks complex vascular structures.
Coordinate descent with random coordinate selection is the current state of the art for many large scale optimization problems. However, greedy selection of the steepest coordinate on smooth problems can yield convergence rates independent of the dimension , and requiring upto times fewer iterations. In this pap…
Following Feynman's prescription for constructing a path integral representation of the propagator of a quantum theory, a short-time approximation to the propagator for imaginary time, N=1 supersymmetric quantum mechanics on a compact, even-dimensional Riemannian manifold is constructed. The path integral is interprete…
This letter proposes a dictionary learning algorithm for blind one bit compressed sensing. In the blind one bit compressed sensing framework, the original signal to be reconstructed from one bit linear random measurements is sparse in an unknown domain. In this context, the multiplication of measurement matrix $\Ab$ an…