Study shows how steepest descent algorithms' geometric margin increases during training.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We develop a progressive training approach for neural networks which adaptively grows the network structure by splitting existing neurons to multiple off-springs. By leveraging a functional steepest descent idea, we derive a simple criterion for deciding the best subset of neurons to split and a splitting gradient for …
This work studies the implicit bias of mini-batch SGD in classification.
Let S be a complete surface of constant curvature K = + 1 or -1, i.e. the sphere S^2 or the Lobachevskij plane L^2, and D a bounded convex subset of S. If S = S^2, assume also diameter (D) < pi/2. It is proved that the length of any steepest descent curve of a quasi-convex function in D is less than or equal to the per…
This letter proposes a sparse diffusion steepest-descent algorithm for one bit compressed sensing in wireless sensor networks. The approach exploits the diffusion strategy from distributed learning in the one bit compressed sensing framework. To estimate a common sparse vector cooperatively from only the sign of measur…
Study shows momentum-based optimizers like Muon and MomentumGD bias towards KKT points in smooth homogeneous models.
New method escapes local optima in neural architecture optimization.
New method accelerates steepest descent for convex optimization.
Improved greedy 2-coordinate updates for optimization problems with constraints.
GANs excel at learning high dimensional distributions, but they can update generator parameters in directions that do not correspond to the steepest descent direction of the objective. Prominent examples of problematic update directions include those used in both Goodfellow's original GAN and the WGAN-GP. To formally d…
Method approximates Riemannian barycenter on manifolds.
Straight lines are a basin of attraction for the elastic flow at least to level 1.9615π.
The paper optimizes regret using covariance between costs and decisions.
Study identifies stable configurations of intertwined threads with repulsive interactions.
SONIA optimizes machine learning problems with a novel algorithm.
In optimization, the negative gradient of a function denotes the direction of steepest descent. Furthermore, traveling in any direction orthogonal to the gradient maintains the value of the function. In this work, we show that these orthogonal directions that are ignored by gradient descent can be critical in equilibri…
Study uses outer metrics for PDE-constrained shape optimization over diffeomorphism group.
Paper improves RLS for sparse outlier detection in linear models.
Logistic regression is one of the most popular methods in binary classification, wherein estimation of model parameters is carried out by solving the maximum likelihood (ML) optimization problem, and the ML estimator is defined to be the optimal solution of this problem. It is well known that the ML estimator exists wh…
New optimizers control network width scaling, improving stability and transfer across different model sizes.
Adaptive step-size improves optimization in complex geometries.
Scalable and robust TR decomposition for large-scale data with missing entries and outliers.
A new algorithm for training generative models using Sinkhorn divergence.
New algorithm optimizes nonlinear SDEs online with convergence guarantees.
Unified signSGD and gradient descent analysis for neural networks.
AdamW optimizes a constrained loss with norm constraint.
A quantum generalization of Natural Gradient Descent is presented as part of a general-purpose optimization framework for variational quantum circuits. The optimization dynamics is interpreted as moving in the steepest descent direction with respect to the Quantum Information Geometry, corresponding to the real part of…
We study the implicit bias of generic optimization methods, such as mirror descent, natural gradient descent, and steepest descent with respect to different potentials and norms, when optimizing underdetermined linear regression or separable linear classification problems. We explore the question of whether the specifi…
We revisit the choice of SGD for training deep neural networks by reconsidering the appropriate geometry in which to optimize the weights. We argue for a geometry invariant to rescaling of weights that does not affect the output of the network, and suggest Path-SGD, which is an approximate steepest descent method with …
New RL algorithms learn policies competitive with best in class without assuming optimal policy.
The affine Grassmannian is a noncompact smooth manifold that parameterizes all affine subspaces of a fixed dimension. It is a natural generalization of Euclidean space, points being zero-dimensional affine subspaces. We will realize the affine Grassmannian as a matrix manifold and extend Riemannian optimization algorit…
Gradient flow expands curves to round shapes.
AWDO trains neural networks for digit classification.
Real flag manifolds are the isotropy orbits of noncompact symmetric spaces . Any such manifold enjoys two very peculiar geometric properties: It carries a transitive action of the (noncompact) Lie group , and it is embedded in euclidean space as a taut submanifold. The aim of the paper is to link these two …
This letter presents an improved version of diffusion least mean ppower (LMP) algorithm for distributed estimation. Instead of sum of mean square errors, a weighted sum of mean square error is defined as the cost function for global and local cost functions of a network of sensors. The weight coefficients are updated b…
New method optimizes SDE models using continuous-time gradient descent.
We present here a new model and algorithm which performs an efficient Natural gradient descent for Multilayer Perceptrons. Natural gradient descent was originally proposed from a point of view of information geometry, and it performs the steepest descent updates on manifolds in a Riemannian space. In particular, we ext…
New data-driven Cartan connection tracks complex vascular structures.
Following Feynman's prescription for constructing a path integral representation of the propagator of a quantum theory, a short-time approximation to the propagator for imaginary time, N=1 supersymmetric quantum mechanics on a compact, even-dimensional Riemannian manifold is constructed. The path integral is interprete…
Paper accelerates Bayesian few-shot classification using mirror descent.
This letter proposes a dictionary learning algorithm for blind one bit compressed sensing. In the blind one bit compressed sensing framework, the original signal to be reconstructed from one bit linear random measurements is sparse in an unknown domain. In this context, the multiplication of measurement matrix $\Ab$ an…
In this paper, as a first step in examining the properties of a feasible portfolio subset that is characterized by budget and risk constraints, we assess the maximum and minimum of the investment concentration using replica analysis. To do this, we apply an analytical approach of statistical mechanics. We note that the…
In this paper we study the steepest descent -gradient flow of the functional $\SW_{λ_1,λ_2}$, which is the the sum of the Willmore energy, -weighted surface area, and -weighted enclosed volume, for surfaces immersed in . This coincides with the Helfrich functional with zero `spontaneous curvature'.…
New optimization method combines gradient clipping and non-Euclidean smoothness.
This paper analyzes Stein variational gradient descent for Bayesian inference.
We regard the real symplectic group as a constraint submanifold of the real matrices endowed with the Euclidean (Frobenius) metric, respectively as a submanifold of the general linear group endowed with the (left) invariant metric. For…
The paper explores maximal destabilizers for both K-stability and Chow-stability in unstable situations.
Many introductory courses in quantum mechanics include Feynman's time-slicing definition of the path integral, with a complete derivation of the propagator in the simplest of cases. However, attempts to generalize this, for instance to non-quadratic potentials, encounter formidable analytic issues in showing the succes…