Straight lines are a basin of attraction for the elastic flow at least to level 1.9615π.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study shows how steepest descent algorithms' geometric margin increases during training.
We develop a progressive training approach for neural networks which adaptively grows the network structure by splitting existing neurons to multiple off-springs. By leveraging a functional steepest descent idea, we derive a simple criterion for deciding the best subset of neurons to split and a splitting gradient for …
Study identifies stable configurations of intertwined threads with repulsive interactions.
This work studies the implicit bias of mini-batch SGD in classification.
Let S be a complete surface of constant curvature K = + 1 or -1, i.e. the sphere S^2 or the Lobachevskij plane L^2, and D a bounded convex subset of S. If S = S^2, assume also diameter (D) < pi/2. It is proved that the length of any steepest descent curve of a quasi-convex function in D is less than or equal to the per…
New method escapes local optima in neural architecture optimization.
This letter proposes a sparse diffusion steepest-descent algorithm for one bit compressed sensing in wireless sensor networks. The approach exploits the diffusion strategy from distributed learning in the one bit compressed sensing framework. To estimate a common sparse vector cooperatively from only the sign of measur…
Gradient flow expands curves to round shapes.
Study shows momentum-based optimizers like Muon and MomentumGD bias towards KKT points in smooth homogeneous models.
New method accelerates steepest descent for convex optimization.
Improved greedy 2-coordinate updates for optimization problems with constraints.
Characterizes corridors in loss surfaces for gradient-based optimization.
Two popular examples of first-order optimization methods over linear spaces are coordinate descent and matching pursuit algorithms, with their randomized variants. While the former targets the optimization by moving along coordinates, the latter considers a generalized notion of directions. Exploiting the connection be…
GANs excel at learning high dimensional distributions, but they can update generator parameters in directions that do not correspond to the steepest descent direction of the objective. Prominent examples of problematic update directions include those used in both Goodfellow's original GAN and the WGAN-GP. To formally d…
Method approximates Riemannian barycenter on manifolds.
Bayesian inference problems require sampling or approximating high-dimensional probability distributions. The focus of this paper is on the recently introduced Stein variational gradient descent methodology, a class of algorithms that rely on iterated steepest descent steps with respect to a reproducing kernel Hilbert …
The paper optimizes regret using covariance between costs and decisions.
In this paper we study the steepest descent -gradient flow of the functional $\SW_{λ_1,λ_2}$, which is the the sum of the Willmore energy, -weighted surface area, and -weighted enclosed volume, for surfaces immersed in . This coincides with the Helfrich functional with zero `spontaneous curvature'.…
Real flag manifolds are the isotropy orbits of noncompact symmetric spaces . Any such manifold enjoys two very peculiar geometric properties: It carries a transitive action of the (noncompact) Lie group , and it is embedded in euclidean space as a taut submanifold. The aim of the paper is to link these two …
Unified signSGD and gradient descent analysis for neural networks.
In this paper we will discuss how one may be able to use mean curvature flow to tackle some of the central problems in topology in 4-dimensions. We will be concerned with smooth closed 4-manifolds that can be smoothly embedded as a hypersurface in R^5. We begin with explaining why all closed smooth homotopy spheres can…
In optimization, the negative gradient of a function denotes the direction of steepest descent. Furthermore, traveling in any direction orthogonal to the gradient maintains the value of the function. In this work, we show that these orthogonal directions that are ignored by gradient descent can be critical in equilibri…
Adaptive step-size improves optimization in complex geometries.
Logistic regression is one of the most popular methods in binary classification, wherein estimation of model parameters is carried out by solving the maximum likelihood (ML) optimization problem, and the ML estimator is defined to be the optimal solution of this problem. It is well known that the ML estimator exists wh…
Study uses outer metrics for PDE-constrained shape optimization over diffeomorphism group.
We study the implicit bias of generic optimization methods, such as mirror descent, natural gradient descent, and steepest descent with respect to different potentials and norms, when optimizing underdetermined linear regression or separable linear classification problems. We explore the question of whether the specifi…
SONIA optimizes machine learning problems with a novel algorithm.
In this paper we consider the steepest descent -gradient flow of the length functional for immersed plane curves, known as the curve diffusion flow. It is known that under this flow there exist both initially immersed curves which develop at least one singularity in finite time and initially embedded curves whi…
A quantum generalization of Natural Gradient Descent is presented as part of a general-purpose optimization framework for variational quantum circuits. The optimization dynamics is interpreted as moving in the steepest descent direction with respect to the Quantum Information Geometry, corresponding to the real part of…
Analyzed a generative model framework through Wasserstein Gradient Flow.
New method optimizes SDE models using continuous-time gradient descent.
We present here a new model and algorithm which performs an efficient Natural gradient descent for Multilayer Perceptrons. Natural gradient descent was originally proposed from a point of view of information geometry, and it performs the steepest descent updates on manifolds in a Riemannian space. In particular, we ext…
New data-driven Cartan connection tracks complex vascular structures.
Paper accelerates Bayesian few-shot classification using mirror descent.
Scalable and robust TR decomposition for large-scale data with missing entries and outliers.
New optimizers control network width scaling, improving stability and transfer across different model sizes.
This paper presents a stochastic behavior analysis of a kernel-based stochastic restricted-gradient descent method. The restricted gradient gives a steepest ascent direction within the so-called dictionary subspace. The analysis provides the transient and steady state performance in the mean squared error criterion. It…
A new algorithm for training generative models using Sinkhorn divergence.
A new algorithm for minimizing functions on Wasserstein space.
New algorithm optimizes nonlinear SDEs online with convergence guarantees.
New RL algorithms learn policies competitive with best in class without assuming optimal policy.
A new method improves Bayesian filtering in nonlinear systems.
AdamW optimizes a constrained loss with norm constraint.
The paper studies ideal flows of closed curves, classifying critical points and proving flow behavior.
We revisit the choice of SGD for training deep neural networks by reconsidering the appropriate geometry in which to optimize the weights. We argue for a geometry invariant to rescaling of weights that does not affect the output of the network, and suggest Path-SGD, which is an approximate steepest descent method with …
This paper presents the application of a newly developed nature-inspired metaheuristic optimization method, namely the Adaptive Wind Driven Optimization (AWDO), to the training of feedforward artificial neural networks (NN) and presents a discussion into the future research of AWDO implementation in Deep Learning (DL).…
The affine Grassmannian is a noncompact smooth manifold that parameterizes all affine subspaces of a fixed dimension. It is a natural generalization of Euclidean space, points being zero-dimensional affine subspaces. We will realize the affine Grassmannian as a matrix manifold and extend Riemannian optimization algorit…