Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

200400599799 · Jun 202019922001200920172026
48 results for Gradient-flow optimization

A new gradient flow framework for distributionally robust optimization.

problem Optimizing under uncertainty with worst-case distributional constraints.
method Gradient flow theory applied to distributionally robust optimization.
result Practical algorithms for sampling from worst-case distributions.

New gradient flows for non-negative and probability measures combining optimal transport and interaction forces.

problem Optimizing non-negative and probability measures using interaction forces and optimal transport.
method Interaction-Force Transport (IFT) gradient flows and their spherical variant, developed via infimal convolution of Wasserstein and spherical MMD tensors, with a particle-based optimization algorithm.
result The spherical IFT gradient flow provides global exponential convergence guarantees for both MMD and KL energy.

The paper constructs optimal confidence bands for kernel gradient flow estimators.

problem Estimating generalization error and constructing confidence bands for kernel gradient flows.
method Established convergence rates and constructed optimal confidence bands under capacity-source condition.
result Optimal confidence bands for kernel gradient flows have shrinkage rates close to minimax optimal rates.

Gradient flows of neural networks converge to optimal values or diverge, with thresholds and asymptotic behaviors.

problem Understanding the convergence and divergence of gradient flows in neural networks.
method Analysis of gradient flows on loss landscapes of neural networks using o-minimal structures.
result Gradient flows either converge to optimal values or diverge to infinity, with thresholds and asymptotic behaviors.

New method optimizes multiple objectives using particle dynamics and gradient flow.

problem Optimizing multiple conflicting objectives in complex scenarios.
method Interacting particle method combining Langevin and birth-death dynamics with a dominance potential.
result Method effectively relocates dominated particles, improving Pareto optimality.

This paper bridges variational inference and Wasserstein gradient flows.

problem Combining variational inference and Wasserstein gradient flows for more efficient approximations.
method Recasting Bures-Wasserstein gradient flow as a Euclidean gradient flow and using path-derivative gradient estimator.
result A new gradient estimator for ff-divergences that can be implemented using machine learning libraries.

This work develops a particle system to approximate Fisher-Rao gradient flows in mean-field optimization.

problem Optimizing probability measures in neural network contexts.
method Constructing an interacting particle system approximating Fisher-Rao gradient flows.
result Propagation of chaos for the Fisher-Rao gradient flow in entropic mean-field optimization.

Paper explores Fisher-Rao gradient flows and their kernel approximations.

problem Understanding and analyzing approximations of Fisher-Rao gradient flows.
method Rigorous investigation of Fisher-Rao and Wasserstein type gradient flows, focusing on kernel approximations.
result Proves evolutionary Γ-convergence for kernel-approximated Fisher-Rao flows, providing theoretical guarantees.

Gradient flow method solves for optimal transport starting distributions.

problem Finding the optimal starting distribution for a martingale in optimal transport.
method Following the gradient flow of the Bass functional's L2-lift.
result Gradient flow converges to a minimizer of the Bass functional.

Sharp results link DLN gradient flow to basis pursuit optimization and GHA phase transitions.

problem Understanding implicit regularization in Diagonal Linear Networks.
method Sharp convergence bounds and characterization of 1\ell_1 minimizers.
result Gradient flow of DLNs with tiny initialization approximates minimizers of basis pursuit optimization problem.

Paper addresses optimization on Hadamard manifolds, generalizing gradient flow.

problem Optimization of convex functions on Hadamard manifolds.
method Introduces a generalized gradient flow to minimize Q(dfx)Q(df_x).
result Gradient flow attains infimum in limit for basic manifolds.

This paper analyzes convergence of large-scale Transformers with weight decay.

problem Understanding optimization guarantees in large-scale Transformer training.
method Construct mean-field limit, show gradient flow convergence to PDE, demonstrate global minimum consistency.
result Gradient flow reaches global minimum in large-scale Transformers with small weight decay.

Policy optimization is a core component of reinforcement learning (RL), and most existing RL methods directly optimize parameters of a policy based on maximizing the expected total reward, or its surrogate. Though often achieving encouraging empirical success, its underlying mathematical principle on {\em policy-distri…

2018-08-09abs ↗pdf ↗

Paper explores stability, regularization, and gradient flows for stochastic inverse problems.

problem Recovering random probability distributions from measurements.
method Direct inversion, variational formulation with regularization, and optimization via gradient flows.
result The choice of metric impacts stability and properties of the optimizer.

Study on multi-head softmax attention dynamics for in-context learning.

problem Understanding and optimizing multi-head softmax attention models for multi-task linear regression.
method Gradient flow analysis and spectral mapping technique.
result Gradient flow converges to optimal multi-head softmax attention model, with task allocation emerging during training.

This paper studies gradient flows for sampling using various metrics and their affine invariance.

problem Sampling from probability distributions with unknown normalizations.
method Gradient flows in the space of probability measures, focusing on Kullback-Leibler divergence and affine invariance of metrics.
result Gradient flows of Kullback-Leibler divergence do not depend on the normalization constant, and affine invariance is achieved for certain metrics.

Unified framework for analyzing gradient flows of measures with exponential decay of entropy.

problem Analyzing exponential decay of entropy functionals in gradient flows of measures.
method Characterization of global exponential decay behaviors using Hellinger-Kantorovich geometry, shape-mass decomposition, and Polyak-Łojasiewicz-type inequalities.
result Unified theoretical framework for gradient flows with complete analysis of exponential decay behaviors.

Gradient-flow optimization is reinterpreted as a statistical inference problem.

problem Optimizing training duration and assessing model performance in deep learning.
method Develops a statistical framework for gradient-flow training, treating it as a random-effects model.
result Establishes asymptotic optimality for prediction and reduces reliance on validation splits.

Paper establishes a generalization bound for gradient flow using a data-dependent kernel.

problem Understanding the generalization properties of gradient-based optimization methods.
method Establishes a generalization bound for gradient flow through a data-dependent kernel called the loss path kernel (LPK).
result The LPK captures the entire training trajectory and leads to tighter generalization guarantees.

Characterizes corridors in loss surfaces for gradient-based optimization.

problem Understanding and mitigating training instabilities in gradient-based optimization.
method Characterizes corridors as regions where gradient descent and gradient flow trajectories are linearly related.
result Corridors indicate regions without implicit regularization effects, leading to better learning rate adaptation schemes.

New method for scalable barycenter computation using Wasserstein gradient flows.

problem Scalability and integration of label information in barycenter computation.
method Gradient flows in Wasserstein space, time discretization, mini-batch optimal transport, modular regularization, task-aware functions, supervised information integration.
result Empirically validated new state-of-the-art barycenter solver with labeled barycenters outperforming unlabeled ones.

Gradient flows on distributions of distributions for machine learning tasks.

problem Designing gradient flows for datasets of probability distributions.
method Representing classes as conditional distributions, modeling datasets as mixture distributions, using Wasserstein over Wasserstein (WoW) distance and gradients.
result Demonstrated gradient flows for dataset transfer and distillation tasks.

We present a short overview on the strongest variational formulation for gradient flows of geodesically λλ-convex functionals in metric spaces, with applications to diffusion equations in Wasserstein spaces of probability measures. These notes are based on a series of lectures given by the second author for the Summer…

2010-09-20abs ↗pdf ↗

Optimizes convex functions in finite vs infinite dimensions, revealing slow convergence rates.

problem Analyzing gradient flows in finite and infinite-dimensional Hilbert spaces.
method Proves convergence rates and optimality conditions for gradient flows and related methods.
result Gradient flow convergence rates in finite dimensions are slower than in infinite dimensions, with optimal rates achievable in Hilbert spaces.

Paper introduces geometry-aware normalizing flows for improved causal inference.

problem Disparity between sample and population distributions in causal inference.
method Integrates continuous normalizing flows with parametric submodels, employing Wasserstein gradient flows and optimal transport.
result Significantly reduces parameter estimation bias and variance in finite-sample settings.

Onflow optimizes portfolio allocation with gradient flows, robust to transaction fees.

problem Optimizing portfolio allocation with transaction costs.
method Gradient flow reinforcement learning method for dynamic asset allocation.
result Onflow outperforms benchmarks in high transaction cost regimes.

Framework uses optimal transport for neural architecture search.

problem Optimizing neural architectures in deep learning.
method Semi-discrete optimization using optimal transport.
result Gradient flow and minimizing movement scheme converge to reaction-diffusion equations.

We show that the concept of H2H^2-gradient flow for the Willmore energy and other functionals that depend at most quadratically on the second fundamental form is well-defined in the space of immersions of Sobolev class W2,pW^{2,p} from a compact, nn-dimensional manifold into Euclidean space, provided that p2p \geq 2 and…

2017-03-19abs ↗pdf ↗

Improved sampling method using regularized Stein Variational Gradient Flow.

problem Improving the accuracy of sampling methods in machine learning.
method Proposed Regularized Stein Variational Gradient Flow to interpolate between SVGD and Wasserstein Gradient Flow.
result Established theoretical properties and provided preliminary numerical evidence of improved performance.

The Sinkhorn flow converges to a Wasserstein mirror gradient flow from the Sinkhorn algorithm.

problem Optimizing joint distributions using the Sinkhorn algorithm.
method Wasserstein mirror gradient flow derived from the Sinkhorn algorithm.
result The Sinkhorn flow converges to a Wasserstein mirror gradient flow.

Gradient flow in phase retrieval escapes spurious minima with high probability.

problem Understanding gradient-based optimization in high-dimensional non-convex functions.
method Analytical and numerical study of gradient dynamics in phase retrieval.
result Gradient flow avoids spurious minima by drifting along unstable directions.

New method optimizes multiple points in Bayesian optimization efficiently.

problem Optimizing multiple points in expensive black-box functions.
method Reformulated BO as probability measure optimization, using convex gradient flows.
result Demonstrated effectiveness on various benchmarks compared to state-of-the-art methods.

New Langevin dynamics samples from entropy-regularized optimal transport.

problem Sampling from entropy-regularized optimal transport.
method Introduced analogous diffusion dynamics constrained to Π(μ,ν)Π(μ,ν).
result Long-time limit is the unique solution of an entropic optimal transport problem.

A new gradient flow for MMD with closed-form implementation.

problem Existing gradient flows either lack tractable numerical implementation or require strong assumptions.
method Introduces a (de)-regularized Maximum Mean Discrepancy (DrMMD) and its gradient flow.
result Guarantees near-global convergence for a broad class of targets in both continuous and discrete time.

Researchers compare different gradient methods for ridge regression, finding conjugate gradients have similar performance.

problem Comparing statistical properties of different gradient methods in ridge regression.
method Explicit non-standard error decomposition to bound prediction error of conjugate gradient iterates.
result Conjugate gradient iterates share optimality properties with gradient flow and ridge regression up to a constant factor.

Softmax policy gradient achieves global optimality in wide neural networks with entropy regularization.

problem Optimizing softmax policies with neural networks in the mean-field regime.
method Modeling neural networks as Wasserstein gradient flows and proving global optimality of fixed points.
result Global optimality of softmax policy gradient in wide single hidden layer neural networks with entropy regularization.