Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

3977116154 · Jun 202019922001200920172026
48 results for descent lemma

Gradient descent converges linearly for overparameterized linear networks.

problem Convergence of gradient descent for overparameterized neural networks.
method Local Polyak-Lojasiewicz and Descent Lemma for overparameterized linear models.
result Gradient descent achieves linear convergence for two-layer linear networks under relaxed assumptions.

Paper proves multiplicative weight updates can train neural networks without learning rate tuning.

problem Vanishing and exploding gradients in gradient descent for compositional functions.
method Proves descent lemma for compositional functions using multiplicative weight updates and derives Madam optimizer.
result Madam optimizer trains state-of-the-art neural networks without learning rate tuning.

Develops a generalized version of Chung's Lemma for stochastic optimization methods.

problem Establishing asymptotic convergence rates for stochastic optimization methods under various step size rules.
method Generalized version of Chung's Lemma for a broader family of step size rules.
result Demonstrates tight non-asymptotic convergence rates for various stochastic methods.

New method improves optimization algorithms without Lipschitz smoothness.

problem Improving optimization algorithms in the absence of Lipschitz smoothness.
method Dual kernel conditioning (DKC) to provide dual Lipschitz continuity.
result First complexity bounds and iterate convergence for random reshuffling mirror descent.

Inexact subgradient methods work well for semialgebraic functions with additive errors.

problem Approximate gradients in machine learning and optimization.
method Inexact subgradient methods with persistent additive errors in semialgebraic functions.
result Iterates eventually fluctuate near the critical set with a proximity of O(ερ)O(ε^ρ), where εε is the magnitude of subgradient evaluation errors.

Paper analyzes SVGD algorithm for non-asymptotic convergence.

problem Optimizing a set of particles to approximate a target probability distribution.
method Finite time analysis of SVGD algorithm, providing descent lemma and convergence rates.
result SVGD algorithm decreases the objective at each iteration and converges to the target distribution.

Gradient descent solves robust mean estimation in high dimensions.

problem High-dimensional robust mean estimation in the presence of adversarial outliers.
method Gradient descent with a structural lemma showing near-optimal solutions.
result Gradient descent can solve the robust mean estimation problem directly.

Derandomization reveals structure in neural networks, reducing sample complexity.

problem Understanding feature learning dynamics in neural networks.
method Derandomization lemma applied to arbitrary NNs with any smooth loss function.
result Optimizing function converges to zero weight matrix, revealing structure.

This paper analyzes convergence of RMSProp and Adam in non-convex optimization with tight complexity bounds.

problem Analyzing convergence of RMSProp and Adam in non-convex optimization with relaxed assumptions.
method Developed new convergence analyses for RMSProp and Adam, considering adaptive learning rates and affine noise variance.
result RMSProp and Adam converge to ε-stationary points with iteration complexities of O(ε^(-4)) under proper hyperparameters.

Formulates Index III lemma and Rauch III theorem with applications.

problem Develops new mathematical theorems based on existing ones.
method Formulation of Index III lemma and Rauch III theorem based on Index I, II lemmas and Rauch I, II theorems.
result Presented Rauch's type theorem and volume comparison result as applications.

Tucker and Ky Fan's lemma are combinatorial analogs of the Borsuk-Ulam theorem (BUT). In 1996, Yu. A. Shashkin proved a version of Fan's lemma, which is a combinatorial analog of the odd mapping theorem (OMT). We consider generalizations of these lemmas for BUT-manifolds, i.e. for manifolds that satisfy BUT. Proofs rel…

2014-09-30abs ↗pdf ↗

Paper proves a discrete Schwarz-Pick lemma for generalized circle packings.

problem Comparing geometric quantities of circle packings with different boundary values.
method Combinatorial Calabi flows and maximum principle.
result Discrete Schwarz-Pick lemma proven for generalized circle packings.

The paper improves Zakalyukin's lemma for frontals and applies it to surface singularities.

problem Improving the conditions under which wave front germs imply map germs.
method Generalization of Zakalyukin's lemma for frontals and applications to surface singularities.
result The paper provides a more general version of Zakalyukin's lemma for map germs.

Meridian lemma extended to fully alternating links in thickened surfaces.

problem Extending Menasco's meridian lemma to fully alternating links in thickened surfaces.
method Developed a new meridian lemma for fully alternating links in thickened orientable surfaces of positive genus.
result The meridian lemma holds for fully alternating links in thickened surfaces.

The paper characterizes when the \partial \overline{\partial}-lemma holds for twistor spaces.

problem Characterizing the \partial \overline{\partial}-lemma for twistor spaces.
method Study Bott-Chern and Aeppli cohomologies of twistor spaces.
result Explicit computation of Dolbeault cohomology for flat torus twistor space.

Positive representations on surfaces have positive cross-ratios and satisfy a collar lemma.

problem Characterizing representations of surface groups with positive properties.
method Proving a collar lemma and showing positivity of cross-ratios for ΘΘ-positive representations.
result Closed subsets of representation varieties are characterized by ΘΘ-positive representations.

Enhanced Schwarz lemma for Hermitian manifolds with new curvature constraints.

problem Improving Schwarz lemma for holomorphic maps between Hermitian manifolds.
method Introducing new curvature constraints on source and target manifolds, controlling by holomorphic sectional curvature.
result Significant improvements on the Wu--Yau theorem and Schwarz lemma for Gauduchon connections.

For the convenience of readers of the article {\em No-arbitrage pricing under systemic risk: accounting for cross-ownership} (Fischer, 2012, arXiv:1005.0768), a full proof of Lemma A.5 and a shorter proof of Lemma A.6 of that paper are provided.

2012-06-21abs ↗pdf ↗

The paper extends the Discrete Schwarz-Pick Lemma to circle packings with obtuse intersections and disjoint packings.

problem Proving the Discrete Schwarz-Pick Lemma for circle packings with various inversive distances.
method Using a variational principle for circle packings with inversive distances, the paper extends the lemma to a broader range of packings.
result The Discrete Schwarz-Pick Lemma holds for circle packings with inversive distances in (1,1](-1,1], provided an additional condition on triangle weights.

The famous Švarc-Milnor Lemma says that a group GG acting properly and cocompactly via isometries on a length space XX is finitely generated and induces a quasi-isometry equivalence ggx0g\to g\cdot x_0 for any x0Xx_0\in X. We redefine the concept of coarseness so that the proof of the Lemma is automatic.

2006-03-21abs ↗pdf ↗

The L2L^2-\partial\overline\partial-Lemma is extended to complete Kähler manifolds with a gap in the spectrum.

problem Extending the L2L^2-\partial\overline\partial-Lemma to non-compact Kähler manifolds.
method Proving the L2L^2-\partial\overline\partial-Lemma on complete Kähler manifolds with a gap in the spectrum.
result The L2L^2-\partial\overline\partial-Lemma is generalized to complete Kähler manifolds.

New method to recover over-parameterized models corrupted during estimation.

problem Recovering statistical models corrupted after initial estimation.
method Robust estimation using over-parameterized models and redundancy.
result Stochastic gradient descent is well-suited for model repair, but sparsity is generally not repairable.