Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

93185278370 · Jun 202019922001200920172026
48 results for rescaled gradients

Rescaled ASGD optimizes distributed learning under heterogeneous data.

problem Vanilla ASGD biases towards a frequency-weighted average of local objectives.
method Rescale worker stepsizes by their computation times.
result Rescaled ASGD converges to the correct global objective in fixed-computation model.

Deterministic GD can behave stochastically in large learning rates for multiscale functions.

problem Understanding deterministic GD's stochastic behavior in large learning rates for multiscale objectives.
method Established a sufficient condition for deterministic GD to converge to a rescaled Gibbs distribution in large learning rates for multiscale functions.
result Deterministic GD can converge to a statistical distribution in large learning rates for multiscale functions.

Riemannian gradient descent escapes some spurious critical points on low-rank matrix manifold.

problem Spurious critical points on the boundary of low-rank matrix manifold.
method Riemannian gradient descent with dynamical low-rank approximation and rescaled gradient flow.
result Riemannian gradient descent escapes some spurious critical points on the boundary of the manifold.

We give two new proofs of Perelman's theorem that shrinking breathers of Ricci flow on closed manifolds are gradient Ricci solitons, using the fact that the singularity models of type I solutions are shrinking gradient Ricci solitons and the fact that non-collapsed type I ancient solutions have rescaled limits being sh…

2017-05-23abs ↗pdf ↗

"Ends of hyperbolic 3-manifolds should support canonical Wick Rotations, so they realize effective interactions of their ending globally hyperbolic spacetimes of constant curvature." We develop a consistent sector of WR-rescaling theory in 3D gravity, that, in particular, concretizes the above guess for many geometrica…

2004-12-23abs ↗pdf ↗

Adaptive stepsizing improves sampling in Bayesian neural networks.

problem Scalable sampling of posterior distributions in Bayesian neural networks.
method SA-SGLD, employing time rescaling to adapt stepsize dynamically.
result SA-SGLD achieves more accurate posterior sampling than SGLD.

Huisken studied asymptotic behavior of a mean curvature flow in a Euclidean space when it develops a singularity of type I, and proved that its rescaled flow converges to a self-shrinker in the Euclidean space. In this paper, we generalize this result for a Ricci-mean curvature flow moving along a Ricci flow constructe…

2015-01-26abs ↗pdf ↗

We prove gradient estimates for hypersurfaces in the hyperbolic space Hn+1,\mathbb{H}^{n+1}, expanding by negative powers of a certain class of homogeneous curvature functions. We obtain optimal gradient estimates for hypersurfaces evolving by certain powers p>1p>1 of F1F^{-1} and smooth convergence of the properly rescale…

2014-10-06abs ↗pdf ↗

PolyNSD improves Neural Sheaf Diffusion with polynomial operators and spectral rescaling.

problem Limitations of common Neural Sheaf Diffusion implementations, including scalability and stability issues.
method Introduces Polynomial Neural Sheaf Diffusion (PolyNSD) with a degree-K polynomial propagation operator and spectral rescaling.
result PolyNSD achieves state-of-the-art results on both homophilic and heterophilic benchmarks with reduced runtime and memory requirements.

The paper studies the free elastic flow of closed curves and finds their asymptotic shape converges to a circle.

problem Challenges in studying the asymptotic behavior of the free elastic flow for closed curves.
method Analysis of the free elastic flow as an L2L^2-gradient flow for Euler's elastic energy.
result An appropriate rescaling of initial curves geometrically close to circles converges to a unique round circle.

Learning with non-modular losses is an important problem when sets of predictions are made simultaneously. The main tools for constructing convex surrogate loss functions for set prediction are margin rescaling and slack rescaling. In this work, we show that these strategies lead to tight convex surrogates iff the unde…

2015-12-24abs ↗pdf ↗

Algorithm optimizes functions without parameters, converging to global minima.

problem Optimizing functions without parameters.
method Follow The Regularized Leader with rescaled gradients and time-varying regularizers.
result Converges to global minimizer for variationally coherent functions.

The study examines 4D steady gradient Ricci solitons with nonnegative curvature away from a compact set.

problem Analyzing noncompact steady gradient Ricci solitons with nonnegative curvature operator.
method Examining the asymptotic behavior of noncompact κ-noncollapsed steady gradient Ricci solitons with nonnegative curvature operator away from a compact set.
result 4D noncompact κ-noncollapsed steady gradient Ricci solitons with nonnegative sectional curvature must be a Bryant Ricci soliton up to scaling.

The paper calculates spectral torsion for rescaled Dirac operators on manifolds.

problem Computing spectral torsion for rescaled Dirac operators.
method Using trilinear Clifford multiplication and functional of differential one-forms.
result Computed spectral torsion for four types of rescaled Dirac operators.

For a Riemannian manifold MM, we determine some curvature properties of a tangent bundle equipped with the rescaled metric.The main aim of this paper is to give explicit formulae for the rescaled metric on TMTM, and investigate the geodesics on the tangent bundle with respect to the rescaled Sasaki metric.

2011-04-29abs ↗pdf ↗

Study steady gradient Ricci solitons with cylindrical tangent flows at infinity.

problem Characterize the geometry of steady gradient Ricci solitons at infinity.
method Analyze the rescaled limits of finite-time singular solutions of the Ricci flow.
result Classify the tangent flows at infinity of 4-dimensional steady soliton singularity models.

Study geometric characterization of asymptotic pseudodifferential calculus on spinor bundles.

problem Geometric characterization of asymptotic pseudodifferential calculus on spinor bundles.
method Groupoid approach to pseudodifferential calculus, rescaled bundle.
result Rescaled bundle provides geometric characterization to asymptotic pseudodifferential calculus on spinor bundles.

The paper calculates the noncommutative residue for a rescaled Dirac operator on 6D manifolds.

problem Computing the noncommutative residue for a specific Dirac operator on 6D manifolds.
method Calculations and proofs for the rescaled Dirac operator fDh on 6D compact manifolds.
result Proof of the Kastler-Kalau-Walze type theorem for the rescaled Dirac operator on 6D compact manifolds with boundary.

We propose a stepsize adaptation scheme for stochastic gradient descent. It operates directly with the loss function and rescales the gradient in order to make fixed predicted progress on the loss. We demonstrate its capabilities by conclusively improving the performance of Adam and Momentum optimizers. The enhanced op…

2018-02-14abs ↗pdf ↗

Infinitesimal gradient boosting is a new algorithm derived from gradient boosting.

problem Improving the efficiency and smoothness of gradient boosting.
method Introduced a new class of randomized regression trees and used a limit process in vanishing-learning-rate asymptotic.
result Convergence of the stochastic algorithm and characterization of the limiting procedure as a unique solution of a nonlinear ODE.

Stochastic gradient descent's long-term fluctuations are described by a diffusion limit.

problem Long-term behavior of stochastic gradient descent in non-smooth settings.
method Functional central limit theorem applied to rescaled trajectory of SGD.
result Characterization of long-term fluctuations around the minimizer.

Fast algorithm for rescaling vectors with clipping, improving training efficiency.

problem Efficiently rescale vectors to a desired length while maintaining them within a domain after clipping.
method Analytical solution for optimal rescaling using fast and differentiable algorithm.
result Optimal rescaling can be found analytically, improving training efficiency for neural networks.

A new method to rescale ReLU neural networks based on path-lifting.

problem Lack of principled ways to leverage rescaling symmetries in ReLU neural networks.
method Introduces a geometrically motivated criterion to rescale neural network parameters, aligning a kernel in the path-lifting space with a chosen reference.
result Proposed method can speed up training and aligns a kernel in the path-lifting space with a chosen reference.

New theory shows predictive coding makes learning landscape easier to navigate.

problem Understanding the impact of predictive coding's inference procedure on learning efficiency.
method Analyzed the geometry of the energy landscape of deep linear networks, proving many non-strict saddles become strict in the equilibrated energy.
result All highly degenerate (non-strict) saddles of the loss become strict in the equilibrated energy, suggesting a more robust learning landscape.

This paper tackles non-vacuous generalization bounds in ReLU networks by resolving rescaling invariances.

problem Non-vacuous generalization guarantees for ReLU networks with rescaling invariances.
method Proposes a lifted representation to resolve rescaling invariances and studies KL-based rescaling-invariant PAC-Bayes bounds.
result KL-based rescaling-invariant PAC-Bayes bounds provide tighter guarantees and resolve discrepancies in network complexity.

A new debiasing method for high-dimensional regression with applications to PCR.

problem Debiasing in high-dimensional statistics with i.i.d. samples and sub-Gaussian covariates.
method Spectrum-Aware Debiasing using rescaled gradient descent with spectral information.
result Achieves debiasing in broader contexts with structured dependencies, heavy tails, and low-rank structures.

We define a second-order neural network stochastic gradient training algorithm whose block-diagonal structure effectively amounts to normalizing the unit activations. Investigating why this algorithm lacks in robustness then reveals two interesting insights. The first insight suggests a new way to scale the stepsizes, …

2017-05-25abs ↗pdf ↗

A new approach simplifies Sliced-Wasserstein distances to improve learning performance.

problem The concentration of measure phenomenon makes random projections uninformative in high dimensions.
method Propose rescaling the 1D Wasserstein distance to make all slices equally informative.
result The classical Sliced-Wasserstein, properly configured, can match or surpass complex variants.

New Lipschitz bound for ReLU networks resists weight rescaling.

problem Lack of robustness guarantees for ReLU networks under weight perturbations.
method Rescaling-invariant Lipschitz bound based on path-metrics.
result The new bound applies to various ReLU-DAG architectures and resists neuron-wise rescalings.

We exploit the spinor description of four-dimensional Walker geometry, and conformal rescalings of such, to describe the local geometry of four-dimensional neutral geometries with algebraically degenerate self-dual Weyl curvature and an integrable distribution of alpha-planes (algebraically special real alpha-geometry)…

2008-08-15abs ↗pdf ↗

The H1(ds)H^1(ds)-gradient flow shrinks circles with radius r0r_0 to a point.

problem The triviality of the L2(ds)L^2(ds) metric topology on immersed planar curves.
method Gradient flow of the length functional with respect to the H1(ds)H^1(ds)-metric.
result Circles shrink to a point under the H1(ds)H^1(ds)-gradient flow.

Symmetry in loss functions constrains model parameters, leading to specific learning outcomes.

problem Understanding and leveraging symmetries in neural networks to improve learning outcomes.
method Analyzing the impact of loss function symmetries on model parameters and learning behavior.
result Mirror-reflection symmetries in loss functions lead to constraints on model parameters, influencing learning outcomes.

In this paper we study the Teichmüller harmonic map flow as introduced by Rupflin and Topping [15]. It evolves pairs of maps and metrics (u,g)(u,g) into branched minimal immersions, or equivalently into weakly conformal harmonic maps, where uu maps from a fixed closed surface MM with metric gg to a general target manif…

2017-11-24abs ↗pdf ↗