Sigmoid-type networks avoid vanishing gradients with regularization and rescaling.
problem Vanishing gradients in sigmoid-type networks.
method Mathematical arguments and two remedies: regularization and rescaling.
result Demonstrates effectiveness of regularization and rescaling in practice.
Rescaled ASGD optimizes distributed learning under heterogeneous data.
problem Vanilla ASGD biases towards a frequency-weighted average of local objectives.
method Rescale worker stepsizes by their computation times.
result Rescaled ASGD converges to the correct global objective in fixed-computation model.
A new method to improve deep neural networks using weight rescaling.
problem Overfitting and sensitivity to hyperparameters in weight decay.
method Weight rescaling (WRS) to control weight norm and prevent overfitting.
result WRS outperforms weight decay and other methods in various applications.
Deterministic GD can behave stochastically in large learning rates for multiscale functions.
problem Understanding deterministic GD's stochastic behavior in large learning rates for multiscale objectives.
method Established a sufficient condition for deterministic GD to converge to a rescaled Gibbs distribution in large learning rates for multiscale functions.
result Deterministic GD can converge to a statistical distribution in large learning rates for multiscale functions.
Riemannian gradient descent escapes some spurious critical points on low-rank matrix manifold.
problem Spurious critical points on the boundary of low-rank matrix manifold.
method Riemannian gradient descent with dynamical low-rank approximation and rescaled gradient flow.
result Riemannian gradient descent escapes some spurious critical points on the boundary of the manifold.
We give two new proofs of Perelman's theorem that shrinking breathers of Ricci flow on closed manifolds are gradient Ricci solitons, using the fact that the singularity models of type I solutions are shrinking gradient Ricci solitons and the fact that non-collapsed type I ancient solutions have rescaled limits being sh…
"Ends of hyperbolic 3-manifolds should support canonical Wick Rotations, so they realize effective interactions of their ending globally hyperbolic spacetimes of constant curvature." We develop a consistent sector of WR-rescaling theory in 3D gravity, that, in particular, concretizes the above guess for many geometrica…
Adaptive stepsizing improves sampling in Bayesian neural networks.
problem Scalable sampling of posterior distributions in Bayesian neural networks.
method SA-SGLD, employing time rescaling to adapt stepsize dynamically.
result SA-SGLD achieves more accurate posterior sampling than SGLD.
We show that a rescale limit at any degenerate singularity of Ricci flow in dimension 3 is a steady gradient soliton. In particular, we give a geometric description of type I and type II singularities.
Huisken studied asymptotic behavior of a mean curvature flow in a Euclidean space when it develops a singularity of type I, and proved that its rescaled flow converges to a self-shrinker in the Euclidean space. In this paper, we generalize this result for a Ricci-mean curvature flow moving along a Ricci flow constructe…
We prove gradient estimates for hypersurfaces in the hyperbolic space Hn+1, expanding by negative powers of a certain class of homogeneous curvature functions. We obtain optimal gradient estimates for hypersurfaces evolving by certain powers p>1 of F−1 and smooth convergence of the properly rescale…
PolyNSD improves Neural Sheaf Diffusion with polynomial operators and spectral rescaling.
problem Limitations of common Neural Sheaf Diffusion implementations, including scalability and stability issues.
method Introduces Polynomial Neural Sheaf Diffusion (PolyNSD) with a degree-K polynomial propagation operator and spectral rescaling.
result PolyNSD achieves state-of-the-art results on both homophilic and heterophilic benchmarks with reduced runtime and memory requirements.
The paper studies the free elastic flow of closed curves and finds their asymptotic shape converges to a circle.
problem Challenges in studying the asymptotic behavior of the free elastic flow for closed curves.
method Analysis of the free elastic flow as an L2-gradient flow for Euler's elastic energy. result An appropriate rescaling of initial curves geometrically close to circles converges to a unique round circle.
Learning with non-modular losses is an important problem when sets of predictions are made simultaneously. The main tools for constructing convex surrogate loss functions for set prediction are margin rescaling and slack rescaling. In this work, we show that these strategies lead to tight convex surrogates iff the unde…
Algorithm optimizes functions without parameters, converging to global minima.
problem Optimizing functions without parameters.
method Follow The Regularized Leader with rescaled gradients and time-varying regularizers.
result Converges to global minimizer for variationally coherent functions.
The study examines 4D steady gradient Ricci solitons with nonnegative curvature away from a compact set.
problem Analyzing noncompact steady gradient Ricci solitons with nonnegative curvature operator.
method Examining the asymptotic behavior of noncompact κ-noncollapsed steady gradient Ricci solitons with nonnegative curvature operator away from a compact set.
result 4D noncompact κ-noncollapsed steady gradient Ricci solitons with nonnegative sectional curvature must be a Bryant Ricci soliton up to scaling.
In this short notes, we discuss monotonicity formulas under various rescaled versions of Ricci flow. The main result is Theorem \ref{theo rescaled}.
The paper calculates spectral torsion for rescaled Dirac operators on manifolds.
problem Computing spectral torsion for rescaled Dirac operators.
method Using trilinear Clifford multiplication and functional of differential one-forms.
result Computed spectral torsion for four types of rescaled Dirac operators.
New spectral torsion defined for rescaled Dirac operators.
problem Defining spectral torsion for rescaled Dirac operators.
method Using three vector fields and noncommutative residue.
result Computed spectral torsion for one form rescaled Dirac operators.
For a Riemannian manifold M, we determine some curvature properties of a tangent bundle equipped with the rescaled metric.The main aim of this paper is to give explicit formulae for the rescaled metric on TM, and investigate the geodesics on the tangent bundle with respect to the rescaled Sasaki metric.
Study steady gradient Ricci solitons with cylindrical tangent flows at infinity.
problem Characterize the geometry of steady gradient Ricci solitons at infinity.
method Analyze the rescaled limits of finite-time singular solutions of the Ricci flow.
result Classify the tangent flows at infinity of 4-dimensional steady soliton singularity models.
Study geometric characterization of asymptotic pseudodifferential calculus on spinor bundles.
problem Geometric characterization of asymptotic pseudodifferential calculus on spinor bundles.
method Groupoid approach to pseudodifferential calculus, rescaled bundle.
result Rescaled bundle provides geometric characterization to asymptotic pseudodifferential calculus on spinor bundles.
The paper calculates the noncommutative residue for a rescaled Dirac operator on 6D manifolds.
problem Computing the noncommutative residue for a specific Dirac operator on 6D manifolds.
method Calculations and proofs for the rescaled Dirac operator fDh on 6D compact manifolds.
result Proof of the Kastler-Kalau-Walze type theorem for the rescaled Dirac operator on 6D compact manifolds with boundary.
Rescaling expansiveness proven for k*-expansive vector fields.
problem Proving rescaling expansiveness for k*-expansive vector fields.
method Introducing and exploring singular-expansive flows.
result Rescaling expansiveness established for k*-expansive vector fields.
In this thesis, we consider the knot energy "integral Menger curvature" which is the triple integral over the inverse of the classic circumradius of three distinct points on the given knot to the power p∈[2,∞). We prove the existence of the first variation for a subset of a certain fractional Sobolev space if…
We propose a stepsize adaptation scheme for stochastic gradient descent. It operates directly with the loss function and rescales the gradient in order to make fixed predicted progress on the loss. We demonstrate its capabilities by conclusively improving the performance of Adam and Momentum optimizers. The enhanced op…
Infinitesimal gradient boosting is a new algorithm derived from gradient boosting.
problem Improving the efficiency and smoothness of gradient boosting.
method Introduced a new class of randomized regression trees and used a limit process in vanishing-learning-rate asymptotic.
result Convergence of the stochastic algorithm and characterization of the limiting procedure as a unique solution of a nonlinear ODE.
Stochastic gradient descent's long-term fluctuations are described by a diffusion limit.
problem Long-term behavior of stochastic gradient descent in non-smooth settings.
method Functional central limit theorem applied to rescaled trajectory of SGD.
result Characterization of long-term fluctuations around the minimizer.
Fast algorithm for rescaling vectors with clipping, improving training efficiency.
problem Efficiently rescale vectors to a desired length while maintaining them within a domain after clipping.
method Analytical solution for optimal rescaling using fast and differentiable algorithm.
result Optimal rescaling can be found analytically, improving training efficiency for neural networks.
A new method to rescale ReLU neural networks based on path-lifting.
problem Lack of principled ways to leverage rescaling symmetries in ReLU neural networks.
method Introduces a geometrically motivated criterion to rescale neural network parameters, aligning a kernel in the path-lifting space with a chosen reference.
result Proposed method can speed up training and aligns a kernel in the path-lifting space with a chosen reference.
Functional central limit theorem for kernel gradient flow and infinitesimal gradient boosting
problem Fluctuations of boosting processes around their deterministic limit
method Stochastic perturbation analysis of ODEs in Banach spaces
result Rescaled deviations converge to a Gaussian process
New theory shows predictive coding makes learning landscape easier to navigate.
problem Understanding the impact of predictive coding's inference procedure on learning efficiency.
method Analyzed the geometry of the energy landscape of deep linear networks, proving many non-strict saddles become strict in the equilibrated energy.
result All highly degenerate (non-strict) saddles of the loss become strict in the equilibrated energy, suggesting a more robust learning landscape.
This paper tackles non-vacuous generalization bounds in ReLU networks by resolving rescaling invariances.
problem Non-vacuous generalization guarantees for ReLU networks with rescaling invariances.
method Proposes a lifted representation to resolve rescaling invariances and studies KL-based rescaling-invariant PAC-Bayes bounds.
result KL-based rescaling-invariant PAC-Bayes bounds provide tighter guarantees and resolve discrepancies in network complexity.
Localizes Wodzicki residue for logarithm of differential operators.
problem Localizing Wodzicki residue for logarithm of differential operators.
method Localisation formula using rescaled differential operators and spinor bundles.
result Expresses index of Dirac operator in terms of local density involving logarithm.
A new debiasing method for high-dimensional regression with applications to PCR.
problem Debiasing in high-dimensional statistics with i.i.d. samples and sub-Gaussian covariates.
method Spectrum-Aware Debiasing using rescaled gradient descent with spectral information.
result Achieves debiasing in broader contexts with structured dependencies, heavy tails, and low-rank structures.
We define a second-order neural network stochastic gradient training algorithm whose block-diagonal structure effectively amounts to normalizing the unit activations. Investigating why this algorithm lacks in robustness then reveals two interesting insights. The first insight suggests a new way to scale the stepsizes, …
We prove a local index theorem of Atiyah-Singer type for Dirac operators on manifolds with a Lie structure at infinity (Lie manifolds for short). With the help of a renormalized supertrace, defined on a suitable class of regularizing operators, the proof of the index theorem relies on a rescaling technique similar in s…
A new approach simplifies Sliced-Wasserstein distances to improve learning performance.
problem The concentration of measure phenomenon makes random projections uninformative in high dimensions.
method Propose rescaling the 1D Wasserstein distance to make all slices equally informative.
result The classical Sliced-Wasserstein, properly configured, can match or surpass complex variants.
New Lipschitz bound for ReLU networks resists weight rescaling.
problem Lack of robustness guarantees for ReLU networks under weight perturbations.
method Rescaling-invariant Lipschitz bound based on path-metrics.
result The new bound applies to various ReLU-DAG architectures and resists neuron-wise rescalings.
We exploit the spinor description of four-dimensional Walker geometry, and conformal rescalings of such, to describe the local geometry of four-dimensional neutral geometries with algebraically degenerate self-dual Weyl curvature and an integrable distribution of alpha-planes (algebraically special real alpha-geometry)…
The H1(ds)-gradient flow shrinks circles with radius r0 to a point.
problem The triviality of the L2(ds) metric topology on immersed planar curves. method Gradient flow of the length functional with respect to the H1(ds)-metric. result Circles shrink to a point under the H1(ds)-gradient flow. The paper constructs bundles and recovers Kirillov character formula.
problem Constructing smooth vector bundles over deformation to the normal cone.
method Rescaling of vector bundles and equivariant constructions.
result Recovery of Kirillov character formula for equivariant index.
Study of Bach flow on specific nilmanifolds, converging to a soliton.
problem Analyzing the Bach flow on specific nilmanifolds.
method Fourth order geometric flow on four-dimensional simply connected nilmanifolds.
result The Bach flow converges to an expanding Bach soliton on these manifolds.
Symmetry in loss functions constrains model parameters, leading to specific learning outcomes.
problem Understanding and leveraging symmetries in neural networks to improve learning outcomes.
method Analyzing the impact of loss function symmetries on model parameters and learning behavior.
result Mirror-reflection symmetries in loss functions lead to constraints on model parameters, influencing learning outcomes.
When optimizing over-parameterized models, such as deep neural networks, a large set of parameters can achieve zero training error. In such cases, the choice of the optimization algorithm and its respective hyper-parameters introduces biases that will lead to convergence to specific minimizers of the objective. Consequ…
We consider in this work a system of two stochastic differential equations named the perturbed compositional gradient flow. By introducing a separation of fast and slow scales of the two equations, we show that the limit of the slow motion is given by an averaged ordinary differential equation. We then demonstrate that…
Study proves existence and uniqueness of ancient flows from cones.
problem Existence and uniqueness of ancient rescaled mean curvature flows.
method Proved existence and uniqueness using strong uniqueness theorem.
result Proved existence and uniqueness of ancient flows from cones.
In this paper we study the Teichmüller harmonic map flow as introduced by Rupflin and Topping [15]. It evolves pairs of maps and metrics (u,g) into branched minimal immersions, or equivalently into weakly conformal harmonic maps, where u maps from a fixed closed surface M with metric g to a general target manif…