SAM improves generalization by operating near the edge of stability.
problem Improving generalization in neural networks.
method Sharpness-Aware Minimization (SAM) approach to training neural networks.
result SAM operates near the 'edge of stability' identified by the analysis.
SGD works well with large learning rates at the edge of stability.
problem Stochasticity at the edge of stability in deep learning.
method Sharp convergence guarantees for SGD with multiclass cross-entropy loss.
result SGD self-stabilizes, ensuring convergence with large learning rates.
Gradient descent on neural nets often operates at the Edge of Stability, where loss behavior is complex but loss decreases over time.
problem Understanding the optimization dynamics of neural networks at the Edge of Stability.
method Empirical demonstration of gradient descent behavior in neural network training.
result Gradient descent on neural networks typically occurs at the Edge of Stability, where loss behavior is non-monotonic but loss decreases over time.
Gradient descent at edge of stability stabilizes implicitly, following projected gradient descent.
problem Gradient descent's stability and sharpness behavior at the edge of instability.
method Cubic Taylor expansion analysis of gradient descent dynamics.
result Gradient descent at edge of stability implicitly follows projected gradient descent.
Rod flow models Adam's behavior at the edge of stability.
problem Modeling adaptive gradient methods like Adam at the edge of stability.
method Extended rod flow to Adam, considering parameters, first moment, and second moment as variables.
result Rod flow accurately tracks Adam's behavior through the edge-of-stability regime.
Kernel networks' stability edge linked to Fisher Information singularity.
problem Understanding the stability edge in high-capacity kernel Hopfield networks.
method Statistical manifold analysis and Riemannian geometry.
result The Ridge of Optimization corresponds to the Edge of Stability, revealing a dual equilibrium.
New findings show mini-batch SGD operates in a 'Edge of Stochastic Stability' regime.
problem Understanding the stability and convergence of mini-batch SGD.
method Analyzing the mini-batch Hessian and its directional curvature.
result Mini-batch SGD operates in a different stability regime (Edge of Stochastic Stability) compared to full-batch GD.
Training avoids edge of stability by aligning Jacobian matrices.
problem Training neural networks on the edge of stability causes inaccuracies.
method Used an exponential Euler solver to prevent entering the edge of stability.
result Alignment of Jacobian matrices causes sharpness increase in Hessian.
Study on planar graph braid groups' second homology.
problem Characterize the second homology of planar graph braid groups.
method Analyzing configuration spaces of planar graphs under specific operations.
result The second homology is generated by three specific graphs.
The study analyzes sharpness dynamics in neural networks, revealing mechanisms and conditions.
problem Understanding sharpness in neural network training.
method Fixed point analysis and edge of stability analysis in a simplified 2-layer linear network.
result Reveals mechanisms behind sharpness trends, conditions for edge of stability, and a period-doubling route to chaos.
Gradient descent forces neural network eigenvalues to a specific threshold.
problem Understanding why gradient descent drives eigenvalues to a specific threshold.
method Introduced edge coupling, a functional on consecutive iterate pairs, to explain the trajectory towards the eigenvalue threshold.
result Gradient descent forces the Hessian eigenvalue to the threshold 2/η from arbitrary initialization. We introduce a novel type of stabilization map on the configuration spaces of a graph, which increases the number of particles occupying an edge. There is an induced action on homology by the polynomial ring generated by the set of edges, and we show that this homology module is finitely generated. An analogue of class…
Momentum affects optimization differently at small vs large batch sizes near instability.
problem Understanding how momentum impacts optimization near the edge of stability.
method Demonstrated through batch-size dependent behavior of SGD with momentum.
result Momentum operates in two distinct regimes: amplifying stochastic fluctuations at small batch sizes and stabilizing at large batch sizes.
A new ODE model explains gradient descent dynamics near edge of stability.
problem Understanding gradient-based training over non-convex landscapes.
method Rod Flow, a new ODE approximation of GD dynamics.
result Rod Flow accurately predicts critical sharpness threshold and self-stabilization in quartic potentials.
This work analyzes the stability of graph filters under large perturbations.
problem Stability of graph filters under large edge rewires.
method Proves a bound on stability using frequency response and community structure.
result Graph filter stability depends on perturbation to community structure.
Gradient descent near stability threshold exhibits sharpness oscillations.
problem Understanding sharpness behavior near stability threshold in non-Euclidean norms.
method Interpreted EoS through Directional Smoothness and generalized sharpness under arbitrary norms.
result Non-Euclidean GD with generalized sharpness shows sharpness oscillations near 2/η. EoS selectively shapes learning, affecting some groups more than others.
problem EoS affects learning differently across the data distribution.
method Branching intervention to enter or exit EoS regime, controlled perturbation to isolate mechanisms.
result EoS redistributes learning, amplifying progress on some groups and suppressing others.
Deep linear networks oscillate beyond the edge of stability in a predictable manner.
problem Understanding oscillations in deep linear networks beyond the edge of stability.
method Theoretical analysis of loss oscillations in deep matrix factorization loss.
result Loss oscillations in deep linear networks follow a period-doubling route to chaos and occur within a small subspace.
VL finds flatter solutions at edge of stability, matching theory with practice.
problem Understanding implicit regularization in deep learning.
method Edge of Stability framework, controlling variational posterior shape and sample number.
result VL finds even flatter solutions than gradient descent.
Sparse connectivity improves generalization in neural networks below the Edge of Stability.
problem Generalization guarantees for fully-connected networks fail at the Edge of Stability.
method Analyzed sparse connectivity's impact on generalization in two-layer ReLU networks.
result Sparse connectivity changes the effective constraint, leading to non-vacuous generalization bounds.
Study on GD and SGD over diagonal networks, focusing on stepsizes and regularisation.
problem Understanding the impact of stochasticity and large stepsizes on gradient descent and SGD solutions.
method Investigation of GD and SGD over diagonal linear networks with macroscopic stepsizes, proving convergence and characterizing solutions.
result Large stepsizes consistently benefit SGD for sparse regression problems, but can hinder GD recovery of sparse solutions, especially in the edge of stability regime.
A simple function shows how neural nets can converge despite high sharpness.
problem Understanding why neural nets converge with high sharpness.
method Constructed a minimal example function and analyzed its training dynamics rigorously.
result Final converging point has sharpness close to 2/η. Weight decay stabilizes training dynamics by slowing progressive sharpening.
problem Understanding how weight decay affects training stability in deep learning models.
method Analyzing weight decay effects at the Edge of Stability, developing a mathematical framework.
result Weight decay dampens oscillations and stabilizes sharpness in CNNs, causing a phase transition in MLPs.
New stability conditions for ZO methods reveal unique regularization effects.
problem Understanding optimization dynamics of ZO methods in deep learning.
method Explicit step size conditions and stability bounds derived for ZO methods.
result ZO methods operate near the edge of stability, with regularization effects specific to Hessian trace vs. eigenvalue.
Cohen et al. (2021) show GD trajectories align on a bifurcation diagram.
problem Understanding the Edge of Stability (EoS) phenomenon in gradient descent.
method Empirical studies and rigorous mathematical proofs for two-layer networks and single-neuron networks.
result GD trajectories align on a specific bifurcation diagram independent of initialization.
Consider a group G and a family A of subgroups of G. We say that vertex finiteness holds for splittings of G over A if, up to isomorphism, there are only finitely many possibilities for vertex stabilizers of minimal G-trees with edge stabilizers in A. We show vertex finiteness when G…
Grokking occurs at numerical stability edge, requiring regularization to prevent.
problem Delayed generalization in deep learning models.
method Identified Softmax Collapse (SC) as the cause of grokking without regularization.
result Mitigating SC enables grokking without regularization.
Characterizes geometric actions on graphs with flexible stabilizers.
problem Understanding geometric actions on flexible stabilizers.
method Defining generalized fine actions and proving relative quasi-convexity criteria.
result Characterizes Bowditch boundary points in relatively geometric actions.
We establish existence of the eta-invariant as well as of the Atiyah-Patodi-Singer and the Cheeger-Gromov rho-invariants for a class of Dirac operators on an incomplete edge space. Our analysis applies in particular to the signature, the Gauss-Bonnet and the spin Dirac operator. We derive an analogue of the Atiyah-Pato…
GD at EoS edge minimizes logistic loss without monotonic convergence.
problem Understanding GD's implicit bias at the edge of stability.
method Theoretical analysis of logistic regression with constant stepsize GD.
result GD with any constant stepsize minimizes logistic loss over long time scales.
GCNs converge and remain stable on large random graphs, revealing geometric insights.
problem Understanding the behavior of GCNs on large, sparse random graphs.
method Analysis of GCNs on random graph models with latent variables and geometric edge probabilities.
result GCNs converge to their continuous counterparts as graph size increases, and are stable to small graph deformations.
Enhances neural network dynamics to boost computational capacity.
problem Improving computational capacity of neural networks.
method Introducing Phase Transition Adaptation to drive system dynamics towards edge of stability.
result Consistently achieves enhancement in computational capacity over multiple datasets.
The paper proves stability and convergence of minimal networks under curvature motion.
problem Stability and convergence of minimal networks under curvature motion.
method Proved Lojasiewicz-Simon gradient inequalities for minimal networks.
result Motion by curvature starting from networks close to minimal ones exists for all times and smoothly converges.
Gradient descent near stability threshold shows sharpness oscillations.
problem Understanding sharpness and stability in non-Euclidean norms during gradient descent.
method Interpreted EoS through Directional Smoothness, defined generalized sharpness for arbitrary norms.
result Non-Euclidean GD exhibits sharpness oscillations around the stability threshold.
The paper explores how data geometry influences generalization in neural networks.
problem Understanding generalization in overparameterized neural networks.
method Theoretical exploration of overparametrized two-layer ReLU networks trained below the edge of stability.
result Generalization bounds adapt to the intrinsic dimension of data distributions and deteriorate as data concentrates towards the unit sphere.
In the framework of homological characterizations of relative hyperbolicity, Groves and Manning posed the question of whether a simply connected 2-complex X with a linear homological isoperimetric inequality, a bound on the length of attaching maps of 2-cells and finitely many 2-cells adjacent to any edge must …
We investigate the relationship between stability and the existence of extremal Kähler metrics on certain toric surfaces. In particular, we consider how log stability depends on weights for toric surfaces whose moment polytope is a quadrilateral. We introduce a space of symplectic potentials for toric manifolds, which …
DAGgr aggregates multiple DAGs to stabilize causal structure learning.
problem Stability in learning causal structure from data.
method Model averaging of candidate DAGs weighted by predictive likelihood, with acyclicity enforced.
result DAGgr consistently outperforms individual DAGs and bootstrap-aggregation baselines.
Proposes a continuous flow model to understand and control instability in gradient descent for deep learning.
problem Understanding and controlling the instability of gradient descent in deep learning.
method Introduces the Principal Flow (PF), a continuous time flow that approximates gradient descent dynamics.
result The PF captures divergent and oscillatory behaviors of gradient descent, including escaping local minima and saddle points.
Stable commutator length scl_G(g) of an element g in a group G is an invariant for group elements sensitive to the geometry and dynamics of G. For any group G acting on a tree, we prove a sharp bound scl_G(g)>=1/2 for any g acting without fixed points, provided that the stabilizer of each edge is relatively torsion-fre…
We prove a general homological stability theorem for certain families of groups equipped with product maps, followed by two theorems of a new kind that give information about the last two homology groups outside the stable range. (These last two unstable groups are the "edge" in our title.) Applying our results to auto…
It has long been suggested that the biological brain operates at some critical point between two different phases, possibly order and chaos. Despite many indirect empirical evidence from the brain and analytical indication on simple neural networks, the foundation of this hypothesis on generic non-linear systems remain…
New findings show GD converges to a linear interpolator even with quadratic loss function under certain conditions.
problem Understanding convergence of Gradient Descent with quadratic loss functions.
method Parameterized linear regression with quadratic loss function, empirical and theoretical analysis.
result Gradient Descent converges to a linear interpolator even with quadratic loss function under the Edge of Stability regime.
Large learning rates lead to various implicit biases in nonconvex optimization.
problem Understanding the conditions under which large learning rates yield edge of stability, balancing, and catapult phenomena.
method Developed a global convergence theory for nonconvex functions without globally Lipschitz continuous gradient, focusing on functions with good regularity.
result These implicit biases are more likely to occur in functions with good regularity, and large learning rates favor flatter regions.
GD converges faster to flatter minima than gradient flow in shallow networks.
problem Understanding the dynamics of gradient descent in shallow linear networks.
method Analyzing the convergence rate and solution of gradient descent in depth-2 linear neural networks.
result GD converges linearly to flatter minima than gradient flow, even with large step sizes.
Debt swaps improve financial networks by optimizing clearing payments and stability.
problem Improving financial network stability and efficiency through debt swaps.
method Analyzing computational complexity of debt swaps, focusing on semi-positive swaps and v-improving swaps.
result Polynomial length of sequences of semi-positive v-improving swaps for ranking-based clearing, but NP-hard for arbitrary v-improving swaps.
Let G be a group acting on a tree with cyclic edge and vertex stabilizers. Then stable commutator length (scl) is rational in G. Furthermore, scl varies predictably and converges to rational limits in so-called "surgery" families. This is a homological analog of the phenomenon of geometric convergence in hyperbolic Deh…
The study proves how groups can be split with limited complexity.
problem Understanding the complexity of group splittings.
method Analyzing trees with finite stabilizers and their quotient structures.
result Deformation spaces of trees have maximal complexity.