In this paper, by using the Bochner technique on almost Hermitian manifolds, we obtain a complex Hessian comparison for almost Hermitian manifolds generalizing the Laplacian comparison for almost Hermitian manifolds by Tossati, and reprove a diameter estimate for almost Hermitian manifolds by Gray. Moreover, we obtain …
We establish integral formulas and sharp two-sided bounds for the Ricci curvature, mean curvature and second fundamental form on a Riemannian manifold with boundary. As applications, sharp gradient and Hessian estimates are derived for the Dirichlet and Neumann eigenfunctions.
We introduce a scalable measure of curvature for analyzing training dynamics of large language models.
problem Analyzing the training dynamics of large language models due to high computational cost of measuring Hessian sharpness.
method We introduce critical sharpness and relative critical sharpness as computationally efficient measures capturing Hessian sharpness phenomena.
result We provide the first demonstration of sharpness phenomena at scale up to 7B parameters.
Study C2 estimates for p-Hessian equations on closed manifolds.
problem Estimating solutions to p-Hessian equations on closed Riemannian manifolds. method Introducing pseudo-solutions to generalize C-subsolution and proving C1 and C2 estimates. result Proves C2 estimates for general p-Hessian equations on closed manifolds under sharp conditions. Paper finds exact Hessian sharpness in deep matrix factorization.
problem Understanding the geometry of loss landscapes in deep matrix factorization.
method Presented the first exact expression for Hessian maximum eigenvalue.
result Spectral-norm balance is a sufficient condition for flatness in deep matrix factorization.
SAM improves neural network generalization by penalizing sharpness, clarifying its exact notion and mechanism.
problem Improving deep neural network generalization for various settings.
method Sharpness-Aware Minimization (SAM) technique that penalizes a notion of sharpness of the model.
result SAM regularizes the third notion of sharpness, most likely preferred for practical performance.
New Hessian estimators for Riemannian manifolds with reduced bias.
problem Estimating Hessians on Riemannian manifolds with reduced bias and computational efficiency.
method Introducing new stochastic zeroth-order Hessian estimators using O(1) function evaluations. result Achieved a bias bound of order O(γδ2) for analytic real-valued functions. New findings challenge the use of flatness measures in neural networks.
problem The validity of flatness measures in assessing generalization in neural networks.
method Analysis of Hessian-based flatness norms and their relation to generalization.
result Solutions with large weights and low loss are often sharper than expected, contradicting flatness measures.
New method accelerates neural network training by focusing on flat directions.
problem Improving neural network training speed and stability.
method Bulk-SGD, interpolated gradient methods.
result Updates along the Dominant subspace can accelerate convergence but compromise stability.
Study reveals flatness of Hessian metrics with non-negative Ricci curvature on foliation leaves.
problem Rigidity of Ricci curvature on Hessian manifold leaves.
method Analysis of Ricci curvature properties of Hessian metrics on foliation leaves.
result Non-negative Ricci curvature on a single leaf forces the Hessian metric to be flat and yields bounds on the first Betti number.
SAM optimizes deep networks by oscillating between sides of the minimum.
problem Improving performance of deep networks.
method Gradient-based optimization method that oscillates between sides of the minimum.
result SAM effectively performs gradient descent on the spectral norm of the Hessian, encouraging drift towards wider minima.
Develops new strategy for Hessian estimates in Lagrangian mean curvature equation.
problem Interior Hessian estimates for solutions with prescribed Lipschitz phases.
method Allard-type regularity theorem, geometric measure theory, geometry of Lagrangian graphs, De Giorgi-Nash-Moser iteration.
result Sharp interior Hessian estimates for solutions with critical and supercritical phases.
Estimates complex Hessian integral for complex Monge-Ampère equations.
problem Improving classical ABP estimate for complex settings.
method De Giorgi iteration method for complex Monge-Ampère equations.
result Sharp gradient estimates for complex Monge-Ampère equations.
Training avoids edge of stability by aligning Jacobian matrices.
problem Training neural networks on the edge of stability causes inaccuracies.
method Used an exponential Euler solver to prevent entering the edge of stability.
result Alignment of Jacobian matrices causes sharpness increase in Hessian.
New findings show mini-batch SGD operates in a 'Edge of Stochastic Stability' regime.
problem Understanding the stability and convergence of mini-batch SGD.
method Analyzing the mini-batch Hessian and its directional curvature.
result Mini-batch SGD operates in a different stability regime (Edge of Stochastic Stability) compared to full-batch GD.
SGD noise helps select flat minima by concentrating in sharp directions and being proportional to loss value.
problem Understanding the implicit regularization of SGD and selecting flat minima in over-parameterized models.
method Relating SGD's linear stability to the Frobenius norm of the Hessian and analyzing the alignment property of SGD noise.
result Flat minima are linearly stable for SGD, and their sharpness is bounded independently of model size and sample size.
The paper studies convexity of products of squared Euclidean distances.
problem Convexity of products of squared Euclidean distances.
method Proved a convexity principle and applied it to products of squared distances, computed Hessian-positive regions and exact convexity levels.
result Computed exact convexity and quasiconvexity truncation levels for the two-centre model.
Established in the 30's, Schauder {\it a priori} estimates are among the most classical and powerful tools in the analysis of problems ruled by 2nd order elliptic PDEs. Since then, a central problem in regularity theory has been to understand Schauder type estimates fashioning particular borderline scenarios. In such c…
ROOT-SGD solves convex optimization problems with optimal nonasymptotic and near-optimal asymptotic performance.
problem Solving strongly convex and smooth unconstrained optimization problems using stochastic first-order algorithms.
method ROOT-SGD: Recursive One-Over-T SGD, averaging past stochastic gradients.
result Achieves state-of-the-art performance in both nonasymptotic and asymptotic senses.
Proves smoothness and estimates for special Lagrangian solutions with semi-convexity.
problem Smoothness and estimates for special Lagrangian solutions.
method Viscosity solutions, smoothness, interior derivative estimates, sharpness of conditions.
result New Liouville theorem and effective Hessian estimates for special Lagrangian solutions.
Paper proposes a new flatness measure for neural networks to improve generalization.
problem Generalization in deep learning models, especially with overparameterization.
method Soft rank measure of the Hessian to assess flatness and generalization.
result Soft rank flatness measure accurately estimates generalization gaps for various models.
New stability conditions for ZO methods reveal unique regularization effects.
problem Understanding optimization dynamics of ZO methods in deep learning.
method Explicit step size conditions and stability bounds derived for ZO methods.
result ZO methods operate near the edge of stability, with regularization effects specific to Hessian trace vs. eigenvalue.
Large learning rates cause parameter instability, leading to better generalization.
problem Understanding why deep neural networks perform well despite operating outside the traditional stability regime.
method Analyzing the effect of large learning rates on the orientation of Hessian eigenvectors and parameter exploration.
result Large learning rates induce parameter instability, leading to better generalization through exploration of flatter regions of the loss landscape.
Noise injection regularizes Hessian, improving neural network training and generalization.
problem Regularizing over-parameterized neural networks with nonconvex and nonlinear geometry.
method Injecting isotropic Gaussian noise into weight matrices and designing a two-point estimate of the Hessian penalty.
result Effective regularization of Hessian improves generalization, achieving up to 2.4% test accuracy increase.
On H-type sub-Riemannian manifolds we establish sub-Hessian and sub-Laplacian comparison theorems which are uniform for a family of approximating Riemannian metrics converging to the sub-Riemannian one. We also prove a sharp sub-Riemannian Bonnet-Myers theorem that extends to this general setting results previously pro…
We find sharp bounds for the norm inequality on a Pseudo-hermitian manifold, where the L^2 norm of all second derivatives of the function involving horizontal derivatives is controlled by the L^2 norm of the sub-Laplacian. Perturbation allows us to get a-priori bounds for solutions to sub-elliptic PDE in non-divergence…
To understand the dynamics of optimization in deep neural networks, we develop a tool to study the evolution of the entire Hessian spectrum throughout the optimization process. Using this, we study a number of hypotheses concerning smoothness, curvature, and sharpness in the deep learning literature. We then thoroughly…
Noise in linear networks minimizes sharpness and leads to shrinkage-thresholding.
problem Minimizing sharpness in diagonal linear networks.
method Stochastic sharpness-aware minimization (SAM) with isotropic noise.
result Noise forces shrinkage-thresholding of true parameters.
The study analyzes sharpness dynamics in neural networks, revealing mechanisms and conditions.
problem Understanding sharpness in neural network training.
method Fixed point analysis and edge of stability analysis in a simplified 2-layer linear network.
result Reveals mechanisms behind sharpness trends, conditions for edge of stability, and a period-doubling route to chaos.
A simple function shows how neural nets can converge despite high sharpness.
problem Understanding why neural nets converge with high sharpness.
method Constructed a minimal example function and analyzed its training dynamics rigorously.
result Final converging point has sharpness close to 2/η. The paper studies the loss landscape of regularized deep matrix factorization, revealing unique and sharp minimizers.
problem Understanding the loss landscape and minimizers of regularized deep matrix factorization problems.
method Theoretical analysis of ℓ2-regularized deep matrix factorization/deep linear network training problems with squared-error loss. result The unique end-to-end minimizer exists for all target matrices except for a set of Lebesgue measure zero.
Let L be an ample bundle over a compact complex manifold X. Fix a Hermitian metric in L whose curvature defines a Kähler metric on X. The Hessian of Mabuchi energy is a fourth-order elliptic operator D on functions which arises in the study of scalar curvature. We quantise D by the Hessian E(k) of balancing energy, a f…
Gradient descent at edge of stability stabilizes implicitly, following projected gradient descent.
problem Gradient descent's stability and sharpness behavior at the edge of instability.
method Cubic Taylor expansion analysis of gradient descent dynamics.
result Gradient descent at edge of stability implicitly follows projected gradient descent.
Zeroth-order methods favor flat minima in machine learning.
problem Finding solutions with small Hessian trace in optimization.
method Zeroth-order optimization with two-point estimator.
result Zeroth-order optimization converges to flat minima.
Gradient descent near stability threshold shows sharpness oscillations.
problem Understanding sharpness and stability in non-Euclidean norms during gradient descent.
method Interpreted EoS through Directional Smoothness, defined generalized sharpness for arbitrary norms.
result Non-Euclidean GD exhibits sharpness oscillations around the stability threshold.
Enhanced estimates for ancient ovals and translators in 3D and 4D.
problem Sharp estimates for ancient ovals and translators.
method Derivation of gradient and Hessian estimates.
result Sharp gradient and Hessian estimates for ancient ovals and translators.
Gradient descent near stability threshold exhibits sharpness oscillations.
problem Understanding sharpness behavior near stability threshold in non-Euclidean norms.
method Interpreted EoS through Directional Smoothness and generalized sharpness under arbitrary norms.
result Non-Euclidean GD with generalized sharpness shows sharpness oscillations near 2/η. Deep linear networks minimize sharpness, avoiding large eigenvalues.
problem Understanding optimization dynamics in deep linear networks for regression.
method Analyzing sharpness (largest eigenvalue of Hessian) of minimizers and gradient flow solutions.
result Gradient flow implicitly regularizes towards flat minima, with sharpness bounded by a constant.
SAM improves generalization by operating near the edge of stability.
problem Improving generalization in neural networks.
method Sharpness-Aware Minimization (SAM) approach to training neural networks.
result SAM operates near the 'edge of stability' identified by the analysis.
Sharp bounds on heat kernel derivatives on incomplete manifolds.
problem Extending bounds on heat kernel derivatives to incomplete Riemannian manifolds.
method Analyzing heat kernels on incomplete Riemannian manifolds with conservative and non-conservative vector fields.
result Sharp bounds on all orders of heat kernel derivatives are established for incomplete manifolds.
SGD favors flat minima exponentially more than sharp minima in deep learning.
problem Understanding how SGD selects flat minima in deep learning.
method Developed a density diffusion theory (DDT) to analyze minima selection.
result SGD exponentially favors flat minima over sharp minima due to Hessian-dependent noise.
Overparameterization enhances SAM's effectiveness in minimizing sharpness.
problem Improving generalization in deep neural networks.
method Analysis of Sharpness-Aware Minimization (SAM) under varying degrees of overparameterization.
result Overparameterization significantly improves SAM's performance, particularly in noisy and sparse settings.
Cohen et al. (2021) show GD trajectories align on a bifurcation diagram.
problem Understanding the Edge of Stability (EoS) phenomenon in gradient descent.
method Empirical studies and rigorous mathematical proofs for two-layer networks and single-neuron networks.
result GD trajectories align on a specific bifurcation diagram independent of initialization.
The main technical result of the paper is a Bochner type formula for the sub-laplacian on a quaternionic contact manifold. With the help of this formula we establish a version of Lichnerowicz' theorem giving a lower bound of the eigenvalues of the sub-Laplacian under a lower bound on the Sp(n)Sp(1) components of the …
It has been empirically observed that the flatness of minima obtained from training deep networks seems to correlate with better generalization. However, for deep networks with positively homogeneous activations, most measures of sharpness/flatness are not invariant to rescaling of the network parameters, corresponding…
SAM improves neural network generalization better than SGD, especially in noisy data.
problem Overfitting in large neural networks with label noise.
method Sharpness-Aware Minimization (SAM) compared to Stochastic Gradient Descent (SGD).
result SAM prevents noise learning and facilitates feature learning better than SGD.
Label noise SGD converges to a simple model with a single linear feature.
problem Understanding the simplicity bias in neural network training.
method Analyzing the convergence of label noise SGD on two-layer neural networks.
result Label noise SGD converges to a model with a single linear feature.
We provide convergence guarantees in Wasserstein distance for a variety of variance-reduction methods: SAGA Langevin diffusion, SVRG Langevin diffusion and control-variate underdamped Langevin diffusion. We analyze these methods under a uniform set of assumptions on the log-posterior distribution, assuming it to be smo…