Study on eigenvalues of complex Hessian operator on pseudoconvex manifolds.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study calculates Hessian of Busemann function on Damek-Ricci spaces.
Study estimates eigenvalues for concave Hessian operators on convex domains.
Abstract: Characterizes special Kähler manifolds with specific properties.
The paper shows infinitely many components in Floer Hessians space.
In this paper, by using the Bochner technique on almost Hermitian manifolds, we obtain a complex Hessian comparison for almost Hermitian manifolds generalizing the Laplacian comparison for almost Hermitian manifolds by Tossati, and reprove a diameter estimate for almost Hermitian manifolds by Gray. Moreover, we obtain …
Analyzes the Hessian of ReLU networks, proving skewed eigenvalue distribution.
Gradient descent forces neural network eigenvalues to a specific threshold.
The loss function of deep networks is known to be non-convex but the precise nature of this nonconvexity is still an active area of research. In this work, we study the loss landscape of deep networks through the eigendecompositions of their Hessian matrix. In particular, we examine how important the negative eigenvalu…
New stability conditions for ZO methods reveal unique regularization effects.
A new optimisation method efficiently scales Hessian-vector products for neural networks.
New insights into matrix factorization show strict saddles have bounded eigenvalues.
We prove a Lichnerowicz type lower bound for the first nontrivial eigenvalue of the -Laplacian on Kähler manifolds. Parallel to the case, the first eigenvalue lower bound is improved by using a decomposition of the Hessian on Kähler manifolds with positive Ricci curvature.
Better Hessian approximations improve influence function attributions in deep learning.
Studying SGD on deep neural networks using diffusion maps.
In this article we show that every geodesic is rank one and the Hessian of Busemann functions is positive definite for a harmonic Damek-Ricci space, a two step solvable Lie group with a left invariant metric. Moreover, the eigenspace of the Hessian of Busemann functions on a Hadamard manifold corresponding to e…
To understand the dynamics of optimization in deep neural networks, we develop a tool to study the evolution of the entire Hessian spectrum throughout the optimization process. Using this, we study a number of hypotheses concerning smoothness, curvature, and sharpness in the deep learning literature. We then thoroughly…
Analyzes Hessian spectrum for neural networks near optimal learning.
We give a lower estimate of the gap of the first two eigenvalues of the Schrodinger operator in the case when the potential is strongly convex. In particular, if the Hessian of the potential is bounded from below by a positive constant, the gap has a lower bound independent of the dimension. We also estimate the gap wh…
New Bethe-Hessian method improves community detection in sparse networks.
Gradient descent performs well on weakly convex losses, offering generalization guarantees.
In a complete Riemannian manifold if the hessian of a real valued function satisfies some suitable conditions then it restricts the geometry of . In this paper we characterize all compact rank-1 symmetric spaces, as those Riemannian manifolds admitting a real valued function such that the …
Paper finds exact Hessian sharpness in deep matrix factorization.
Characterizes Hessian eigenspectra for realistic nonlinear models.
We consider self-similar solutions to mean curvature evolution of entire Lagrangian graphs. When the Hessian of the potential function has eigenvalues strictly uniformly between -1 and 1, we show that on the potential level all the shrinking solitons are quadratic polynomials while the expanding solitons are in one…
Data structure affects deep learning performance, study finds.
We prove Hessian comparison theorems, Laplacian comparison theorems and volume comparison theorems of Finsler manifolds under various curvature conditions. As applications, we derive Mckean type theorems for the first eigenvalue of Finsler manifolds, as well as generalize a result on fundamental group due to Milnor to …
The paper finds geodesics in Kähler potentials with no degeneration.
Empirical study on SGD hyperparameters and adversarial robustness.
We prove a CR Obata type result that if the first positive eigenvalue of the sub-Laplacian on a compact strictly pseudoconvex pseudohermitian manifold with a divergence free pseudohermitian torsion takes the smallest possible value then, up to a homothety of the pseudohermitian structure, the manifold is the standart S…
We prove Obata's rigidity theorem for metric measure spaces that satisfy a Riemannian curvature-dimension condition. Additionally, we show that a lower bound for the generalized Hessian of a sufficiently regular function holds if and only if is -convex. A corollary is also a rigidity result for higher or…
Noise injection regularizes Hessian, improving neural network training and generalization.
New findings challenge the use of flatness measures in neural networks.
In this work, we will verify some comparison results on Kahler manifolds. They are complex Hessian comparison for the distance function from a closed complex submanifold of a Kahler manifold with holomorphic bisectional curvature bounded below by a constant, eigenvalue comparison and volume comparison in terms of scala…
Let L be an ample bundle over a compact complex manifold X. Fix a Hermitian metric in L whose curvature defines a Kähler metric on X. The Hessian of Mabuchi energy is a fourth-order elliptic operator D on functions which arises in the study of scalar curvature. We quantise D by the Hessian E(k) of balancing energy, a f…
Study on convex capillary hypersurfaces with Lp curvature in half-space.
New findings cast doubt on the role of in generalizing neural networks.
Study eigenvalues of a nonlinear operator and apply to submanifolds with bounded mean curvature.
The Hessian of neural networks can be decomposed into a sum of two matrices: (i) the positive semidefinite generalized Gauss-Newton matrix G, and (ii) the matrix H containing negative eigenvalues. We observe that for wider networks, minimizing the loss with the gradient descent optimization maneuvers through surfaces o…
PWGF escapes saddle points in nonconvex optimization.
Spectral clustering is a standard approach to label nodes on a graph by studying the (largest or lowest) eigenvalues of a symmetric real matrix such as e.g. the adjacency or the Laplacian. Recently, it has been argued that using instead a more complicated, non-symmetric and higher dimensional operator, related to the n…
Enhances deep learning robustness to noise without sacrificing clean data accuracy.
A C^2 function on C^n is called (n-1)-plurisubharmonic in the sense of Harvey-Lawson if the sum of any n-1 eigenvalues of its complex Hessian is nonnegative. We show that the associated Monge-Ampere equation can be solved on any compact Kahler manifold. As a consequence we prove the existence of solutions to an equatio…
We consider large scale empirical risk minimization (ERM) problems, where both the problem dimension and variable size is large. In these cases, most second order methods are infeasible due to the high cost in both computing the Hessian over all samples and computing its inverse in high dimensions. In this paper, we pr…
We introduce a scalable measure of curvature for analyzing training dynamics of large language models.
The main technical result of the paper is a Bochner type formula for the sub-laplacian on a quaternionic contact manifold. With the help of this formula we establish a version of Lichnerowicz' theorem giving a lower bound of the eigenvalues of the sub-Laplacian under a lower bound on the components of the …
Training neural networks involves finding minima of a high-dimensional non-convex loss function. Knowledge of the structure of this energy landscape is sparse. Relaxing from linear interpolations, we construct continuous paths between minima of recent neural network architectures on CIFAR10 and CIFAR100. Surprisingly, …
Large learning rates cause parameter instability, leading to better generalization.