Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

14294357 · May 202619922001200920172026
48 results for Hessian eigenvalue

Study on eigenvalues of complex Hessian operator on pseudoconvex manifolds.

problem Eigenvalue problem for complex Hessian operator on pseudoconvex manifolds.
method Established C1,1C^{1,1}-regularity and uniqueness of the first eigenfunction, derived variational formula for the first eigenvalue.
result Derivation of a bifurcation-type theorem and geometric bounds for the eigenvalue.

Study estimates eigenvalues for concave Hessian operators on convex domains.

problem Estimating eigenvalues for concave elliptic Hessian operators.
method Investigates Dirichlet eigenvalue problem for a broad class of concave elliptic Hessian operators.
result Existence and properties of the first nonzero eigenvalue and eigenfunction.

Gradient descent forces neural network eigenvalues to a specific threshold.

problem Understanding why gradient descent drives eigenvalues to a specific threshold.
method Introduced edge coupling, a functional on consecutive iterate pairs, to explain the trajectory towards the eigenvalue threshold.
result Gradient descent forces the Hessian eigenvalue to the threshold 2/η2/η from arbitrary initialization.

The loss function of deep networks is known to be non-convex but the precise nature of this nonconvexity is still an active area of research. In this work, we study the loss landscape of deep networks through the eigendecompositions of their Hessian matrix. In particular, we examine how important the negative eigenvalu…

2019-02-06abs ↗pdf ↗

New stability conditions for ZO methods reveal unique regularization effects.

problem Understanding optimization dynamics of ZO methods in deep learning.
method Explicit step size conditions and stability bounds derived for ZO methods.
result ZO methods operate near the edge of stability, with regularization effects specific to Hessian trace vs. eigenvalue.

A new optimisation method efficiently scales Hessian-vector products for neural networks.

problem Challenges in applying second-order quasi-Newton methods due to large Hessian and non-convexity.
method Proposes an optimisation algorithm that asymptotically uses the exact inverse Hessian with modified eigenvalues.
result Demonstrates scalability and comparable performance to other optimisation methods in neural networks.

We prove a Lichnerowicz type lower bound for the first nontrivial eigenvalue of the pp-Laplacian on Kähler manifolds. Parallel to the p=2p = 2 case, the first eigenvalue lower bound is improved by using a decomposition of the Hessian on Kähler manifolds with positive Ricci curvature.

2018-04-29abs ↗pdf ↗

Better Hessian approximations improve influence function attributions in deep learning.

problem Influence functions are difficult to compute due to ill-conditioned Hessians, leading to poor data attribution performance.
method Investigated the impact of Hessian approximation quality on influence-function attributions in a controlled setting.
result Better Hessian approximations consistently yield better influence score quality.

In this article we show that every geodesic is rank one and the Hessian of Busemann functions is positive definite for a harmonic Damek-Ricci space, a two step solvable Lie group with a left invariant metric. Moreover, the eigenspace of the Hessian of Busemann functions on a Hadamard manifold (M,g)(M,g) corresponding to e…

2017-02-13abs ↗pdf ↗

Analyzes Hessian spectrum for neural networks near optimal learning.

problem Understanding learning dynamics near optimal points in neural networks.
method Characterizes Hessian eigenspectrum for teacher-student problems, using analytical and numerical methods.
result The rank of the Hessian matrix determines effective number of parameters for non-linear networks.

In a complete Riemannian manifold (M,g)(M, g) if the hessian of a real valued function satisfies some suitable conditions then it restricts the geometry of (M,g)(M, g). In this paper we characterize all compact rank-1 symmetric spaces, as those Riemannian manifolds (M,g)(M, g) admitting a real valued function uu such that the …

1995-11-23abs ↗pdf ↗

We consider self-similar solutions to mean curvature evolution of entire Lagrangian graphs. When the Hessian of the potential function uu has eigenvalues strictly uniformly between -1 and 1, we show that on the potential level all the shrinking solitons are quadratic polynomials while the expanding solitons are in one…

2009-05-24abs ↗pdf ↗

We prove Hessian comparison theorems, Laplacian comparison theorems and volume comparison theorems of Finsler manifolds under various curvature conditions. As applications, we derive Mckean type theorems for the first eigenvalue of Finsler manifolds, as well as generalize a result on fundamental group due to Milnor to …

2005-12-29abs ↗pdf ↗

Empirical study on SGD hyperparameters and adversarial robustness.

problem Effect of SGD hyperparameters on adversarial robustness and generalization.
method Empirical observation of learning rate, batch size, and momentum effects on adversarial robustness and generalization.
result Constant learning rate to batch size ratio leads to good generalization and almost constant adversarial robustness.

We prove Obata's rigidity theorem for metric measure spaces that satisfy a Riemannian curvature-dimension condition. Additionally, we show that a lower bound KK for the generalized Hessian of a sufficiently regular function uu holds if and only if uu is KK-convex. A corollary is also a rigidity result for higher or…

2014-10-20abs ↗pdf ↗

Noise injection regularizes Hessian, improving neural network training and generalization.

problem Regularizing over-parameterized neural networks with nonconvex and nonlinear geometry.
method Injecting isotropic Gaussian noise into weight matrices and designing a two-point estimate of the Hessian penalty.
result Effective regularization of Hessian improves generalization, achieving up to 2.4% test accuracy increase.

New findings challenge the use of flatness measures in neural networks.

problem The validity of flatness measures in assessing generalization in neural networks.
method Analysis of Hessian-based flatness norms and their relation to generalization.
result Solutions with large weights and low loss are often sharper than expected, contradicting flatness measures.

In this work, we will verify some comparison results on Kahler manifolds. They are complex Hessian comparison for the distance function from a closed complex submanifold of a Kahler manifold with holomorphic bisectional curvature bounded below by a constant, eigenvalue comparison and volume comparison in terms of scala…

2010-07-09abs ↗pdf ↗

Let L be an ample bundle over a compact complex manifold X. Fix a Hermitian metric in L whose curvature defines a Kähler metric on X. The Hessian of Mabuchi energy is a fourth-order elliptic operator D on functions which arises in the study of scalar curvature. We quantise D by the Hessian E(k) of balancing energy, a f…

2010-09-23abs ↗pdf ↗

New findings cast doubt on the role of λmaxλ_{max} in generalizing neural networks.

problem The role of λmaxλ_{max} in neural network generalization remains unclear.
method Experiments with various training interventions and batch sizes.
result Generalization benefits can vanish at larger batch sizes, challenging the role of λmaxλ_{max}.

Study eigenvalues of a nonlinear operator and apply to submanifolds with bounded mean curvature.

problem Eigenvalue of a nonlinear operator and submanifolds with bounded mean curvature.
method Lower estimate for eigenvalue using generalized Hausdorff measure.
result Improves understanding of the spectrum of submanifolds in R^n.

Spectral clustering is a standard approach to label nodes on a graph by studying the (largest or lowest) eigenvalues of a symmetric real matrix such as e.g. the adjacency or the Laplacian. Recently, it has been argued that using instead a more complicated, non-symmetric and higher dimensional operator, related to the n…

2014-06-07abs ↗pdf ↗

Enhances deep learning robustness to noise without sacrificing clean data accuracy.

problem Robustness of deep neural networks to input noise.
method Discriminative loss at penultimate layer and class-wise feature alignment with Gaussian noise.
result Improves robustness to various perturbations without degrading clean data accuracy.

We introduce a scalable measure of curvature for analyzing training dynamics of large language models.

problem Analyzing the training dynamics of large language models due to high computational cost of measuring Hessian sharpness.
method We introduce critical sharpness and relative critical sharpness as computationally efficient measures capturing Hessian sharpness phenomena.
result We provide the first demonstration of sharpness phenomena at scale up to 7B parameters.

Training neural networks involves finding minima of a high-dimensional non-convex loss function. Knowledge of the structure of this energy landscape is sparse. Relaxing from linear interpolations, we construct continuous paths between minima of recent neural network architectures on CIFAR10 and CIFAR100. Surprisingly, …

2018-03-02abs ↗pdf ↗

Large learning rates cause parameter instability, leading to better generalization.

problem Understanding why deep neural networks perform well despite operating outside the traditional stability regime.
method Analyzing the effect of large learning rates on the orientation of Hessian eigenvectors and parameter exploration.
result Large learning rates induce parameter instability, leading to better generalization through exploration of flatter regions of the loss landscape.