Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

326395126 · May 202619922001200920172026
48 results for Hessian alignment

Hessian alignment improves OOD generalization in deep learning.

problem Improving deep learning models' ability to generalize to out-of-distribution data.
method Analyzed Hessian and gradient alignment for domain generalization using recent OOD theory.
result Hessian alignment methods achieve promising performance on various OOD benchmarks.

Unified approach to domain generalization by aligning gradients and Hessians.

problem Developing models that generalize well across unseen domains.
method Moment Alignment, extending transfer measure to DG, aligning derivatives across domains.
result Moment Alignment unifies gradient and Hessian matching approaches, improving generalizability.

SGD's training dynamics align with Hessian and gradient spectra in high-dimensional classification tasks.

problem Understanding the spectra of Hessian and gradient matrices in high-dimensional classification tasks.
method Rigorous analysis of SGD dynamics and spectra of Hessian and gradient matrices.
result SGD trajectory and emergent outlier eigenspaces align with a common low-dimensional subspace in multi-class high-dimensional mixtures and neural networks.

SGD updates align with a low-rank subspace but do not lead to further loss reduction.

problem Understanding the training dynamics of deep neural networks, particularly the role of the dominant subspace.
method Exploring whether neural networks can be trained within the dominant subspace of the loss Hessian.
result SGD updates, when projected onto the dominant subspace, do not decrease the training loss further, suggesting spurious alignment.

We analyze the Hessian spectra of large models up to 100B parameters.

problem Accurate Hessian spectra of large foundation models are difficult to obtain.
method We use shard-local finite-difference Hessian vector products and stochastic Lanczos quadrature.
result We produce the first large-scale spectral density estimates of foundation models.

Our work connects parameter magnitudes and Hessian eigenspaces in deep neural nets.

problem Understanding the relationship between parameter magnitudes and Hessian curvature in deep learning models.
method Developed a matrix-free algorithm based on sketched SVDs to measure similarity between parameter masks and Hessian eigenspaces.
result Top Hessian eigenvectors tend to be concentrated around larger parameters, indicating a connection between parameter magnitudes and loss curvature.

Cohen et al. (2021) show GD trajectories align on a bifurcation diagram.

problem Understanding the Edge of Stability (EoS) phenomenon in gradient descent.
method Empirical studies and rigorous mathematical proofs for two-layer networks and single-neuron networks.
result GD trajectories align on a specific bifurcation diagram independent of initialization.

The paper solves a conjecture about spacelike hypersurfaces in de Sitter space.

problem Proving an Alexandrov-Fenchel inequality for closed 2-convex spacelike hypersurfaces in de Sitter space.
method Investigating the locally constrained inverse curvature flow to establish the inequality.
result Established an Alexandrov-Fenchel inequality for closed 2-convex spacelike hypersurfaces in de Sitter space.

The paper studies hybrid connections on Hessian manifolds and their properties.

problem Investigating hybrid connections on Hessian manifolds.
method Defining and analyzing hybrid connections as incompressible affine connections projective to a flat connection DD.
result The difference ablaD abla - D is determined by the logarithmic differential of a Hessian potential function.

SGD noise helps select flat minima by concentrating in sharp directions and being proportional to loss value.

problem Understanding the implicit regularization of SGD and selecting flat minima in over-parameterized models.
method Relating SGD's linear stability to the Frobenius norm of the Hessian and analyzing the alignment property of SGD noise.
result Flat minima are linearly stable for SGD, and their sharpness is bounded independently of model size and sample size.

SAIL-RevKL improves SAIL's convergence by regularizing the objective function.

problem Convergence of self-improving online LLM alignment algorithms.
method Proposed SAIL-RevKL, a regularized objective function to improve optimization landscape.
result Proved SAIL-RevKL satisfies the Polyak-Lojasiewicz (PL) condition with near-linear sample complexity.

Enhances deep learning robustness to noise without sacrificing clean data accuracy.

problem Robustness of deep neural networks to input noise.
method Discriminative loss at penultimate layer and class-wise feature alignment with Gaussian noise.
result Improves robustness to various perturbations without degrading clean data accuracy.

Improved robustness in optimization methods using second-order information.

problem Scalability and sensitivity to mini-batch size in optimization methods.
method Mini-Batch Stochastic Variance-Reduced Newton (extttMbSVRN exttt{Mb-SVRN}) algorithm incorporating partial second-order information.
result Achieves a fast linear convergence rate independent of mini-batch size for large data sizes.

Paper solves a new Minkowski problem for a specific type of rigidity.

problem Solving a new Minkowski problem for a specific type of rigidity.
method Developed a nonlinear partial differential equation and used a curvature flow method.
result Existence of smooth non-even solutions to the p-th dual Minkowski problem for p < n-2.

EoS selectively shapes learning, affecting some groups more than others.

problem EoS affects learning differently across the data distribution.
method Branching intervention to enter or exit EoS regime, controlled perturbation to isolate mechanisms.
result EoS redistributes learning, amplifying progress on some groups and suppressing others.

We analyze why some models resist unlearning using linear stability theory.

problem Understanding and predicting when machine learning models resist unlearning.
method Linear stability theory applied to machine learning models, focusing on data coherence and optimization dynamics.
result Data coherence and signal-to-noise ratio (SNR) influence unlearning resistance; lower SNR makes unlearning easier.

We analyze the performance of a class of manifold-learning algorithms that find their output by minimizing a quadratic form under some normalization constraints. This class consists of Locally Linear Embedding (LLE), Laplacian Eigenmap, Local Tangent Space Alignment (LTSA), Hessian Eigenmaps (HLLE), and Diffusion maps.…

2008-06-16abs ↗pdf ↗

New findings show mini-batch SGD operates in a 'Edge of Stochastic Stability' regime.

problem Understanding the stability and convergence of mini-batch SGD.
method Analyzing the mini-batch Hessian and its directional curvature.
result Mini-batch SGD operates in a different stability regime (Edge of Stochastic Stability) compared to full-batch GD.

We present local ensembles, a method for detecting underspecification -- when many possible predictors are consistent with the training data and model class -- at test time in a pre-trained model. Our method uses local second-order information to approximate the variance of predictions across an ensemble of models from…

2019-10-21abs ↗pdf ↗

The study proves that certain noncompact Hessian manifolds are diffeomorphic to R^n.

problem Characterizing complete noncompact Hessian manifolds with nonnegative Hessian sectional curvature.
method Using a geometric flow on noncompact affine Riemannian manifolds, constructing Hessian metrics, and proving diffeomorphism.
result Complete noncompact Hessian manifolds with nonnegative Hessian sectional curvature are diffeomorphic to R^n if their tangent bundle has maximal volume growth.

New Hessian estimates for heat equations on manifolds.

problem Estimating Hessian matrices for heat-type equations on Riemannian manifolds.
method Using Bismut-Stroock Hessian formula, with explicit coefficients and delay/growth rate functions.
result Novel backward weak Harnack inequality and precise pointwise Hessian estimates for eigenfunctions.

New framework to understand and exploit curvature in deep learning loss landscapes.

problem Understanding and optimizing the loss landscape in deep learning models.
method New conceptual framework and techniques to estimate and exploit curvature of expected loss changes.
result Alice algorithm optimizes training by incorporating curvature terms and step bounds.

SAM improves neural network generalization by penalizing sharpness, clarifying its exact notion and mechanism.

problem Improving deep neural network generalization for various settings.
method Sharpness-Aware Minimization (SAM) technique that penalizes a notion of sharpness of the model.
result SAM regularizes the third notion of sharpness, most likely preferred for practical performance.

We prove that, in dimensions greater than 2, the generic metric is not a Hessian metric and find a curvature condition on Hessian metrics in dimensions greater than 3. In particular we prove that the forms used to define the Pontryagin classes in terms of the curvature vanish on a Hessian manifold. By contrast all anal…

2013-12-04abs ↗pdf ↗

Curved Frobenius manifolds link to Hessian metrics in geometry.

problem Understanding curved Frobenius manifolds and their relation to Hessian metrics.
method Analyzing the relationship between curved Frobenius structures and Hessian metrics on spaces with non-vanishing curvature.
result Consistent curved Frobenius structures on constant curvature spaces are linked to Hessian metrics.

Criterion for solvability of complex 2-Hessian equation on compact Kähler manifolds.

problem Solvability of complex 2-Hessian equation on compact Kähler manifolds.
method Nakai--Moishezon-type criterion associated with the complex 2-Hessian equation.
result Criterion equivalent to existence of a smooth 2-admissible representative in complex dimension three.

Constructs homogeneous Kähler structures on tangent bundles of Hessian manifolds.

problem Creating Kähler structures on tangent bundles of Hessian manifolds.
method Endowing Hessian manifolds with Kähler structures using group actions and homothetic vector fields.
result Homogeneous conformally Kähler structures on tangent bundles of selfsimilar Hessian manifolds.

Paper proves inequalities on Hermitian manifolds with applications to bounded solutions.

problem Establishing mixed Hessian inequalities on Hermitian manifolds.
method Weak convergence theorem of complex Hessian operators and general mixed Hessian inequality.
result Existence of bounded solutions of complex Hessian equations.

A selfsimiar manifold is a Riemannian manifold (M,g)\left(M,g\right) endowed with a homothetic vector field ξξ. We characterize global selfsimilar manifolds and describe the structure of local selfsimilar manifolds. We prove that any selfsimilar manifold with a potential homothetic vector field is a conical Riemannian ma…

2019-08-05abs ↗pdf ↗

This paper uncovers the low-rank structure of neural network Hessians.

problem Understanding the structure of Hessians in neural networks.
method Proposes a decoupling conjecture to decompose layer-wise Hessians into Kronecker products of smaller matrices.
result Proves the structure of top eigenspaces in 2-layer networks and shows high overlap in top eigenvectors across different models.