Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

60119179238 · May 202619922001200920172026
48 results for Principal Hessian Directions

Survey of SDR methods for high-dimensional regression and embedding.

problem Reducing dimensionality in high-dimensional data.
method Involves both statistical and machine learning approaches, covering inverse and forward regression methods.
result Supervised Kernel Dimension Reduction is equivalent to supervised PCA.

Study geometric properties of loss functions to understand neural network performance.

problem Understanding the geometric properties of high-dimensional loss functions to improve neural network performance.
method Combine concepts from high-dimensional probability and differential geometry to study curvature properties in lower-dimensional loss representations.
result Mean curvature in the original loss space determines if saddle points appear as minima, maxima, or flat regions.

New algorithms estimate Hessians using random directions for faster stochastic optimization.

problem Efficiently estimating Hessians for stochastic optimization.
method Generalized Hessian estimators using random directions and noisy function measurements.
result Asymptotically unbiased estimators with lower bias for more measurements.

A new dimension reduction method based on Gaussian finite mixtures is proposed as an extension to sliced inverse regression (SIR). The model-based SIR (MSIR) approach allows the main limitation of SIR to be overcome, i.e., failure in the presence of regression symmetric relationships, without the need to impose further…

2015-08-10abs ↗pdf ↗

Study eigenvalues of a nonlinear operator and apply to submanifolds with bounded mean curvature.

problem Eigenvalue of a nonlinear operator and submanifolds with bounded mean curvature.
method Lower estimate for eigenvalue using generalized Hausdorff measure.
result Improves understanding of the spectrum of submanifolds in R^n.

Large learning rates cause parameter instability, leading to better generalization.

problem Understanding why deep neural networks perform well despite operating outside the traditional stability regime.
method Analyzing the effect of large learning rates on the orientation of Hessian eigenvectors and parameter exploration.
result Large learning rates induce parameter instability, leading to better generalization through exploration of flatter regions of the loss landscape.

Establishes log-concavity estimates for convex domains' first Dirichlet eigenfunctions.

problem Quantifying the Hessian of log-concave eigenfunctions on convex domains.
method Analyzes log-concavity properties of the first Dirichlet eigenfunction on convex domains.
result Obtains quantitative estimates for the Hessian of logu\log u.

A new optimisation method efficiently scales Hessian-vector products for neural networks.

problem Challenges in applying second-order quasi-Newton methods due to large Hessian and non-convexity.
method Proposes an optimisation algorithm that asymptotically uses the exact inverse Hessian with modified eigenvalues.
result Demonstrates scalability and comparable performance to other optimisation methods in neural networks.

The study examines principal directions and curvatures of Lagrangian submanifolds.

problem Understanding the geometry of Lagrangian submanifolds.
method Recalling and analyzing the extrinsic principal tangential and normal directions, and their corresponding curvatures for Lagrangian submanifolds in complex Euclidean spaces.
result Established natural relationships between distinguished tangential and normal directions and their curvatures for Lagrangian submanifolds.

We analyze the Hessian spectra of large models up to 100B parameters.

problem Accurate Hessian spectra of large foundation models are difficult to obtain.
method We use shard-local finite-difference Hessian vector products and stochastic Lanczos quadrature.
result We produce the first large-scale spectral density estimates of foundation models.

New method shows Hessian estimator from random samples converges to true Hessian on complex manifolds.

problem Uncertainty in Hessian estimator accuracy on complex manifolds with boundaries and nonuniform sampling.
method Locally fitting quadratic polynomials, rigorous theoretical analysis under mild conditions.
result The Hessian estimator asymptotically converges to the true Hessian, even near boundaries.

New method computes affine normal directions efficiently for sparse polynomials.

problem Computing affine normal directions is computationally expensive in high dimensions.
method Reduces third-order tensor contraction to matrix-free formulation using log-determinant gradient.
result Scalable implementations with near-linear scaling in dimension and sparsity.

Given a vector field XX in a Riemannian manifold, a hypersurface is said to have a canonical principal direction relative to XX if the projection of XX onto the tangent space of the hypersurface gives a principal direction. We give different ways for building these hypersurfaces, as well as a number of useful charac…

2011-10-10abs ↗pdf ↗

Study extends Hausdorff dimension Hessian results to new hyperconvex representations.

problem Extending classical results on Hausdorff dimension Hessian.
method Analyzes (1,1,2)-hyperconvex representations and small complex deformations.
result Positive definiteness of Hessian of Hausdorff dimension for co-compact Γ in PO(n,1).

Paper develops efficient methods for estimating Hessian inverses in stochastic optimization.

problem Estimating the inverse Hessian for convex function minimization.
method Robbins-Monro procedure for recursive estimation of the inverse Hessian.
result Develops universal stochastic Newton methods with improved efficiency.

Study on discrete surfaces with constant principal curvature for nanocarbon applications.

problem Understanding discrete geometry properties of nanocarbon materials.
method Developed discrete surface theory on 3-ary oriented trees, defined discrete principal directions, constructed examples of discrete CPC surfaces.
result Construction of discrete constant principal curvature surfaces, including discrete CPC tori.

Hessian-free (HF) optimization has been successfully used for training deep autoencoders and recurrent networks. HF uses the conjugate gradient algorithm to construct update directions through curvature-vector products that can be computed on the same order of time as gradients. In this paper we exploit this property a…

2013-01-16abs ↗pdf ↗

Autoencoders are a deep learning model for representation learning. When trained to minimize the distance between the data and its reconstruction, linear autoencoders (LAEs) learn the subspace spanned by the top principal directions but cannot learn the principal directions themselves. In this paper, we prove that $L_2…

2019-01-23abs ↗pdf ↗

CWGD measures gradient diversity weighted by curvature, improving SGD convergence.

problem Gradient noise in high-curvature directions is underestimated by standard methods.
method CWGD weights gradient diversity by the inverse square root of the Hessian.
result CWGD-Cosine reduces optimization error by up to 20% compared to standard cosine annealing.

We establish a classification of cubic minimal cones in case of the so-called radial eigencubics. Our principal result states that any radial eigencubic is either a member of the infinite family of eigencubics of Clifford type, or belongs to one of 18 exceptional families. We prove that at least 12 of the 18 families a…

2010-09-27abs ↗pdf ↗

Our work connects parameter magnitudes and Hessian eigenspaces in deep neural nets.

problem Understanding the relationship between parameter magnitudes and Hessian curvature in deep learning models.
method Developed a matrix-free algorithm based on sketched SVDs to measure similarity between parameter masks and Hessian eigenspaces.
result Top Hessian eigenvectors tend to be concentrated around larger parameters, indicating a connection between parameter magnitudes and loss curvature.

We consider surfaces in Euclidean space parametrized on an annular domain such that the first fundamental form and the principal curvatures are rotationally invariant, and the principal curvature directions only depend on the angle of rotation (but not the radius). Such surfaces generalize the Enneper surface. We show …

2016-07-28abs ↗pdf ↗

A new PCR method using SVD with sparse regularization.

problem Lack of response variable information in traditional PCR.
method One-stage SVD approach with two loss functions and sparse regularization.
result Obtains principal component loadings with response variable information.

Approximate Newton methods are a standard optimization tool which aim to maintain the benefits of Newton's method, such as a fast rate of convergence, whilst alleviating its drawbacks, such as computationally expensive calculation or estimation of the inverse Hessian. In this work we investigate approximate Newton meth…

2015-07-29abs ↗pdf ↗

Attention learns PCA on Gaussian data, proving its connection to principal component analysis.

problem Principal component analysis on Gaussian data.
method Analysis of attention mechanisms through PCA, covering finite and infinite prompt regimes.
result Attention aligns with principal eigenvectors of covariance matrices, converging to optimal solutions in the infinite-prompt limit.

Noise injection regularizes Hessian, improving neural network training and generalization.

problem Regularizing over-parameterized neural networks with nonconvex and nonlinear geometry.
method Injecting isotropic Gaussian noise into weight matrices and designing a two-point estimate of the Hessian penalty.
result Effective regularization of Hessian improves generalization, achieving up to 2.4% test accuracy increase.

A parallel optimization method for convex functions using Hessian sketching and debiasing.

problem Massively parallel optimization of convex functions with limited communication.
method Newton method with Hessian sketching and debiasing by workers, server averages descent directions.
result Approximation of Newton step with low-complexity adaptive sketching scheme.

The paper analyzes stability and convergence rates of entropic and Sinkhorn potentials.

problem Stability and convergence rates of entropic and Sinkhorn potentials.
method Semiconcavity properties of entropic potentials and Schrödinger bridges.
result Exponential convergence rates for gradient and Hessian of Sinkhorn iterates.

We introduce a scalable measure of curvature for analyzing training dynamics of large language models.

problem Analyzing the training dynamics of large language models due to high computational cost of measuring Hessian sharpness.
method We introduce critical sharpness and relative critical sharpness as computationally efficient measures capturing Hessian sharpness phenomena.
result We provide the first demonstration of sharpness phenomena at scale up to 7B parameters.

We explore the geometrical interpretation of the PCA based clustering algorithm Principal Direction Divisive Partitioning (PDDP). We give several examples where this algorithm breaks down, and suggest a new method, gap partitioning, which takes into account natural gaps in the data between clusters. Geometric features …

2012-11-17abs ↗pdf ↗

A neural network models pressure-Hessian from local velocity gradients in turbulent flows.

problem Modeling the pressure-Hessian from local velocity gradients in turbulent flows.
method Tensor basis neural network (TBNN) trained on DNS data.
result Neural network accurately captures key alignment statistics of the pressure-Hessian tensor.

We study submanifolds whose principal curvatures, counted with multiplicities, do not depend on the normal direction. Such submanifolds, which we briefly call CPC submanifolds, are always austere, hence minimal, and have constant principal curvatures. Well-known classes of examples include totally geodesic submanifolds…

2018-05-25abs ↗pdf ↗

In this paper we introduce a notion of parallel transport for principal bundles with connections over differentiable stacks. We show that principal bundles with connections over stacks can be recovered from their parallel transport thereby extending the results of Barrett, Caetano and Picken, and Schreiber and Waldof f…

2015-09-16abs ↗pdf ↗

Gradient descent forces neural network eigenvalues to a specific threshold.

problem Understanding why gradient descent drives eigenvalues to a specific threshold.
method Introduced edge coupling, a functional on consecutive iterate pairs, to explain the trajectory towards the eigenvalue threshold.
result Gradient descent forces the Hessian eigenvalue to the threshold 2/η2/η from arbitrary initialization.

A new algorithm reduces the time and space complexity for multinomial logistic bandits.

problem High-dimensional feedback in multinomial logistic bandits makes existing algorithms inefficient.
method Integrates frequent directions matrix sketching into OFUL-MLogB to reduce time and space complexity.
result Achieves a regret bound of ildeO(ΔT(KdlnΔT+m)T) ilde{\mathcal{O}}(Δ_T(Kd\lnΔ_T+m)\sqrt{T}).

Study on null hypersurfaces with constant angle in Lorentzian manifolds.

problem Understanding constant angle null hypersurfaces in Lorentzian manifolds.
method Introduced constant angle null hypersurfaces, analyzed with respect to a given ambient vector field, and provided classification results.
result Null hypersurfaces have a canonical principal direction when the vector field is closed and conformal.

New superintegrable systems derived from Frobenius structures.

problem Constructing second-order superintegrable systems.
method Using conification and direct product construction, applying to semi-simple and nilpotent algebras.
result Explicitly constructed second-order superintegrable systems in three dimensions.

SAM optimizes deep networks by oscillating between sides of the minimum.

problem Improving performance of deep networks.
method Gradient-based optimization method that oscillates between sides of the minimum.
result SAM effectively performs gradient descent on the spectral norm of the Hessian, encouraging drift towards wider minima.