Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

3468102136 · May 202619922001200920172026
48 results for Hessian rank 1

This paper uncovers the low-rank structure of neural network Hessians.

problem Understanding the structure of Hessians in neural networks.
method Proposes a decoupling conjecture to decompose layer-wise Hessians into Kronecker products of smaller matrices.
result Proves the structure of top eigenspaces in 2-layer networks and shows high overlap in top eigenvectors across different models.

New perspective on CNNs using Hessian maps reveals their structure.

problem Understanding the nature of Convolutional Neural Networks (CNNs).
method Developed a framework using Toeplitz representation of CNNs to reveal Hessian structure and prove rank bounds.
result Proved that the Hessian rank of CNNs grows as the square root of the number of parameters.

In this article we show that every geodesic is rank one and the Hessian of Busemann functions is positive definite for a harmonic Damek-Ricci space, a two step solvable Lie group with a left invariant metric. Moreover, the eigenspace of the Hessian of Busemann functions on a Hadamard manifold (M,g)(M,g) corresponding to e…

2017-02-13abs ↗pdf ↗

Locally conformally Hessian manifolds are dense in radiant ones of rank 1.

problem Characterizing locally conformally Hessian manifolds and their properties.
method Analyzing quotient spaces of Hessian manifolds and using statistical manifold theory.
result The set of radiant l.c.H. metrics of rank 1 is dense in all radiant l.c.H. metrics.

Analyzes Hessian spectrum for neural networks near optimal learning.

problem Understanding learning dynamics near optimal points in neural networks.
method Characterizes Hessian eigenspectrum for teacher-student problems, using analytical and numerical methods.
result The rank of the Hessian matrix determines effective number of parameters for non-linear networks.

The paper proves constant rank theorems for special Lagrangian equations.

problem Understanding saddle solutions and Liouville type results for special Lagrangian equations.
method Argument based on saddle solutions and Liouville type results for the special Lagrangian equation.
result Obtained constant rank theorems for saddle solutions to the special Lagrangian equation and the quadratic Hessian equation.

The study classifies Hessian rank 1 hypersurfaces in dimensions 2, 3, and 4.

problem Classifying Hessian rank 1 affinely homogeneous hypersurfaces in specific dimensions.
method Power Series Method of Equivalence, infinitesimal calculations.
result Identified all non-product constant Hessian rank 1 affinely homogeneous hypersurfaces in dimensions 2, 3, and 4.

No non-product Hessian rank 1 affine homogeneous hypersurfaces exist in dimensions 5 and above.

problem Identifying non-product Hessian rank 1 affine homogeneous hypersurfaces in higher dimensions.
method Developed a normal form for hypersurfaces under the affine group, up to order ≤ n+5, in any dimension n ≥ 2.
result Non-existence of non-product Hessian rank 1 affine homogeneous hypersurfaces in dimensions 5 and above.

The study proves the existence of kk-convex hypersurfaces for specific curvature equations.

problem Proving the existence of kk-convex hypersurfaces for Hessian curvature equations.
method Combining a priori estimates with the continuity method, and establishing a constant rank theorem.
result Existence and uniqueness of kk-convex hypersurfaces for both nonhomogeneous and homogeneous Hessian curvature equations.

New method proves asymptotic normality for matrix sensing problems.

problem Proving asymptotic normality for matrix sensing under general convex losses.
method Riemannian geometry to handle degeneracy of the Hessian due to rotational symmetry.
result Proves n(φ0φ)DN(0,(H)1)\sqrt{n}(φ^0-φ^*)\xrightarrow{D}N(0,(H^*)^{-1}) as non o\infty.

In this work we develop Curvature Propagation (CP), a general technique for efficiently computing unbiased approximations of the Hessian of any function that is computed using a computational graph. At the cost of roughly two gradient evaluations, CP can give a rank-1 approximation of the whole Hessian, and can be repe…

2012-06-27abs ↗pdf ↗

The goal of tensor completion is to fill in missing entries of a partially known tensor (possibly including some noise) under a low-rank constraint. This may be formulated as a least-squares problem. The set of tensors of a given multilinear rank is known to admit a Riemannian manifold structure, thus methods of Rieman…

2017-03-29abs ↗pdf ↗

A new quasi-Newton method uses cubic regularization to avoid saddle points in deep learning.

problem Avoiding saddle points and poor local minima in deep learning models.
method Limited-memory symmetric rank-one quasi-Newton approach with adaptive regularized cubics.
result The method effectively avoids saddle points and converges to better local minima.

New ACV method speeds up CV in high dimensions with approximate low-rank data.

problem Accurate model assessment in high-dimensional, large data settings with expensive algorithms.
method Developed a new ACV algorithm that uses low-rank approximations of the Hessian matrix.
result The new method is fast and accurate in the presence of approximate low-rank data.

Flat minima lead to better generalization in low-rank matrix recovery models.

problem Understanding why flat minima generalize well in overparameterized models.
method Analysis of overparameterized matrix and bilinear sensing, robust PCA, covariance matrix estimation, and neural networks with quadratic activation functions.
result Flat minima, measured by the trace of the Hessian, exactly recover the ground truth in low-rank matrix recovery models under standard statistical assumptions.

SGD's training dynamics align with Hessian and gradient spectra in high-dimensional classification tasks.

problem Understanding the spectra of Hessian and gradient matrices in high-dimensional classification tasks.
method Rigorous analysis of SGD dynamics and spectra of Hessian and gradient matrices.
result SGD trajectory and emergent outlier eigenspaces align with a common low-dimensional subspace in multi-class high-dimensional mixtures and neural networks.

In a complete Riemannian manifold (M,g)(M, g) if the hessian of a real valued function satisfies some suitable conditions then it restricts the geometry of (M,g)(M, g). In this paper we characterize all compact rank-1 symmetric spaces, as those Riemannian manifolds (M,g)(M, g) admitting a real valued function uu such that the …

1995-11-23abs ↗pdf ↗

Equivalent formulations for low-rank matrix optimization are proven.

problem Low-rank matrix optimization with rank constraints.
method Established geometric landscape connections between manifold and factorization formulations.
result Equivalence between manifold and factorization formulations at FOSPs, SOSPs, and strict saddles.

The paper explores the geometric structure of cost functions in multiple dimensions.

problem Understanding the geometric properties of cost functions in multidimensional settings.
method Analyzes the Hessian metric and geodesics in logarithmic and original coordinates.
result The geometry is one-dimensional in logarithmic coordinates but effectively (n1)(n-1)-dimensional in original coordinates.

New findings challenge the use of flatness measures in neural networks.

problem The validity of flatness measures in assessing generalization in neural networks.
method Analysis of Hessian-based flatness norms and their relation to generalization.
result Solutions with large weights and low loss are often sharper than expected, contradicting flatness measures.

Paper proposes HCDC to improve hyperparameter search efficiency.

problem Poor generalizability of dataset condensation across different hyperparameters.
method HCDC algorithm that matches hyperparameter gradients for synthetic validation dataset.
result HCDC effectively maintains validation-performance rankings of models.

Reinforcement Learning (RL) algorithms allow artificial agents to improve their action selections so as to increase rewarding experiences in their environments. Deep Reinforcement Learning algorithms require solving a nonconvex and nonlinear unconstrained optimization problem. Methods for solving the optimization probl…

2018-11-06abs ↗pdf ↗

Finite-sum optimization problems are ubiquitous in machine learning, and are commonly solved using first-order methods which rely on gradient computations. Recently, there has been growing interest in \emph{second-order} methods, which rely on both gradients and Hessians. In principle, second-order methods can require …

2016-11-15abs ↗pdf ↗

Muon optimizer outperforms GD in neural networks.

problem Optimizing matrix-structured parameters in neural networks.
method Muon optimizer specifically designed for matrix parameters, analyzing convergence rate and low-rank Hessian structure.
result Muon can outperform Gradient Descent due to its ability to leverage the low-rank structure of Hessian matrices.

Let ΩRnΩ\subset\mathbb R^n be a Lipschitz domain. Given 1p<kn1\leq p<k\leq n and any uW2,p(Ω)u\in W^{2,p}(Ω) belonging to the little Hölder class c1,αc^{1,α}, we construct a sequence uju_j in the same space with rankD2uj<k\operatorname{rank}D^2u_j<k almost everywhere such that ujuu_j\to u in C1,αC^{1,α} and weakly in W2,pW^{2,p}. This result i…

2017-10-25abs ↗pdf ↗

New algorithms estimate Jacobian matrices for large-scale machine learning.

problem Efficiently computing search directions for large nonlinear least squares.
method Exploit low-rank structure in Hessian to estimate Jacobian matrices.
result Two algorithms perform well compared to state-of-the-art methods.

A new algorithm reduces the time and space complexity for multinomial logistic bandits.

problem High-dimensional feedback in multinomial logistic bandits makes existing algorithms inefficient.
method Integrates frequent directions matrix sketching into OFUL-MLogB to reduce time and space complexity.
result Achieves a regret bound of ildeO(ΔT(KdlnΔT+m)T) ilde{\mathcal{O}}(Δ_T(Kd\lnΔ_T+m)\sqrt{T}).

FedLoRU improves FL efficiency by using low-rank updates.

problem Communication inefficiency and performance reduction in Federated Learning.
method Proposes FedLoRU, a low-rank update framework for FL, which reduces communication costs while maintaining performance.
result FedLoRU achieves convergence rates similar to FedAvg and is robust to heterogeneous and large numbers of clients.

Deep ReLU networks escape from the origin via saddle points with a low-rank bias.

problem Understanding the dynamics of gradient descent in deep ReLU networks.
method Analysis of escape directions and singular values of weight matrices.
result The first singular value of the \ell-th layer weight matrix is at least 14\ell^{\frac{1}{4}} larger than any other singular value.

EpiMer merges models by solving Fréchet mean on a Riemannian manifold.

problem Integrating knowledge from multiple models without retraining.
method EpiMer casts model merging as solving the Fréchet mean on a Riemannian manifold, restricting computation to a low-rank subspace.
result EpiMer outperforms flat-geometry methods on image classification tasks.

SGD updates align with a low-rank subspace but do not lead to further loss reduction.

problem Understanding the training dynamics of deep neural networks, particularly the role of the dominant subspace.
method Exploring whether neural networks can be trained within the dominant subspace of the loss Hessian.
result SGD updates, when projected onto the dominant subspace, do not decrease the training loss further, suggesting spurious alignment.

Novel Newton method for large-scale kernel methods using random features.

problem Efficiently solving large-scale finite-sum minimization problems in RKHS.
method Randomized feature-based Newton method for empirical risk minimization.
result Local superlinear and global linear convergence of the method.

The study proves that certain noncompact Hessian manifolds are diffeomorphic to R^n.

problem Characterizing complete noncompact Hessian manifolds with nonnegative Hessian sectional curvature.
method Using a geometric flow on noncompact affine Riemannian manifolds, constructing Hessian metrics, and proving diffeomorphism.
result Complete noncompact Hessian manifolds with nonnegative Hessian sectional curvature are diffeomorphic to R^n if their tangent bundle has maximal volume growth.