Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

63125188250 · Jun 202019922001200920172026
48 results for Hessian regularization

New algorithm adds Hessian regularization to improve neural network robustness.

problem Improving neural network robustness against adversarial attacks.
method Proposes an efficient algorithm to train neural networks with Hessian operator-norm regularization.
result Hessian operator-norm regularization increases neural network robustness over input gradient regularization.

Noise injection regularizes Hessian, improving neural network training and generalization.

problem Regularizing over-parameterized neural networks with nonconvex and nonlinear geometry.
method Injecting isotropic Gaussian noise into weight matrices and designing a two-point estimate of the Hessian penalty.
result Effective regularization of Hessian improves generalization, achieving up to 2.4% test accuracy increase.

The paper studies the smoothness of critical points of variational integrals on Hessian spaces.

problem The study focuses on the regularity of critical points of variational integrals defined on Hessian spaces.
method The approach involves solving a fourth order nonlinear equation and analyzing the Hessian of the critical points.
result Smooth critical points with bounded Hessian are shown to be smooth provided their Hessian has small BMO.

Hessian alignment improves OOD generalization in deep learning.

problem Improving deep learning models' ability to generalize to out-of-distribution data.
method Analyzed Hessian and gradient alignment for domain generalization using recent OOD theory.
result Hessian alignment methods achieve promising performance on various OOD benchmarks.

Stochastic Variance-Reduced Cubic regularization (SVRC) algorithms have received increasing attention due to its improved gradient/Hessian complexities (i.e., number of queries to stochastic gradient/Hessian oracles) to find local minima for nonconvex finite-sum optimization. However, it is unclear whether existing SVR…

2019-01-31abs ↗pdf ↗

Study on eigenvalues of complex Hessian operator on pseudoconvex manifolds.

problem Eigenvalue problem for complex Hessian operator on pseudoconvex manifolds.
method Established C1,1C^{1,1}-regularity and uniqueness of the first eigenfunction, derived variational formula for the first eigenvalue.
result Derivation of a bifurcation-type theorem and geometric bounds for the eigenvalue.

New stability conditions for ZO methods reveal unique regularization effects.

problem Understanding optimization dynamics of ZO methods in deep learning.
method Explicit step size conditions and stability bounds derived for ZO methods.
result ZO methods operate near the edge of stability, with regularization effects specific to Hessian trace vs. eigenvalue.

We study Hessian fully nonlinear uniformly elliptic equations and show that the second derivatives of viscosity solutions of those equations (in 12 or more dimensions) can blow up in an interior point of the domain. We prove that the optimal interior regularity of such solutions is no more than C^{1+ε}, showing the opt…

2008-05-17abs ↗pdf ↗

Unified framework for understanding and optimizing training acceleration.

problem Challenges in optimizing training with regularization and acceleration techniques.
method Explains how AdaGrad, RMSProp, and Adam accelerate training, and derives a generalization for L1L_1-regularization.
result Derives a unified mathematical framework for understanding and optimizing training acceleration.

Adapts PDE method to prove LL^\infty estimates for complex Hessian equations.

problem Proving LL^\infty estimates for complex Hessian equations on transverse Kähler manifolds.
method Adapts PDE approach of Guo-Phong-Tong and Guo-Phong-Tong-Wang [17, 18].
result Obtains LL^\infty estimate for transverse complex Monge-Ampère equations.

Develops new strategy for Hessian estimates in Lagrangian mean curvature equation.

problem Interior Hessian estimates for solutions with prescribed Lipschitz phases.
method Allard-type regularity theorem, geometric measure theory, geometry of Lagrangian graphs, De Giorgi-Nash-Moser iteration.
result Sharp interior Hessian estimates for solutions with critical and supercritical phases.

In this paper, complex Hessian equation over Kähler manifold was studied. Under the condition that the underline Kähler manifold has non-negative holomorphic bisectional curvature, the existence and regularity of the solution was proved.

2008-12-24abs ↗pdf ↗

The paper describes a fine representation of the Ricci tensor and Hessian on RCD spaces.

problem Understanding the structure of the Ricci tensor and Hessians on RCD spaces.
method Polar decomposition of the Ricci tensor and Hessians, providing regularity results.
result The Ricci tensor and Hessians on RCD spaces can be represented by a polar decomposition, revealing their regularity.

The paper studies the loss landscape of regularized deep matrix factorization, revealing unique and sharp minimizers.

problem Understanding the loss landscape and minimizers of regularized deep matrix factorization problems.
method Theoretical analysis of 2\ell^2-regularized deep matrix factorization/deep linear network training problems with squared-error loss.
result The unique end-to-end minimizer exists for all target matrices except for a set of Lebesgue measure zero.

A new quasi-Newton method uses cubic regularization to avoid saddle points in deep learning.

problem Avoiding saddle points and poor local minima in deep learning models.
method Limited-memory symmetric rank-one quasi-Newton approach with adaptive regularized cubics.
result The method effectively avoids saddle points and converges to better local minima.

The Hessian geometry is the real analogue of the Kähler one. Sasakian geometry is an odd-dimensional counterpart of the Kähler geometry. In the paper, we study the connection between projective Hessian and Sasakian manifolds analogous to the one between Hessian and Kähler manifolds. In particular, we construct a Sasaki…

2018-03-07abs ↗pdf ↗

New algorithm improves convergence of gradient boosting trees.

problem Global convergence of Newton boosting in tabular machine learning.
method Introduces Gradient Regularized Newton Descent for GBDTs, proving linear convergence for smooth, strongly convex losses and O(1k2)\mathcal{O}(\frac{1}{k^2}) rate for general convex losses.
result Achieves globally convergent second-order GBDT algorithm with rate matching first-order boosting.

Study shows how to balance memory and learning efficiency in continual learning.

problem Balancing memory and learning efficiency in continual learning.
method Structural regularization with Hessian-based regularization.
result Structural regularization improves statistical performance at the cost of increased memory complexity.

The nonzero level sets of a homogeneous, logarithmically homogeneous, or translationally homogeneous function are affine spheres if and only if the Hessian determinant of the function is a multiple of a power or an exponential of the function. In particular, the nonzero level sets of a homogeneous polynomial are proper…

2013-07-20abs ↗pdf ↗

RES, a regularized stochastic version of the Broyden-Fletcher-Goldfarb-Shanno (BFGS) quasi-Newton method is proposed to solve convex optimization problems with stochastic objectives. The use of stochastic gradient descent algorithms is widespread, but the number of iterations required to approximate optimal arguments c…

2014-01-29abs ↗pdf ↗

New technique debiases distributed optimization, improving convergence rate.

problem Bias in local estimates limits effectiveness of distributed second order optimization.
method Surrogate sketching and scaled regularization to eliminate bias.
result The debiased local estimates lead to faster convergence in distributed optimization.

The rapid development of computer hardware and Internet technology makes large scale data dependent models computationally tractable, and opens a bright avenue for annotating images through innovative machine learning algorithms. Semi-supervised learning (SSL) has consequently received intensive attention in recent yea…

2019-04-23abs ↗pdf ↗

Studied how SGD's stability regularization affects generalization in neural networks.

problem Understanding why SGD often generalizes better than GD in neural networks.
method Analyzed stability of SGD and GD through Frobenius norm and trace of Hessian, and compared their generalization properties.
result Stable minima of SGD generalize well, while GD's stability-induced regularization is too weak.

Large scale optimization problems are ubiquitous in machine learning and data analysis and there is a plethora of algorithms for solving such problems. Many of these algorithms employ sub-sampling, as a way to either speed up the computations and/or to implicitly implement a form of statistical regularization. In this …

2016-01-18abs ↗pdf ↗

We reparametrize ReLU NNs as splines to understand their learning dynamics.

problem Understanding the learning dynamics and inductive bias of neural networks.
method Reparametrize ReLU NNs as continuous piecewise linear splines to study learning dynamics.
result Standard weight initializations yield very flat functions, leading to strength and type of implicit regularization.

The purpose of this article is to present a new regularization technique of quasi-plurisubharmoinc functions on a compact Kaehler manifold. The idea is to regularize the function on local coordinate balls first, and then glue each piece together. Therefore, all the higher order terms in the complex Hessian of this regu…

2017-05-20abs ↗pdf ↗

Subsampled Newton methods approximate Hessian matrices through subsampling techniques, alleviating the cost of forming Hessian matrices but using sufficient curvature information. However, previous results require Ω(d)Ω(d) samples to approximate Hessians, where dd is the dimension of data points, making it less practica…

2019-02-13abs ↗pdf ↗

Introduces HTV to measure function complexity in learning schemes.

problem Assessing the complexity of supervised-learning schemes.
method Defines Hessian-Schatten total variation (HTV) as a seminorm to quantify function complexity.
result HTV is invariant to rotations, scalings, and translations, and its minimum value is achieved for linear mappings.

Constructs flows on manifolds with small curvature, proving Euclidean topology.

problem Geometric structure of manifolds with unbounded curvature.
method Distance like functions with integral hessian bound, Ricci flows.
result Manifolds with Ricci lower bound, non-negative scalar curvature, bounded entropy, Ahlfors nn-regular and small curvature concentration are topologically Euclidean.

This paper tackles catastrophic forgetting in neural networks by providing a unified framework for regularization-based continual learning.

problem Catastrophic forgetting in neural networks trained sequentially on multiple tasks.
method Formulates regularization-based continual learning as a second-order Taylor approximation of the loss function, leading to a unified framework.
result Theoretical results indicate the importance of accurate approximation of the Hessian matrix for optimization and generalization.

The paper analyzes stability and convergence rates of entropic and Sinkhorn potentials.

problem Stability and convergence rates of entropic and Sinkhorn potentials.
method Semiconcavity properties of entropic potentials and Schrödinger bridges.
result Exponential convergence rates for gradient and Hessian of Sinkhorn iterates.

Hamiltonian Monte Carlo (HMC) is a widely deployed method to sample from high-dimensional distributions in Statistics and Machine learning. HMC is known to run very efficiently in practice and its popular second-order "leapfrog" implementation has long been conjectured to run in d1/4d^{1/4} gradient evaluations. Here we …

2018-02-24abs ↗pdf ↗

SGD without replacement decouples into curvature-following and flatness-regularizing steps.

problem Theoretical analysis of SGD without replacement for large-scale neural networks.
method Analysis of SGD without replacement in a realistic regime, considering high curvature and flatness.
result Optimizing with SGD without replacement is locally equivalent to an additional regularizer step.