New criterion for solving inverse Hessian equations, including J-equation.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
ASTRA improves TDA by more accurately approximating iHVP.
WoodFisher improves neural network compression efficiency and accuracy.
In this paper, we consider the inverse hessian quotient curvature flow with star-shaped initial hypersurface in anti-de Sitter-Schwarzschild manifold. We prove that the solution exists for all time, and the second fundamental form converges to identity exponentially fast.
Proves smooth solutions for generalised Monge-Ampère equations on projective manifolds.
Study on curvature flow in Minkowski space for cocompact hypersurfaces.
Paper proposes HCDC to improve hyperparameter search efficiency.
Recently, Stochastic Gradient Markov Chain Monte Carlo (SG-MCMC) methods have been proposed for scaling up Monte Carlo computations to large data problems. Whilst these approaches have proven useful in many applications, vanilla SG-MCMC might suffer from poor mixing rates when random variables exhibit strong couplings …
We derive a priori estimates for solutions of a general class of fully non-linear equations on compact Hermitian manifolds. Our method is based on ideas that have been used for different specific equations, such as the complex Monge-Ampère, Hessian and inverse Hessian equations. As an application we solve a class of He…
FedNew improves federated learning efficiency and privacy.
In distributed optimization and distributed numerical linear algebra, we often encounter an inversion bias: if we want to compute a quantity that depends on the inverse of a sum of distributed matrices, then the sum of the inverses does not equal the inverse of the sum. An example of this occurs in distributed Newton's…
The paper solves a conjecture about spacelike hypersurfaces in de Sitter space.
Establishes statistical and computational bounds for influence diagnostics.
We provide a pointwise confidence bound for non-linear least-squares with fixed design.
We propose an algorithm for inexpensive gradient-based hyperparameter optimization that combines the implicit function theorem (IFT) with efficient inverse Hessian approximations. We present results about the relationship between the IFT and differentiating through optimization, motivating our algorithm. We use the pro…
Paper develops efficient methods for estimating Hessian inverses in stochastic optimization.
We provide a numerically robust and fast method capable of exploiting the local geometry when solving large-scale stochastic optimisation problems. Our key innovation is an auxiliary variable construction coupled with an inverse Hessian approximation computed using a receding history of iterates and gradients. It is th…
We propose an L-BFGS optimization algorithm on Riemannian manifolds using minibatched stochastic variance reduction techniques for fast convergence with constant step sizes, without resorting to linesearch methods designed to satisfy Wolfe conditions. We provide a new convergence proof for strongly convex functions wit…
We present two sampled quasi-Newton methods (sampled LBFGS and sampled LSR1) for solving empirical risk minimization problems that arise in machine learning. Contrary to the classical variants of these methods that sequentially build Hessian or inverse Hessian approximations as the optimization progresses, our proposed…
New techniques extend certified unlearning to deep neural networks.
Standard gradient descent methods are susceptible to a range of issues that can impede training, such as high correlations and different scaling in parameter space.These difficulties can be addressed by second-order approaches that apply a pre-conditioning matrix to the gradient to improve convergence. Unfortunately, s…
The paper characterizes when numerical criteria for PDE solvability fail and provides effective criteria for existence.
A new optimisation method efficiently scales Hessian-vector products for neural networks.
We develop a new modeling framework for Inter-Subject Analysis (ISA). The goal of ISA is to explore the dependency structure between different subjects with the intra-subject dependency as nuisance. It has important applications in neuroscience to explore the functional connectivity between brain regions under natural …
This paper investigates different vector step-size adaptation approaches for non-stationary online, continual prediction problems. Vanilla stochastic gradient descent can be considerably improved by scaling the update with a vector of appropriately chosen step-sizes. Many methods, including AdaGrad, RMSProp, and AMSGra…
We propose a fast second-order method that can be used as a drop-in replacement for current deep learning solvers. Compared to stochastic gradient descent (SGD), it only requires two additional forward-mode automatic differentiation operations per iteration, which has a computational cost comparable to two standard for…
Approximate Newton methods are a standard optimization tool which aim to maintain the benefits of Newton's method, such as a fast rate of convergence, whilst alleviating its drawbacks, such as computationally expensive calculation or estimation of the inverse Hessian. In this work we investigate approximate Newton meth…
Paper develops a distributed debiased estimator for sparse statistical inference.
The Delta method is a classical procedure for quantifying epistemic uncertainty in statistical models, but its direct application to deep neural networks is prevented by the large number of parameters . We propose a low cost variant of the Delta method applicable to -regularized deep neural networks based on th…
Pathfinder uses quasi-Newton optimization for variational inference.
Progress in deep learning is slowed by the days or weeks it takes to train large models. The natural solution of using more hardware is limited by diminishing returns, and leads to inefficient use of additional resources. In this paper, we present a large batch, stochastic optimization algorithm that is both faster tha…
New technique debiases distributed optimization, improving convergence rate.
A parallel optimization method for convex functions using Hessian sketching and debiasing.
Influence functions help study large language model generalization, revealing surprising decay patterns.
Optimization geometrodynamics simplifies adaptive optimizer dynamics.
New ACV method speeds up CV in high dimensions with approximate low-rank data.