ASTRA improves TDA by more accurately approximating iHVP.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
WoodFisher improves neural network compression efficiency and accuracy.
New SGD algorithm finds critical points faster with second-order corrections.
A new optimisation method efficiently scales Hessian-vector products for neural networks.
Establishes statistical and computational bounds for influence diagnostics.
HessFormer enables distributed Hessian computation for large models.
The Hessian-vector product has been utilized to find a second-order stationary solution with strong complexity guarantee (e.g., almost linear time complexity in the problem's dimensionality). In this paper, we propose to further reduce the number of Hessian-vector products for faster non-convex optimization. Previous a…
Stochastic Variance-Reduced Cubic regularization (SVRC) algorithms have received increasing attention due to its improved gradient/Hessian complexities (i.e., number of queries to stochastic gradient/Hessian oracles) to find local minima for nonconvex finite-sum optimization. However, it is unclear whether existing SVR…
The paper studies efficient Hessian fitting methods for stochastic optimization.
We propose a reduction for non-convex optimization that can (1) turn an stationary-point finding algorithm into an local-minimum finding one, and (2) replace the Hessian-vector product computations with only gradient computations. It works both in the stochastic and the deterministic settings, without hurting the algor…
A new algorithm reduces communication rounds for distributed convex optimization.
Improves understanding of neural network predictions using influence functions.
This paper proposes a stochastic variant of a classic algorithm---the cubic-regularized Newton method [Nesterov and Polyak 2006]. The proposed algorithm efficiently escapes saddle points and finds approximate local minima for general smooth, nonconvex functions in only stochastic gradien…
HTE improves PINNs for high-dimensional, high-order PDEs by reducing computational cost and memory usage.
Progress in deep learning is slowed by the days or weeks it takes to train large models. The natural solution of using more hardware is limited by diminishing returns, and leads to inefficient use of additional resources. In this paper, we present a large batch, stochastic optimization algorithm that is both faster tha…
New method computes affine normal directions efficiently for sparse polynomials.
New method for efficient sketching of gradients and Hessians.
This paper proposes a family of online second order methods for possibly non-convex stochastic optimizations based on the theory of preconditioned stochastic gradient descent (PSGD), which can be regarded as an enhance stochastic Newton method with the ability to handle gradient noise and non-convexity simultaneously. …
New algorithms tackle complex multi-block optimization problems in machine learning.
How can we explain the predictions of a black-box model? In this paper, we use influence functions -- a classic technique from robust statistics -- to trace a model's prediction through the learning algorithm and back to its training data, thereby identifying training points most responsible for a given prediction. To …
Influence functions help study large language model generalization, revealing surprising decay patterns.
The question of how to incorporate curvature information in stochastic approximation methods is challenging. The direct application of classical quasi- Newton updating techniques for deterministic optimization leads to noisy curvature estimates that have harmful effects on the robustness of the iteration. In this paper…
We present a novel statistical inference framework for convex empirical risk minimization, using approximate stochastic Newton steps. The proposed algorithm is based on the notion of finite differences and allows the approximation of a Hessian-vector product from first-order information. In theory, our method efficient…
Score matching is a popular method for estimating unnormalized statistical models. However, it has been so far limited to simple, shallow models or low-dimensional data, due to the difficulty of computing the Hessian of log-density functions. We show this difficulty can be mitigated by projecting the scores onto random…
New preconditioners speed up SGD on Lie groups.
We explore inverse and quanto inverse crypto options, their pricing, and applications.
New algorithm finds approximate stationary points in non-convex optimization.
We consider the warped product manifold, , with Riemannian metric , where is a smooth closed Riemannian -manifold. We investigate what sufficient curvature condition is required of to ensure that a solution to the inverse mean cur…
We study inverse mean curvature flows of starshaped, mean convex hypersurfaces in warped product manifolds with a positive warping factor . If and , we show that these flows exist for all times, remain starshaped and mean convex. Plus the positivity of and …
A new unbiased Hessian estimator for expectation-based objectives.
The long-time existence and umbilicity estimates for compact, graphical solutions to expanding curvature flows are deduced in Riemannian warped products of a real interval with a compact fibre. Notably we do not assume the ambient manifold to be rotationally symmetric, nor the radial curvature to converge, nor a lower …
This paper tackles efficient optimization for nonlinear embeddings in similarity learning.
Given a convex cone in the \emph{prescribed} warped product, we consider hypersurfaces with boundary which are star-shaped with respect to the center of the cone and which meet the cone perpendicularly. If those hypersurfaces inside the cone evolve along the inverse mean curvature flow, then, by using the convexity of …
Abstracts a theorem for non-smooth maps in infinite dimensions.
Machine learning improves solving inverse problems and integrating data.
This paper is devoted to an inverse Steklov problem for a particular class of n-dimensional manifolds having the topology of a hollow sphere and equipped with a warped product metric. We prove that the knowledge of the Steklov spectrum determines uniquely the associated warping function up to a natural invariance.
EiGLasso speeds up sparse Kronecker-sum covariance estimation.
Discrete Green's functions are the inverses or pseudo-inverses of combinatorial Laplacians. We present compact formulas for discrete Green's functions, in terms of the eigensystems of corresponding Laplacians, for products of regular graphs with or without boundary. Explicit formulas are derived for the cycle, torus, a…
We consider inverse curvature flows in warped product manifolds, which are constrained subject to local terms of lower order, namely the radial coordinate and the generalized support function. Under various assumptions we prove longtime existence and smooth convergence to a coordinate slice. We apply this result to ded…
We study solutions to the inverse mean curvature flow which evolve by homotheties of a given submanifold with arbitrary dimension and codimension. We first show that the closed ones are necessarily spherical minimal immersions and so we reveal the strong rigidity of the Clifford torus in this setting. Mainly we focus o…
Tensor decomposition, a collection of factorization techniques for multidimensional arrays, are among the most general and powerful tools for scientific analysis. However, because of their increasing size, today's data sets require more complex tensor decomposition involving factorization with multiple matrices and dia…
The paper analyzes convergence rates of bilevel optimization algorithms and introduces a new stochastic algorithm.
The paper studies area-preserving and length-preserving inverse curvature flow for planar curves with singularities.
NG+ method improves deep learning efficiency and accuracy.
We analyze the Hessian spectra of large models up to 100B parameters.
New pruning method captures global correlations for efficient neural network inference.
In this paper we study Inverse Mean Curvature Flow (IMCF) on manifolds that are conformal to a warped product manifold. To this end, we show how the gradient conformal vector field in warped product manifolds is related to the conformal vector field on the conformal metric and use this to gain control of the flow in or…
A new method tackles bilevel optimization using Lanczos process for efficient hyper-gradient computation.