Gradient oversmoothing and expansion hinder deep GNN training, solved with normalization.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Gradient-enhanced GSA uses Poincaré chaos expansions for accurate sensitivity analysis.
The paper studies geometric properties of group equivariant operators and their Riemannian structure.
This study provides an explicit expansion of KL divergence's gradient flow in Fisher-Rao geometry.
Develops a new framework to analyze gradient flow regimes and derive explicit solutions.
Efficient method for high-dimensional American option pricing and hedging.
Iterative tilting fine-tunes diffusion models for reward-tilted distributions.
Study on neural networks' performance under different normalizations as N grows.
Gradient descent at edge of stability stabilizes implicitly, following projected gradient descent.
New method estimates SDE parameters efficiently using WCE and SGD.
We show that at the level of formal expansions, any compact Riemannian manifold is the sphere at infinity of an asymptotically conical gradient expanding Ricci soliton.
Proposes efficient model for continual learning that grows model over task-specific parameters.
We consider estimating a piecewise-constant image, or a gradient-sparse signal on a general graph, from noisy linear measurements. We propose and study an iterative algorithm to minimize a penalized least-squares objective, with a penalty given by the "l_0-norm" of the signal's discrete graph gradient. The method proce…
Projection pursuit model improves Gaussian process regression for high-dimensional data.
A new method reduces variance in training discrete latent variable models.
Secure Aggregation protocols allow a collection of mutually distrust parties, each holding a private value, to collaboratively compute the sum of those values without revealing the values themselves. We consider training a deep neural network in the Federated Learning model, using distributed stochastic gradient descen…
By adapting methods of \cite{AC} we prove a sharp estimate on the expansion modulus of the gradient of the log of the parabolic kernel to the Schördinger operator with convex potential, which improves an earlier work of Brascamp-Lieb. We also include alternate proofs to the improved log-concavity estimate, and to the f…
EigenVI uses orthogonal function expansions for efficient variational inference.
Sharp privacy bounds for sequential analysis of sensitive data.
New findings show neural network training loss follows a power law over time.
New neural network models for functional data.
We obtain expressions for the shear and the vorticity tensors of perfect-fluid spacetimes, in terms of the divergence of the Weyl tensor. For such spacetimes, we prove that if the gradient of the energy density is parallel to the velocity, then either the expansion rate is zero, or the vorticity vanishes. This statemen…
The choice of activation function can significantly influence the performance of neural networks. The lack of guiding principles for the selection of activation function is lamentable. We try to address this issue by introducing our variational neural networks, where the activation function is represented as a linear c…
Developing a differentially private deep learning algorithm is challenging, due to the difficulty in analyzing the sensitivity of objective functions that are typically used to train deep neural networks. Many existing methods resort to the stochastic gradient descent algorithm and apply a pre-defined sensitivity to th…
In this paper we introduce a family of stochastic gradient estimation techniques based of the perturbative expansion around the mean of the sampling distribution. We characterize the bias and variance of the resulting Taylor-corrected estimators using the Lagrange error formula. Furthermore, we introduce a family of va…
We study the small time asymptotics of the gradient and Hessian of the logarithm of the heat kernel at the cut locus, giving, in principle, complete expansions for both quantities. We relate the leading terms of the expansions to the structure of the cut locus, especially to conjugacy, and we provide a probabilistic in…
McKernel introduces a framework to use kernel approximates in the mini-batch setting with Stochastic Gradient Descent (SGD) as an alternative to Deep Learning. Based on Random Kitchen Sinks [Rahimi and Recht 2007], we provide a C++ library for Large-scale Machine Learning. It contains a CPU optimized implementation of …
Diffusion approximation provides weak approximation for stochastic gradient descent algorithms in a finite time horizon. In this paper, we introduce new tools motivated by the backward error analysis of numerical stochastic differential equations into the theoretical framework of diffusion approximation, extending the …
We consider the minimization of an objective function given access to unbiased estimates of its gradient through stochastic gradient descent (SGD) with constant step-size. While the detailed analysis was only performed for quadratic functions, we provide an explicit asymptotic expansion of the moments of the averaged S…
GUESS improves surrogate model accuracy with adaptive sampling.
Rescaling expansiveness proven for k*-expansive vector fields.
A large class of machine learning techniques requires the solution of optimization problems involving spectral functions of parametric matrices, e.g. log-determinant and nuclear norm. Unfortunately, computing the gradient of a spectral function is generally of cubic complexity, as such gradient descent methods are rath…
An expansion is developed for the Weil-Petersson Riemann curvature tensor in the thin region of the Teichmüller and moduli spaces. The tensor is evaluated on the gradients of geodesic-lengths for disjoint geodesics. A precise lower bound for sectional curvature in terms of the surface systole is presented. The curvatur…
The paper derives expansions for Green's operators and resolvents using Hadamard methods.
With a f-left-invariant Riemannian metric on a Lie group , we mean a Riemannian metric which is conformally equivalent to a left-invariant Riemannian metric, with the conformal factor . In this article, we study the geometry of such metrics and give a necessary and sufficient condition for an f-left-invariant Rie…
This paper presents a novel technique based on gradient boosting to train the final layers of a neural network (NN). Gradient boosting is an additive expansion algorithm in which a series of models are trained sequentially to approximate a given function. A neural network can also be seen as an additive expansion where…
Let be the Teichmüller space of marked genus , punctured Riemann surfaces with its bordification $\Tbar$ the {\em augmented Teichmüller space} of marked Riemann surfaces with nodes, \cite{Abdegn, Bersdeg}. Provided with the WP metric $\Tbar$ is a complete CAT(0) metric space, \cite{DW2, Wlcomp, Yam2…
Develops a martingale expansion for stochastic volatility models.
Paper develops a gradient-like proposal for discrete distributions without requiring natural differentiability.
Taylor expansions improve reinforcement learning policies.
Analytic torsion expansions for symmetric and complex homogeneous spaces.
Proved cyclotomic expansion for double twist knots' HOMFLY-PT invariants.
Study of hypersurfaces with specific expansion properties.
A new hypergraph expansion method treats vertices and hyperedges equally, improving node classification.
Paper introduces SGD for nonparametric additive models with optimal risk.
TEAM generates more powerful adversarial examples for DNNs.
Asymmetric expansion preserves convexity in hyperbolic geometry.
The paper calculates asymptotic expansions for specific types of oscillatory integrals.