A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
In this paper we introduce a novel method of gradient normalization and decay with respect to depth. Our method leverages the simple concept of normalizing all gradients in a deep neural network, and then decaying said gradients with respect to their depth in the network. Our proposed normalization and decay techniques…
In this paper we study potential function of gradient steady Ricci solitons. We prove that infimum of potential function decays linearly; in particular, potential function of rectifiable gradient steady Ricci solitons decays linearly. As a consequence, we show that a gradient steady Ricci soliton with bounded potential…
Unified framework for analyzing gradient flows of measures with exponential decay of entropy.
problem Analyzing exponential decay of entropy functionals in gradient flows of measures.
method Characterization of global exponential decay behaviors using Hellinger-Kantorovich geometry, shape-mass decomposition, and Polyak-Łojasiewicz-type inequalities.
result Unified theoretical framework for gradient flows with complete analysis of exponential decay behaviors.
In this paper, we give a description for steady Ricci solitons with a linear decay of sectional curvature. In particular, we classify all 3-dimensional steady Ricci solitons and 4-dimensional κ-noncollpased steady Ricci solitons with nonnegative sectional curvature under the linear curvature decay.
Regularization in the optimization of deep neural networks is often critical to avoid undesirable over-fitting leading to better generalization of model. One of the most popular regularization algorithms is to impose L-2 penalty on the model parameters resulting in the decay of parameters, called weight-decay, and the …
We study stability of non-compact gradient Kaehler-Ricci flow solitons with positive holomorphic bisectional curvature. Our main result is that any compactly supported perturbation and appropriately decaying perturbations of the Kaehler potential of the soliton will converge to the original soliton under Kaehler-Ricci …
In this paper, we derive certain curvature estimates for 4-dimensional gradient steady Ricci solitons either with positive Ricci curvature or with scalar curvature decay.
In this paper, we study the gradient Ricci soliton equation on a complete Riemannian manifold. We show that under a natural decay condition on the Ricci curvature, the Ricci soliton is Ricci-flat and ALE.
We first investigate the asymptotics of conical expanding gradient Ricci solitons by proving sharp decay rates to the asymptotic cone both in the generic and the asymptotically Ricci flat case. We then establish a compactness theorem concerning nonnegatively curved expanding gradient Ricci solitons.
We show that the only complete shrinking gradient Ricci solitons with vanishing Weyl tensor are quotients of the standard ones. This gives a new proof of the Hamilton-Ivey-Perel'man classification of 3-dimensional shrinking gradient solitons. We also prove a classification for expanding gradient Ricci solitons with con…
A long-standing obstacle to progress in deep learning is the problem of vanishing and exploding gradients. Although, the problem has largely been overcome via carefully constructed initializations and batch normalization, architectures incorporating skip-connections such as highway and resnets perform much better than …
Minimax optimal convergence rates for classes of stochastic convex optimization problems are well characterized, where the majority of results utilize iterate averaged stochastic gradient descent (SGD) with polynomially decaying step sizes. In contrast, SGD's final iterate behavior has received much less attention desp…
SGD and weight decay encourage neural networks to learn low-rank weight matrices.
problem The bias of SGD towards low-rank weight matrices in neural networks.
method The study investigates the effect of SGD and weight decay on the rank of weight matrices in neural networks, both theoretically and empirically.
result Training with SGD and weight decay induces a bias towards rank minimization in weight matrices, which becomes more pronounced with smaller batch sizes and stronger weight decay.
We develop heat kernel and Green's function estimates for manifolds with positive bottom spectrum. The results are then used to establish existence and sharp estimates of the solution to the Poisson equation on such manifolds with Ricci curvature bounded below. As an application, we show that the curvature of a steady …
A recent breakthrough in deep learning theory shows that the training of over-parameterized deep neural networks can be characterized by a kernel function called \textit{neural tangent kernel} (NTK). However, it is known that this type of results does not perfectly match the practice, as NTK-based analysis requires the…
The role of L2 regularization, in the specific case of deep neural networks rather than more traditional machine learning models, is still not fully elucidated. We hypothesize that this complex interplay is due to the combination of overparameterization and high dimensional phenomena that take place during training …
In this paper, we study the minimax optimization problem in the smooth and strongly convex-strongly concave setting when we have access to noisy estimates of gradients. In particular, we first analyze the stochastic Gradient Descent Ascent (GDA) method with constant stepsize, and show that it converges to a neighborhoo…
We study minimal graphic functions on complete Riemannian manifolds $\Si$ with non-negative Ricci curvature, Euclidean volume growth and quadratic curvature decay. We derive global bounds for the gradients for minimal graphic functions of linear growth only on one side. Then we can obtain a Liouville type theorem with …
New method trains neural networks with threshold activation functions efficiently.
problem Training neural networks with threshold activation functions is challenging due to zero gradients.
method We study weight decay regularized training problems of deep neural networks with threshold activations, showing they can be formulated as convex optimization problems.
result Regularized deep threshold network training problems can be formulated as standard convex optimization problems, paralleling the LASSO method.