This paper bounds the Lipschitz constants of neural networks and their gradients.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New methods for convex optimization with locally Lipschitz gradient, achieving faster convergence.
We focus our attention on the notion of intrinsic Lipschitz graphs, inside a special class of metric spaces i.e. the Carnot groups. More precisely, we provide a characterization of locally intrinsic Lipschitz functions in Carnot groups of step 2 in terms of their intrinsic distributional gradients.
Introduces gradient decay in Softmax for better generalization.
New method improves optimization algorithms without Lipschitz smoothness.
We study distributed optimization algorithms for minimizing the average of \emph{heterogeneous} functions distributed across several machines with a focus on communication efficiency. In such settings, naively using the classical stochastic gradient descent (SGD) or its variants (e.g., SVRG) with a uniform sampling of …
Study Ricci-Deturck flow from rough metrics, proving short-time existence.
Adaptive methods improve gradient descent and proximal gradient for convex optimization.
HALO uses local Lipschitz constants to optimize functions efficiently.
The paper analyzes neural network dynamics after weights escape the origin.
Constructs a map with prescribed local Lipschitz constants on a subset of a manifold.
New bounds explain neural network generalization by considering local Lipschitz properties.
New method balances multivariate model fitting for mixed likelihoods.
Gradient flow of curve length on Sobolev metrics preserves convexity.
Maps and measures on surfaces link best Lipschitz and least gradient functions.
The mean curvature flow is the gradient flow of volume functionals on the space of submanifolds. We prove a fundamental regularity result of the mean curvature flow in this paper: a Lipschitz submanifold with small local Lipschitz norm becomes smooth instantly along the mean curvature flow. This generalizes the regular…
We compute the local Lipschitz constant of ReLU networks precisely.
GD converges in unstable regimes, even with oscillatory behavior.
New method for differentially private optimization with general Lipschitz conditions.
Extends Lipschitz functions while preserving local constants.
Let be a complete metric measure space, with a locally doubling measure, that supports a local weak -Poincaré inequality. By assuming a heat semigroup type curvature condition, we prove that Cheeger-harmonic functions are Lipschitz continuous on . Gradient estimates for Cheeger-harmonic func…
Optimistic method adapted for faster convex-concave min-max problems.
The curse of dimensionality affects neural network optimization, especially with smooth functions.
The Jacobian matrix (or the gradient for single-output networks) is directly related to many important properties of neural networks, such as the function landscape, stationary points, (local) Lipschitz constants and robustness to adversarial attacks. In this paper, we propose a recursive algorithm, RecurJac, to comput…
This paper analyzes convergence of large-scale Transformers with weight decay.
New Transformers maintain Lipschitz continuity for robustness.
We construct short retractions of a CAT(1) space to its small convex subsets. This construction provides an alternative geometric description of an analytic tool introduced by Wilfrid Kendall. Our construction uses a tractrix flow which can be defined as a gradient flow for a family of functions of certain type. In an …
In this paper, we are concerned with a non-asymptotic analysis of sampling algorithms used in nonconvex optimization. In particular, we obtain non-asymptotic estimates in Wasserstein-1 and Wasserstein-2 distances for a popular class of algorithms called Stochastic Gradient Langevin Dynamics (SGLD). In addition, the afo…
New method estimates Riemannian derivatives from noisy function evaluations.
Lipschitz continuity recently becomes popular in generative adversarial networks (GANs). It was observed that the Lipschitz regularized discriminator leads to improved training stability and sample quality. The mainstream implementations of Lipschitz continuity include gradient penalty and spectral normalization. In th…
Lipschitz constraints under L2 norm on deep neural networks are useful for provable adversarial robustness bounds, stable training, and Wasserstein distance estimation. While heuristic approaches such as the gradient penalty have seen much practical success, it is challenging to achieve similar practical performance wh…
We give sufficient conditions for a -local diffeomorphism between Fréchet spaces to be a global one. We extend the Clarke's theory of generalized gradients to the more general setting of Fréchet spaces. As a consequence, we define the Chang Palais-Smale condition for Lipschitz functions and show that a functio…
We model how Lipschitz continuity changes during neural network training.
ModHiFi identifies critical components for model modification without gradients or loss function.
In this paper we describe the notion of a weak lipschitzianity of a mapping on a stratification. We also distinguish a class of regularity conditions that are in some sense invariant under definable, locally Lipschitz and weakly bi-Lipschitz homeomorphisms. This class includes the Whitney (B) condition and the …
Training neural networks under a strict Lipschitz constraint is useful for provable adversarial robustness, generalization bounds, interpretable gradients, and Wasserstein distance estimation. By the composition property of Lipschitz functions, it suffices to ensure that each individual affine transformation or nonline…
We investigate the challenge of multi-output learning, where the goal is to learn a vector-valued function based on a supervised data set. This includes a range of important problems in Machine Learning including multi-target regression, multi-class classification and multi-label classification. We begin our analysis b…
Gradient descent converges linearly in finite-width networks with positive NTK and compatible conditions.
This paper analyzes the convergence of Federated Average under relaxed assumptions.
New shuffling methods improve convergence without Lipschitz smoothness.
In this paper, we study the convergence of generative adversarial networks (GANs) from the perspective of the informativeness of the gradient of the optimal discriminative function. We show that GANs without restriction on the discriminative function space commonly suffer from the problem that the gradient produced by …
Generative adversarial networks (GANs) are one of the most popular approaches when it comes to training generative models, among which variants of Wasserstein GANs are considered superior to the standard GAN formulation in terms of learning stability and sample quality. However, Wasserstein GANs require the critic to b…
A Lipschitz hypersurface is a hypersurface which locally is the graph of a Lipschitz function. A Lipschitz (or C^1) hypersurface is said to be Levi-flat if it is locally foliated by complex manifolds of complex dimension (n-1). We shall prove that there exist no Lipschitz Levi-flat hypersurfaces in CP^n with n >= 3. Ou…
The paper examines partial regularity of Lipschitz solutions to minimal surface system.
GraN-GAN normalizes gradients for better GAN performance.
Sharp Lipschitz bounds and gradient estimates for fully nonlinear parabolic equations.
New algorithms sample from log concave distributions without gradient Lipschitz continuity.
Wasserstein GAN(WGAN) is a model that minimizes the Wasserstein distance between a data distribution and sample distribution. Recent studies have proposed stabilizing the training process for the WGAN and implementing the Lipschitz constraint. In this study, we prove the local stability of optimizing the simple gradien…