The paper analyzes RLVR's training dynamics, proving convergence depends on aligning update direction with Gradient Gap.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Gradient steady Ricci solitons are natural generalizations of Ricci-flat manifolds. In this article, we prove a curvature gap theorem for gradient steady Ricci solitons with nonconstant potential functions; and a curvature gap theorem for Ricci-flat manifolds, removing the volume growth assumptions in known results.
Cloud computing is becoming increasingly popular as a platform for distributed training of deep neural networks. Synchronous stochastic gradient descent (SSGD) suffers from substantial slowdowns due to stragglers if the environment is non-dedicated, as is common in cloud computing. Asynchronous SGD (ASGD) methods are i…
This work analyzes the gap between off-policy and on-policy policy gradient methods and provides conditions to reduce this gap.
Estimates generalization gap for overparameterized models using Langevin approximation.
We study integral and pointwise bounds on the curvature of gradient shrinking Ricci solitons. As applications we discuss gap and compactness results for gradient shrinkers.
The paper proves various inequalities on gradient shrinking Ricci solitons.
New convergence rates for shuffling gradient methods without strong convexity.
This work is a part of ICLR Reproducibility Challenge 2019, we try to reproduce the results in the conference submission PADAM: Closing The Generalization Gap of Adaptive Gradient Methods In Training Deep Neural Networks. Adaptive gradient methods proposed in past demonstrate a degraded generalization performance than …
This note provides a novel, simple analysis of the method of conjugate gradients for the minimization of convex quadratic functions. In contrast with standard arguments, our proof is entirely self-contained and does not rely on the existence of Chebyshev polynomials. Another advantage of our development is that it clar…
New algorithm closes empirical gap in PFSGD performance.
New analysis improves understanding of bilevel optimization stability and generalization.
In this note, we prove an -energy gap result for Yang-Mills connections on a principal -bundle over a compact manifold without using Lojasiewicz-Simon gradient inequality (arXiv:1502.00668).
Optimal privacy-preserving algorithm for solving saddle point problems.
This paper bridges variational inference and Wasserstein gradient flows.
In this short note, using Günther's volume comparison theorem and Yokota's gap theorem on complete shrinking gradient Ricci solitons, we prove that for any complete shrinking gradient Ricci soliton with sectional curvature and for some uniform constant , there exists…
In this paper, we will prove a gap theorem for four-dimensional gradient shrinking soliton. More precisely, we will show that any complete four-dimensional gradient shrinking soliton with nonnegative and bounded Ricci curvature, satisfying a pinched Weyl curvature, either is flat, or everywhere f…
Gradient descent methods for deep ReLU networks achieve optimal generalization rates.
DARTS is a popular algorithm for neural architecture search (NAS). Despite its great advantage in search efficiency, DARTS often suffers weak stability, which reflects in the large variation among individual trials as well as the sensitivity to the hyper-parameters of the search process. This paper owes such instabilit…
Adaptive gradient methods, which adopt historical gradient information to automatically adjust the learning rate, despite the nice property of fast convergence, have been observed to generalize worse than stochastic gradient descent (SGD) with momentum in training deep neural networks. This leaves how to close the gene…
Tax evasion is the illegal evasion of taxes by individuals, corporations, and trusts. The revenue loss from tax avoidance can undermine the effectiveness and equity of the government policies. A standard measure of tax evasion is the tax gap, that can be estimated as the difference between the total amounts of tax theo…
AdAdaGrad optimizes batch sizes for deep learning models, reducing the generalization gap.
Random feature model shows slow self-correction of generalization gap.
We derive a sharp lower bound for the scalar curvature of non-flat and non-compact expanding gradient Ricci soliton provided that the scalar curvature is non-negative and the potential function is proper. We also give an upper bound for the scalar curvature of noncompact expander when the Ricci curvature is nonpositive…
In this paper, we prove that a gradient shrinking compact Kähler-Ricci soliton cannot have too large Ricci curvature unless it is Kähler-Einstein.
The study examines how averaging data improves model performance.
Stochastic gradient descent is the method of choice for large scale optimization of machine learning objective functions. Yet, its performance is greatly variable and heavily depends on the choice of the stepsizes. This has motivated a large body of research on adaptive stepsizes. However, there is currently a gap in o…
In this paper, we show that any ancient solution to the Ricci flow with the reduced volume whose asymptotic limit is sufficiently close to that of the Gaussian soliton is isometric to the Euclidean space for all time. This is a generalization of Anderson's result for Ricci-flat manifolds. As a corollary, a gap theorem …
Background: Deep learning models are typically trained using stochastic gradient descent or one of its variants. These methods update the weights using their gradient, estimated from a small fraction of the training data. It has been observed that when using large batch sizes there is a persistent degradation in genera…
Epoch-GDA achieves optimal convergence rate for SCSC min-max problems.
GSP improves global average pooling for deep metric learning by learning weights and selecting semantic entities.
Study on policy gradient for stochastic bandits using diffusion approximation.
Policy gradients methods apply to complex, poorly understood, control problems by performing stochastic gradient descent over a parameterized class of polices. Unfortunately, even for simple control problems solvable by standard dynamic programming techniques, policy gradient algorithms face non-convex optimization pro…
SGD reduces test error by decorrelating updates.
We study curvature dimension inequalities for the sub-Laplacian on contact Riemannian manifolds. This new curvature dimension condition is then used to obtain: 1) Geometric conditions ensuring the compactness of the underlying manifold (Bonnet-Myers type results); 2) Volume estimates of metric balls; 3) Gradient bounds…
Effective Gram matrix predicts deep network generalization.
Distributed learning with random features and gradient descent improves performance and reduces memory usage.
Paper analyzes complexity of solving nonconvex-strongly-concave problems.
Recently, there is a growing interest in the study of median-based algorithms for distributed non-convex optimization. Two prominent such algorithms include signSGD with majority vote, an effective approach for communication reduction via 1-bit compression on the local gradients, and medianSGD, an algorithm recently pr…
A new algorithm COVA-FC improves subgroup-fair clustering efficiency.
We provide tight upper and lower bounds on the complexity of minimizing the average of convex functions using gradient and prox oracles of the component functions. We show a significant gap between the complexity of deterministic vs randomized optimization. For smooth functions, we show that accelerated gradient de…
The natural gradient of ELBO vanishes in unconstrained optimization, simplifying learning.
Adversarial training is a training scheme designed to counter adversarial attacks by augmenting the training dataset with adversarial examples. Surprisingly, several studies have observed that loss gradients from adversarially trained DNNs are visually more interpretable than those from standard DNNs. Although this phe…
Paper explores PG for MCR, finding suboptimal policies but providing bounds.
In this sequel to [arXiv:1412.4114], we prove an energy gap result for Yang-Mills connections on principal -bundles, , over arbitrary, closed, Riemannian, smooth manifolds of dimension . We apply our version of the Lojasiewicz-Simon gradient inequality [arXiv:1409.1525, arXiv:1510.03815] to rem…
We propose and analyze a new type of stochastic first order method: gradient descent with compressed iterates (GDCI). GDCI in each iteration first compresses the current iterate using a lossy randomized compression technique, and subsequently takes a gradient step. This method is a distillation of a key ingredient in t…
This study tightens bounds on how GD and SGD generalize in smooth convex optimization problems.
Paper generalizes Bakry-Émery calculus for curvature and applies to Markov chains.