Adaptive batch sizes improve local gradient methods in distributed training.
problem Communication bottlenecks in distributed deep learning.
method Adaptive batch size strategies for local gradient methods.
result Adaptive batch sizes reduce minibatch gradient variance and improve training efficiency.
It is shown that locally conformally flat Lorentzian gradient Ricci solitons are locally isometric to a Robertson-Walker warped product, if the gradient of the potential function is non null, and to a plane wave, if the gradient of the potential function is null. The latter gradient Ricci solitons are necessarily stead…
Paper proposes faster method to find local minima in nonconvex optimization.
problem Escaping saddle points and finding local minima in nonconvex optimization.
method LENA (Last stEp shriNkAge) framework for faster perturbed stochastic gradient methods.
result LENA finds (ε,εH)-approximate local minima within ildeO(ε−3+εH−6) evaluations. Proves convergence of gradient Ricci shrinkers with uniform bounds.
problem Compactness and energy concentration in gradient Ricci shrinkers.
method Bubble-tree convergence and local energy analysis.
result No energy concentrates in neck regions, leading to a local diffeomorphism finiteness theorem.
Adaptive methods improve gradient descent and proximal gradient for convex optimization.
problem Improving efficiency of gradient descent and proximal gradient methods.
method Adaptive versions of GD and ProxGD using local curvature information.
result Proved convergence with local Lipschitz gradient assumptions.
Local Gradient Descent with local steps converges to the centralized model in the interpolation regime.
problem Understanding the implicit bias of Local Gradient Descent in the interpolation regime.
method Analyzing the implicit bias of Local Gradient Descent for classification tasks with linearly separable data.
result The aggregated global model from Local-GD converges exactly to the centralized model in the interpolation regime.
Gradients help find global optima in complex functions.
problem Finding global optima in functions with many local minima.
method A principle for generating search directions from non-local quadratic approximants based on gradients.
result The proposed algorithm and CMA-ES perform better than random reinitialized BFGS.
Local GD proves effective for heterogeneous data in federated learning.
problem Minimizing functions from private, heterogeneous data in federated learning.
method Local gradient descent for smooth, convex functions.
result Communication complexity similar to gradient descent in low accuracy regime.
Stochastic gradient methods are dominant in nonconvex optimization especially for deep models but have low asymptotical convergence due to the fixed smoothness. To address this problem, we propose a simple yet effective method for improving stochastic gradient methods named predictive local smoothness (PLS). First, we …
Adapts Hölder smoothness with normalized gradients.
problem Improving smoothness adaptation methods.
method Black-box adaptation of Levy's method using normalized gradients.
result Bound depends on local Hölder smoothness.
Gradient-based methods find saddle points, not critical points, in neural networks.
problem Gradient-based optimization methods converge to saddle points rather than critical points in deep neural networks.
method Critical point-finding methods used to analyze neural network losses.
result Gradient-based methods often converge to or pass through gradient-flat regions, where gradient norm has a stationary point.
Optimize black-box simulators with local generative models.
problem Optimizing non-differentiable, stochastic simulators with intractable likelihoods.
method Differentiable local surrogate models based on deep generative models.
result Local surrogates enable gradient-based optimization, faster than baseline methods.
In this paper, we classify n-dimensional (n>2) complete noncompact locally conformally flat gradient steady solitons. In particular, we prove that a complete noncompact non-flat conformally flat gradient steady Ricci soliton is, up to scaling, the Bryant soliton.
In this paper, we first apply an integral identity on Ricci solitons to prove that closed locally conformally flat gradient Ricci solitons are of constant sectional curvature. We then generalize this integral identity to complete noncompact gradient shrinking Ricci solitons, under the conditions that the Ricci curvatur…
This paper concerns local gradient estimates to solutions of general conformally invariant fully nonlinear elliptic equations of second order.
Locally Accelerated Conditional Gradients improve convergence rates for smooth convex optimization problems.
problem Achieving optimal convergence rates for smooth convex optimization problems over polytopes.
method Locally Accelerated Conditional Gradients, coupling accelerated steps with conditional gradient steps.
result Achieves optimal accelerated local convergence for smooth strongly convex problems.
Forward gradients improve neural network training without backpropagation issues.
problem Training neural networks without backpropagation's locking and memorization problems.
method Using directional derivatives in forward differentiation mode, with biased guesses based on feedback from small auxiliary networks.
result Using gradients from a local loss as a candidate direction improves Forward Gradient methods.
Gradient method converges locally linearly for overparameterized Gaussian mixtures.
problem Learning Gaussian mixtures under overparameterization.
method Gradient-based method alternating short descent steps and long Polyak steps.
result Gradient method converges locally linearly to minimizers.
This study proves the local existence of a symplectic gradient flow on a flat torus.
problem Proving the local existence of a symplectic gradient flow on a flat torus.
method Using a moment map and a DeTurck trick to make the flow strictly parabolic and showing local existence and regularity.
result The group of symplectomorphisms of the real four-dimensional torus is locally contractible.
We describe the local structure of self-dual gradient Ricci solitons in neutral signature. If the Ricci soliton is non-isotropic then it is locally conformally flat and locally isometric to a warped product of the form I×φN(c), where N(c) is a space of constant curvature. If the Ricci soliton is isotro…
New techniques improve distributed training with compressed gradients.
problem Gradient mismatch problem in local error feedback.
method Step-ahead error feedback and error averaging techniques.
result Our methods handle gradient mismatch and train faster than full-precision training.
New definition of regret for nonconvex online learning models.
problem Intractability of standard regret measures for nonconvex models.
method Introduced a local gradient based regret definition.
result Our definition provides more interpretable bounds for forecasting.
Gradient descent learns useful features even in the NTK regime.
problem The ability of neural networks to learn useful features.
method Local convergence analysis of gradient descent with regularization.
result Gradient descent can capture ground-truth directions for feature learning even after the loss threshold is reached.
ABHT boosts regression by filtering regions with different smoothness.
problem Improving regression performance through local adaptivity.
method Gradient boosting with adaptive histogram transform.
result ABHT converges faster than PEHT in Hölder continuous spaces.
Simple rules ensure gradient descent adapts to local geometry, converging for convex and nonconvex problems.
problem Minimizing convex and nonconvex functions efficiently.
method Two rules: don't increase stepsize too fast and don't overstep local curvature.
result Method converges for convex and nonconvex problems, even with infinite global smoothness.
The local structure of half conformally flat gradient Ricci almost solitons is investigated, showing that they are locally conformally flat in a neighborhood of any point where the gradient of the potential function is non-null. In opposition, if the gradient of the potential function is null, then the soliton is a ste…
In this paper, we study two kind of L^2 norm preserved non-local heat flows on closed manifolds. We first study the global existence, stability and asymptotic behavior to such non-local heat flows. Next we give the gradient estimates of positive solutions to these heat flows.
Local constancy of index for certain gradient mappings proved.
problem Proving the local constancy of the index for specific gradient mappings.
method Using a more general theorem for quasiregular gradient mappings, deducing the result from the Hessian's properties.
result The index is locally constant for C1,1 functions with uniformly positive determinant Hessian almost everywhere. Estimates Kähler metrics with noncollapsing volume under complex Monge-Ampère constraints.
problem Volume noncollapsing for Kähler metrics induced by complex Monge-Ampère equations.
method Proves local volume noncollapsing estimate with Ricci curvature lower bound.
result Establishes diameter and gradient estimates for Kähler metrics.
We provide the classification of locally conformally flat gradient Yamabe solitons with positive sectional curvature. We first show that locally conformally flat gradient Yamabe solitons with positive sectional curvature have to be rotationally symmetric and then give the classification and asymptotic behavior of all r…
Paper proposes FR algorithm to solve minimax optimization locally.
problem Gradient descent fails to find local minimax in minimax optimization.
method Follow-the-Ridge (FR) algorithm, addressing rotational behavior of gradient dynamics.
result FR algorithm provably converges to local minimax.
Paper proposes a federated learning method for quantile inference with local differential privacy.
problem Federated learning of quantile inference under local differential privacy constraints.
method Local stochastic gradient descent with randomized mechanism for privacy and efficiency.
result Asymptotic normality and functional central limit theorem for the proposed estimator.
Natural gradient simplification for deep learning networks.
problem Efficiency in training deep Bayesian networks.
method Analysis of two geometries of Fisher information matrix and development of a method to simplify natural gradient for the second geometry.
result A method to simplify natural gradient for deep networks using an auxiliary recognition model.
SAVO actor improves reinforcement learning by avoiding local optima in complex Q-functions.
problem Gradient ascent in complex Q-functions leads to suboptimal solutions.
method SAVO actor generates multiple action proposals and truncates poor local optima.
result SAVO actor finds optimal actions more frequently and outperforms other architectures.
Gradient bounds and Liouville theorems for quasi-linear equations on manifolds with nonnegative Ricci curvature.
problem Establishing bounds and theorems for solutions to quasi-linear elliptic equations on compact manifolds with nonnegative Ricci curvature.
method Gradient bounds, Liouville-type theorems, local splitting theorem, Harnack-type inequality, ABP estimate.
result Gradient bounds and Liouville-type theorems for solutions to quasi-linear equations on compact manifolds with nonnegative Ricci curvature.
New k-step policy gradient method avoids local optima in restricted policy classes.
problem Suboptimal local optima in policy gradient methods for restricted policy classes.
method Proposes a k-step policy gradient method to escape myopic local optima. result The method converges to near optimal solutions exponentially close to the optimal deterministic policy.
This paper analyzes saddle points and minimax points in non-convex smooth games.
problem Understanding local optimal points in non-convex smooth games.
method Comprehensive analysis of local minimax points, including their optimality conditions and stability.
result Local saddle points are uniformly local minimax points under mild continuity assumptions.
Qsparse-local-SGD reduces communication in large-scale learning models.
problem Communication bottleneck in distributed optimization of large-scale models.
method Combines sparsification, quantization, and local computation with error compensation.
result Converges at the same rate as vanilla distributed SGD for many sparsifiers and quantizers.
GradSkip reduces local training steps for better communication efficiency.
problem High communication costs in distributed optimization.
method GradSkip redesigns ProxSkip to allow clients with less important data to take fewer local training steps.
result GradSkip converges linearly with reduced local training steps and same accelerated communication complexity.
Paper proposes MCTSPO for better reinforcement learning policy optimization.
problem Local optima and saddle points in gradient-based methods and poor initialization in gradient-free methods.
method Monte-Carlo tree search combined with gradient-free optimization.
result Improved performance on reinforcement learning tasks with deceptive or sparse reward functions.
Paper analyzes SGLD for nonconvex optimization with local conditions.
problem Analyzing sampling algorithms for nonconvex optimization.
method Non-asymptotic estimates for SGLD under local conditions.
result Establishes error bounds for expected excess risk.
This paper evaluates and compares gradient leakage attacks in federated learning.
problem Gradient leakage attacks compromise client privacy in federated learning.
method Formal and experimental analysis of gradient leakage attacks, evaluation of attack effectiveness and cost.
result Gradient leakage attacks can reconstruct private local training data from shared parameter updates.
Improves reinforcement learning by combining off-policy data and exploration.
problem Data inefficiency and local optima in policy gradient methods.
method Combines off-policy data reuse, exploration, and deterministic policies with stochastic optimization.
result Successfully learns solutions using fewer interactions than standard methods.
Proves an analytical analogue of Morse's lemma for gradient fields near critical points.
problem Understanding the behavior of gradient fields near critical points of Morse functions.
method Proves an analytical analogue of Morse's lemma showing unique linear vector fields.
result Shows that gradient fields near critical points have a natural standard form.
In this paper, we prove the local gradient estimate for harmonic functions on complete, noncompact Finsler measure spaces under the condition that the weighted Ricci curvature has a lower bound. As applications, we obtain Liouville type theorem on Finsler manifolds with nonnegative Ricci curvature.
Study on gradient ρ-Einstein solitons with radially nonnegative Bach tensor.
problem Characterizing gradient ρ-Einstein solitons with specific tensor properties.
method Analyzing the properties of Bach tensor and using local warping to classify solitons.
result Gradient ρ-Einstein solitons with radially nonnegative Bach tensor are locally warped products of an interval and an Einstein manifold.
Introduces gradient decay in Softmax for better generalization.
problem Improving generalization performance in neural networks.
method Gradient decay hyperparameter in Softmax for varying gradient rates based on probability.
result Gradient decay rate affects generalization performance and can be tuned for better optimization.
In this paper, we investigate some new local Aronson-Bénilan type gradient estimates for positive solutions of the porous medium equation ut=Δum, under Ricci flow. As application, the related Harnack inequalities are derived. Our results generalize known results. These results in the paper can be regard as …