Proposes SME for ASGD, revealing dynamics and optimal mini-batching.
problem Understanding and optimizing ASGD algorithms.
method Develops SME for ASGD, proving convergence and solving optimal control problem.
result ASGD converges to SME in continuous time limit and predicts ASGD trajectories.
ASGD outperforms SGD in overparameterized linear regression, especially in subspaces of small eigenvalues.
problem Generalization of ASGD for overparameterized linear regression.
method Established instance-dependent excess risk bound for ASGD in each eigen-subspace of the data covariance matrix.
result ASGD outperforms SGD in subspaces of small eigenvalues, exhibiting faster decay of bias error.
Rescaled ASGD optimizes distributed learning under heterogeneous data.
problem Vanilla ASGD biases towards a frequency-weighted average of local objectives.
method Rescale worker stepsizes by their computation times.
result Rescaled ASGD converges to the correct global objective in fixed-computation model.
The paper provides statistical guarantees for SGD and ASGD in high-dimensional settings.
problem Theoretical understanding of SGD and ASGD in high-dimensional settings.
method Transfer of tools from high-dimensional time series to online learning, using coupling techniques.
result Established geometric-moment contraction and q-th moment convergence of SGD and ASGD. K-AVG improves convergence for nonconvex optimization problems.
problem Improving the convergence of ASGD for nonconvex optimization.
method K-step averaging stochastic gradient descent (K-AVG) for nonconvex objectives.
result K-AVG converges faster and achieves better accuracies than ASGD.
Ringmaster ASGD improves Asynchronous SGD's efficiency under varying worker times.
problem Suboptimal performance of Asynchronous SGD under heterogeneous worker computation times.
method Ringmaster ASGD, a novel Asynchronous SGD method with optimal time complexity.
result Ringmaster ASGD achieves optimal time complexity under arbitrary worker heterogeneity.
This paper analyzes ASGD using SDEs for a more intuitive convergence rate.
problem Theoretical analysis of ASGD is limited by discrete methods and complex proofs.
method Continuous approximation of ASGD using SDDEs and convergence rate analysis methods.
result Continuous view provides better convergence rates and insights into ASGD.
Stochastic momentum methods trade compute efficiency for serial runtime.
problem Stochastic momentum methods trade compute efficiency for serial runtime.
method Stochastic HB and ASGD for consistent linear regression with Gaussian covariates.
result HB preserves SGD-level CE over a larger batch-size window, allowing larger batches to reduce serial runtime until HB reaches its deterministic accelerated scale.
Ringleader ASGD optimizes SGD for diverse edge devices with varying data and computation speeds.
problem Scalable distributed optimization with heterogeneous devices and data.
method Ringleader ASGD, an asynchronous SGD algorithm.
result Achieves optimal time complexity under data heterogeneity and arbitrary computation speeds.
The paper analyzes SGD with dropout regularization in linear models, proving asymptotic properties and providing inference tools.
problem Analyzing the behavior of SGD with dropout regularization in linear models.
method Establishing geometric-moment contraction (GMC) and proving quenched central limit theorems (CLT).
result The existence of a unique stationary distribution and asymptotic normality results for SGD with dropout.
Ringmaster LMO accelerates training in distributed systems by asynchronously updating neural networks.
problem Asynchronous training in distributed systems where workers compute gradients at different speeds.
method Introduces an asynchronous LMO-based momentum method for unconstrained stochastic nonconvex optimization.
result Establishes convergence guarantees and time complexity bounds for asynchronous LMO-based updates.
New research shows existing momentum schemes often fail for stochastic optimization problems.
problem Existing momentum schemes like HB and NAG fail to outperform SGD in certain stochastic optimization problems.
method The study provides a new algorithm, ASGD, which improves over existing schemes on specific problem instances.
result ASGD significantly outperforms HB, NAG, and SGD on certain stochastic optimization problems.
Paper proposes an online estimator for covariance matrix of SGD iterates.
problem Quantifying variability and randomness of SGD-based estimates in online learning.
method Proposes a fully online estimator for covariance matrix of ASGD using SGD iterates.
result Establishes consistency of the online estimator and shows comparable convergence rate to offline methods.
Adapts SGD to noise and problem specifics for faster convergence.
problem Minimizing smooth, strongly-convex functions with varying noise and problem constants.
method Adaptive SGD with exponentially decreasing step-sizes, Nesterov acceleration, and stochastic line-search.
result Achieves near-optimal convergence rates without knowing noise or problem specifics.
New approach splits deep models for parallel training.
problem Training deep models requires significant computing resources.
method Split the model into parts and train them separately.
result High-performance model-parallel training achieved.
New asynchronous SGD algorithms achieve optimal performance in distributed learning.
problem Asynchronous training introduces staleness, complicating optimization analysis.
method Developed rigorous framework for asynchronous first order stochastic optimization.
result Asynchronous SGD can achieve optimal time complexity, matching synchronous methods.
Proposes DC-S3GD for efficient large-scale decentralized neural network training.
problem Training large-scale decentralized neural networks efficiently.
method Decentralized stale-synchronous version of DC-ASGD with gradient correction.
result Achieves state-of-the-art results in training Convolutional Neural Networks.
GA method reduces gradient staleness in cloud computing.
problem Gradient staleness in asynchronous SGD methods.
method Gap-Aware (GA) method that penalizes stale gradients linearly to the Gap.
result GA outperforms existing methods in final test accuracy.