The present paper develops a novel aggregated gradient approach for distributed machine learning that adaptively compresses the gradient communication. The key idea is to first quantize the computed gradients, and then skip less informative quantized gradient communications by reusing outdated gradients. Quantizing and…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper studies fundamental limits of communication in distributed learning.
LASG improves communication efficiency in distributed learning.
New method reduces communication costs in distributed nonconvex optimization.
This paper develops a communication-efficient algorithm to solve the stochastic optimization problem defined over a distributed network, aiming at reducing the burdensome communication in applications such as distributed machine learning.Different from the existing works based on quantization and sparsification, we int…
Optimal gradient quantization reduces communication costs in distributed deep learning.
A2SGD reduces distributed SGD communication to O(1) per worker.
This paper presents a new class of gradient methods for distributed machine learning that adaptively skip the gradient calculations to learn with reduced communication and computation. Simple rules are designed to detect slowly-varying gradients and, therefore, trigger the reuse of outdated gradients. The resultant gra…
New method reduces communication in deep learning training.
A new gradient quantization scheme improves communication efficiency in distributed training.
Paper proposes an adaptive gradient method for federated learning.
A new algorithm reduces communication in deep learning training.
A new hybrid-ordered SGD method reduces communication and complexity for non-convex optimization.
The high cost of communicating gradients is a major bottleneck for federated learning, as the bandwidth of the participating user devices is limited. Existing gradient compression algorithms are mainly designed for data centers with high-speed network and achieve per-iteration communication cost at…
New approach uses SPG for semantic communication without a known channel model.
This paper deals with distributed policy optimization in reinforcement learning, which involves a central controller and a group of learners. In particular, two typical settings encountered in several applications are considered: multi-agent reinforcement learning (RL) and parallel RL, where frequent information exchan…
GradSkip reduces local training steps for better communication efficiency.
DESTRESS optimizes decentralized nonconvex optimization with optimal IFO complexity and efficient communication.
Recently, researchers proposed various low-precision gradient compression, for efficient communication in large-scale distributed optimization. Based on these work, we try to reduce the communication complexity from a new direction. We pursue an ideal bijective mapping between two spaces of gradient distribution, so th…
Distributed Lion optimizes large model training by reducing communication costs.
Communication overhead is a major bottleneck hampering the scalability of distributed machine learning systems. Recently, there has been a surge of interest in using gradient compression to improve the communication efficiency of distributed neural network training. Using 1-bit quantization, signSGD with majority vote …
COMP-AMS optimizes distributed training with compressed gradients, achieving similar accuracy with less communication.
Flexible framework improves communication efficiency across various systems.
Large-scale distributed training of neural networks is often limited by network bandwidth, wherein the communication time overwhelms the local computation time. Motivated by the success of sketching methods in sub-linear/streaming algorithms, we introduce Sketched SGD, an algorithm for carrying out distributed SGD by c…
To reduce the long training time of large deep neural network (DNN) models, distributed synchronous stochastic gradient descent (S-SGD) is commonly used on a cluster of workers. However, the speedup brought by multiple workers is limited by the communication overhead. Two approaches, namely pipelining and gradient spar…
Adaptive batch sizes improve local gradient methods in distributed training.
Modern large scale machine learning applications require stochastic optimization algorithms to be implemented on distributed computational architectures. A key bottleneck is the communication overhead for exchanging information such as stochastic gradients among different workers. In this paper, to reduce the communica…
Modern distributed training of machine learning models suffers from high communication overhead for synchronizing stochastic gradients and model parameters. In this paper, to reduce the communication complexity, we propose \emph{double quantization}, a general scheme for quantizing both model parameters and gradients. …
With the rapid growth of data, distributed momentum stochastic gradient descent~(DMSGD) has been widely used in distributed learning, especially for training large-scale deep models. Due to the latency and limited bandwidth of the network, communication has become the bottleneck of distributed learning. Communication c…
With the increase in the amount of data and the expansion of model scale, distributed parallel training becomes an important and successful technique to address the optimization challenges. Nevertheless, although distributed stochastic gradient descent (SGD) algorithms can achieve a linear iteration speedup, they are l…
Paper tackles hyper-gradient estimation in decentralized FL over time-varying networks.
We study distributed optimization algorithms for minimizing the average of convex functions. The applications include empirical risk minimization problems in statistical machine learning where the datasets are large and have to be stored on different machines. We design a distributed stochastic variance reduced gradien…
New sparsification technique for SGD reduces communication costs.
Unified analysis of federated learning with compression for various data distributions.
As the size and complexity of models and datasets grow, so does the need for communication-efficient variants of stochastic gradient descent that can be deployed to perform parallel model training. One popular communication-compression method for data-parallel SGD is QSGD (Alistarh et al., 2017), which quantizes and en…
Large-scale distributed training requires significant communication bandwidth for gradient exchange that limits the scalability of multi-node training, and requires expensive high-bandwidth network infrastructure. The situation gets even worse with distributed training on mobile devices (federated learning), which suff…
AdaQuantFL reduces communication in federated learning by adaptively quantizing model updates.
A new method for distributed optimization reduces communication rounds without minibatches.
We focus on the commonly used synchronous Gradient Descent paradigm for large-scale distributed learning, for which there has been a growing interest to develop efficient and robust gradient aggregation strategies that overcome two key system bottlenecks: communication bandwidth and stragglers' delays. In particular, R…
Cyclic Data Parallelism reduces memory usage and balances gradient communications.
FedSKETCH and FedSKETCHGATE improve privacy and efficiency in federated learning.
New method improves FL efficiency by shuffling data, balancing privacy and accuracy.
Clapping reduces memory usage in distributed optimization by reusing data samples.
New techniques improve distributed training with compressed gradients.
A new approach for cooperative multi-agent reinforcement learning with limited communication, reducing the number of communication rounds.
A new method reduces communication costs in decentralized optimization.
New analysis shows Local SGD can achieve error scaling with only fixed number of communications.
One of the most significant bottleneck in training large scale machine learning models on parameter server (PS) is the communication overhead, because it needs to frequently exchange the model gradients between the workers and servers during the training iterations. Gradient quantization has been proposed as an effecti…