Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

101201302402 · Jun 202019922001200920182026
48 results for asynchronous parameter server

Paper introduces an asynchronous optimization algorithm for parameter servers.

problem Solving optimization problems with asynchrony and general regularizers.
method Asynchronous incremental aggregated gradient algorithm in a parameter server framework.
result Established linear convergence rate and step-size choices for strong convex data loss.

AD-PSGD is an asynchronous decentralized parallel SGD that converges as fast as AllReduce-SGD but is much faster in a heterogeneous environment.

problem Designing an efficient and robust asynchronous decentralized parallel SGD algorithm.
method Proposes AD-PSGD, an asynchronous decentralized parallel SGD algorithm.
result AD-PSGD converges at the optimal O(1/K)O(1/\sqrt{K}) rate and has linear speedup w.r.t. number of workers.

DSCOVR improves distributed optimization for big data with less communication and synchronization.

problem Efficiently optimizing large linear models with convex loss functions over distributed systems.
method Randomized primal-dual block coordinate algorithms with doubly stochastic coordinate optimization and variance reduction.
result DSCOVR algorithms require less overall computation and communication compared to other first-order distributed algorithms.

Enhanced federated learning reduces communication costs and improves model accuracy.

problem Reducing communication costs in federated learning.
method Asynchronous model update and temporally weighted aggregation.
result The proposed algorithm outperforms baseline in terms of communication cost and model accuracy.

This work tackles resource allocation in asynchronous and stochastic systems.

problem Distributed resource allocation in asynchronous and stochastic settings.
method Approximate stochastic primal-dual approach with asynchronous updates.
result The Asynchronous stochastic Primal-Dual (Asyn-PD) algorithm converges to the saddle point solution at a rate of O(1/t)O(1/t).

New algorithm reduces distributed optimization time with stochastic delays.

problem Optimizing distributed data with stochastic delays.
method Developed ADSAGA, a variant of SAGA for distributed-data settings with stochastic delays.
result ADSAGA converges in $ ilde{O}\left(\left(n + \sqrt{m}κ ight)\log(1/ε) ight)$ iterations under mean delay mm.

Asynchronous distributed stochastic gradient descent methods have trouble converging because of stale gradients. A gradient update sent to a parameter server by a client is stale if the parameters used to calculate that gradient have since been updated on the server. Approaches have been proposed to circumvent this pro…

2016-01-15abs ↗pdf ↗

A(DP)2^2SGD improves federated learning privacy and efficiency.

problem Privacy and efficiency in federated learning with asynchronous decentralized parallel SGD.
method Differentially private asynchronous decentralized parallel SGD (A(DP)2^2SGD) using R{é}nyi differential privacy.
result Achieves optimal convergence rate and comparable model accuracy to SSGD but faster.

In distributed ML applications, shared parameters are usually replicated among computing nodes to minimize network overhead. Therefore, proper consistency model must be carefully chosen to ensure algorithm's correctness and provide high throughput. Existing consistency models used in general-purpose databases and moder…

2013-12-30abs ↗pdf ↗

A simple algorithm reduces federated contextual linear bandits' regret efficiently.

problem Solving federated contextual linear bandits with asynchronous agents.
method Proposed a simple algorithm exttt{FedLinUCB} based on optimism principle.
result Proved exttt{FedLinUCB} has bounded regret ildeO(dm=1MTm) ilde{O}(d\sqrt{\sum_{m=1}^M T_m}) and communication complexity ildeO(dM2) ilde{O}(dM^2).

Triadic-OCD detects changes in data streams robustly and optimally, even in asynchronous settings.

problem Online change detection in data streams with practical constraints.
method Triadic-OCD framework for asynchronous online change detection with provable robustness, optimality, and convergence.
result The proposed triadic-OCD algorithm achieves optimal performance and convergence in asynchronous settings.

New algorithm reduces federated learning rounds and improves privacy.

problem Inefficient synchronous federated learning causing scalability issues.
method Asynchronous federated learning with reduced communication and differential privacy via Gaussian noise.
result The algorithm reduces waiting times and network communication, making federated learning more scalable and private.

AsyB-ProxSGD parallelizes model updates and stochastic gradient descent for large models and data.

problem Efficiently training large models and handling large datasets in parallel.
method AsyB-ProxSGD: model parallel proximal stochastic gradient algorithm for asynchronous systems.
result Achieves linear speedup with O(K1/4)O(K^{1/4}) number of workers for nonconvex problems.

We study the problem of stochastic optimization for deep learning in the parallel computing environment under communication constraints. A new algorithm is proposed in this setting where the communication and coordination of work among concurrent processes (local workers), is based on an elastic force which links the p…

2014-12-20abs ↗pdf ↗

Unified analysis of asynchronous-SGD algorithms for distributed learning.

problem Analyzing asynchronous-SGD in heterogeneous settings with varying speeds and data distributions.
method Unified convergence theory for non-convex smooth functions, including pure asynchronous SGD and its modifications.
result Unified convergence rates for various asynchronous algorithms, including novel methods.

Lapse improves parameter servers by dynamically allocating parameters, achieving near-linear scaling.

problem Efficiently managing distributed training with reduced communication overhead.
method Integrate dynamic parameter allocation into parameter servers, proposing Lapse.
result Lapse provides near-linear scaling and can be orders of magnitude faster than existing parameter servers.

Proposes DC-S3GD for efficient large-scale decentralized neural network training.

problem Training large-scale decentralized neural networks efficiently.
method Decentralized stale-synchronous version of DC-ASGD with gradient correction.
result Achieves state-of-the-art results in training Convolutional Neural Networks.

DSSP improves deep learning training speed by dynamically adjusting staleness thresholds.

problem Time-consuming deep learning training on large datasets.
method Dynamic Stale Synchronous Parallel (DSSP) framework that adapts staleness threshold at runtime.
result DSSP converges faster and achieves higher accuracy than other paradigms.

Paper proposes double quantization to reduce communication in distributed machine learning.

problem High communication overhead in synchronizing stochastic gradients and model parameters in distributed training.
method Proposes double quantization for model parameters and gradients, and three communication-efficient algorithms.
result Established performance guarantees and demonstrated effective bit reduction without performance degradation.

Asynchronous framework improves distributed learning performance.

problem Heterogeneous computing machines hinder synchronous learning strategies.
method Asynchronous distributed framework with parameter exchanges.
result Convergence of consistency in distributed asynchronous methods for gradient iterations.

A distributed algorithm learns patterns in large images and signals.

problem High-dimensional optimization in large images and signals.
method Distributed asynchronous algorithm with locally greedy coordinate descent.
result Patterns can be learned on large scales images from the Hubble Space Telescope.

Distributed Collaborative Hashing improves recommendation efficiency in big data.

problem Efficiency in offline model training and online recommendation for collaborative filtering.
method Distributed Learning Framework + Hashing Technique.
result DCH model achieves comparable recommendation accuracy with fast convergence and real-time efficiency.

Asynchronous algorithms reduce privacy costs in distributed machine learning.

problem Privacy concerns in training machine learning models on scattered private data.
method Differentially-private asynchronous algorithms for collaborative training.
result Cost of privacy is inversely proportional to dataset size and privacy budgets.

OL4EL optimizes edge learning on resource-constrained servers.

problem Resource constraints on edge servers hinder effective distributed machine learning.
method Online Learning for EL (OL4EL) framework using budget-limited multi-armed bandit model.
result OL4EL significantly improves learning performance while conserving resources.

Rescaled ASGD optimizes distributed learning under heterogeneous data.

problem Vanilla ASGD biases towards a frequency-weighted average of local objectives.
method Rescale worker stepsizes by their computation times.
result Rescaled ASGD converges to the correct global objective in fixed-computation model.

Paper proposes a GPU-based system for training massive deep learning models in ads systems.

problem Training massive deep learning models with terabyte-scale parameters in ads systems.
method Hierarchical GPU parameter server with 3-layer storage (GPU High-Bandwidth Memory, CPU main memory, SSD).
result 4-node hierarchical GPU parameter server trains a model 2X faster than a 150-node in-memory system.

Stanza separates convolutional and fully connected layers for faster deep learning training.

problem Heavy data transfer between workers and servers in distributed deep learning.
method Layer separation: most nodes train convolutional layers, others train fully connected layers only.
result Significant acceleration of training time (1.34x--13.9x) over current systems.

DANA mitigates gradient staleness in asynchronous distributed SGD with momentum.

problem Gradient staleness in asynchronous distributed SGD with momentum.
method DANA: a novel technique for asynchronous distributed SGD with momentum that computes the gradient on an estimated future position of the model's parameters.
result DANA fully incorporates momentum in asynchronous training with almost no ramifications to final accuracy.

Paper tackles fault tolerance in distributed linear regression.

problem Fault tolerance in distributed linear regression with Byzantine faulty agents.
method Robustified distributed gradient descent using norm-based filters.
result Server can determine linear relationship deterministically in a log-linear computation cost.

This work provides bounds on generalization error and privacy leakage in federated learning.

problem Bounding generalization error and privacy leakage in federated learning.
method Information-theoretic framework for classical, distributed, and federated learning.
result Upper and lower bounds on generalization error and privacy leakage.

LiuBei is a resilient ML algorithm that tolerates Byzantine workers and servers without trusting any component.

problem Byzantine failures in distributed ML solutions.
method Byzantine-resilient ML algorithm that aggregates gradients and replicates parameter servers, using a filtering mechanism and scatter/gather protocol.
result LiuBei achieves Byzantine resilience to both servers and workers and guarantees convergence, with an accuracy loss of around 5% and a 24% convergence overhead.

Study improves distributed linear estimation under adversarial conditions.

problem Mean estimation of a random vector with adversarial measurements and asynchrony.
method Two-timescale ℓ1-minimization algorithm with tight convergence rates.
result Unified finite-time characterization of robustness, identifiability, and statistical efficiency.

DFedAvgM is a decentralized FedAvg with momentum for privacy and communication efficiency.

problem Efficiently train models with privacy and communication efficiency in federated learning.
method Decentralized Federated Averaging with Momentum (DFedAvgM) on clients connected by an undirected graph, using stochastic gradient descent with momentum and quantization.
result DFedAvgM converges under trivial assumptions and can be improved with the PŁ property, numerically verified.