LocalNewton reduces communication in distributed learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A novel MCMC method clusters data faster and more accurately.
In this paper, we focus on approaches to parallelizing stochastic gradient descent (SGD) wherein data is farmed out to a set of workers, the results of which, after a number of updates, are then combined at a central master node. Although such synchronized SGD approaches parallelize well in idealized computing environm…
Asynchronous parallel optimization algorithms for solving large-scale machine learning problems have drawn significant attention from academia to industry recently. This paper proposes a novel algorithm, decoupled asynchronous proximal stochastic gradient descent (DAP-SGD), to minimize an objective function that is the…
The event-driven and elastic nature of serverless runtimes makes them a very efficient and cost-effective alternative for scaling up computations. So far, they have mostly been used for stateless, data parallel and ephemeral computations. In this work, we propose using serverless runtimes to solve generic, large-scale …
In this paper, we propose a bootstrap method applied to massive data processed distributedly in a large number of machines. This new method is computationally efficient in that we bootstrap on the master machine without over-resampling, typically required by existing methods \cite{kleiner2014scalable,sengupta2016subsam…
We propose communication-efficient distributed estimation and inference methods for the transelliptical graphical model, a semiparametric extension of the elliptical distribution in the high dimensional regime. In detail, the proposed method distributes the -dimensional data of size generated from a transellipti…
Low complexity decentralized neural net with centralized performance.
We study the problem of stochastic optimization for deep learning in the parallel computing environment under communication constraints. A new algorithm is proposed in this setting where the communication and coordination of work among concurrent processes (local workers), is based on an elastic force which links the p…
Paper tackles Byzantine attacks in distributed learning with a new ADMM method.
Large-scale machine learning models are often trained by parallel stochastic gradient descent algorithms. However, the communication cost of gradient aggregation and model synchronization between the master and worker nodes becomes the major obstacle for efficient learning as the number of workers and the dimension of …
Adaptive distributed SGD reduces delay in slow workers.
Many popular distributed optimization methods for training machine learning models fit the following template: a local gradient estimate is computed independently by each worker, then communicated to a master, which subsequently performs averaging. The average is broadcast back to the workers, which use it to perform a…
We propose a novel, efficient approach for distributed sparse learning in high-dimensions, where observations are randomly partitioned across machines. Computationally, at each round our method only requires the master machine to solve a shifted ell_1 regularized M-estimation problem, and other workers to compute the g…
New algorithm resists Byzantine attacks in distributed SGD for heterogeneous data.
Proposes a method to improve Byzantine-robustness in compressed federated learning.
We focus on the commonly used synchronous Gradient Descent paradigm for large-scale distributed learning, for which there has been a growing interest to develop efficient and robust gradient aggregation strategies that overcome two key system bottlenecks: communication bandwidth and stragglers' delays. In particular, R…
New optimization algorithm for mixed-variable problems improves efficiency.
Crowdsourced labeling recovers task types with minimal queries.
New model optimizes worker-task specialization for crowdsourcing.
We prove an explicit characterization of the points in Thurston's Master Teapot. This description can be implemented algorithmically to test whether a point in belongs to the complement of the Master Teapot. As an application, we show that the intersection of the Master Teapot with the un…
Master algorithm fails to detect non-stationarity in practical settings.
Crowdsourcing provides a popular paradigm for data collection at scale. We study the problem of selecting subsets of workers from a given worker pool to maximize the accuracy under a budget constraint. One natural question is whether we should hire as many workers as the budget allows, or restrict on a small number of …
A new algorithm reduces training time for distributed machine learning by dynamically assigning backup workers.
Due to concerns about human error in crowdsourcing, it is standard practice to collect labels for the same data point from multiple internet workers. We here show that the resulting budget can be used more effectively with a flexible worker assignment strategy that asks fewer workers to analyze easy-to-label data and m…
We extend the Chern-Simons perturbative invariant of Axelrod and Singer to non-acyclic connections. We construct a solution of the quantum master equation on the space of functions on the cohomology of the connection. We prove that this solution is well defined up to master homotopy. We discuss also invariants of links…
Consider unsupervised clustering of objects drawn from a discrete set, through the use of human intelligence available in crowdsourcing platforms. This paper defines and studies the problem of universal clustering using responses of crowd workers, without knowledge of worker reliability or task difficulty. We model sto…
In distributed optimization and machine learning, multiple nodes coordinate to solve large problems. To do this, the nodes need to compress important algorithm information to bits so that it can be communicated over a digital channel. The communication time of these algorithms follows a complex interplay between a) the…
We construct master spaces for oriented torsion free sheaves coupled with morphisms into a fixed reference sheaf. These spaces are projective varieties endowed with a natural $\C^*$-action. The fixed point set of this action contains the moduli space of semistable oriented torsion free sheaves and the quot scheme assoc…
Generalizes Hodge correlators using quantum master equation concepts.
Two neural network methods solve the master equation for MFGs.
We propose Zeno++, a new robust asynchronous Stochastic Gradient Descent~(SGD) procedure which tolerates Byzantine failures of the workers. In contrast to previous work, Zeno++ removes some unrealistic restrictions on worker-server communications, allowing for fully asynchronous updates from anonymous workers, arbitrar…
Master-slave architecture tackles combinatorial multi-armed bandits with diversity constraints.
New algorithm outperforms existing ones by focusing on mastering rate.
While training a machine learning model using multiple workers, each of which collects data from their own data sources, it would be most useful when the data collected from different workers can be {\em unique} and {\em different}. Ironically, recent analysis of decentralized parallel stochastic gradient descent (D-PS…
The number of Italian firms in function of the number of workers is well approximated by an inverse power law up to 15 workers but shows a clear downward deflection beyond this point, both when using old pre-1999 data and when using recent (2014) data. This phenomenon could be associated with employent protection legis…
We analyze the convergence of gradient-based optimization algorithms that base their updates on delayed stochastic gradient information. The main application of our results is to the development of gradient-based distributed optimization algorithms where a master node performs parameter updates while worker nodes compu…
A master equation approach to the numerical solution of option pricing models is developed. The basic idea of the approach is to consider the Black--Scholes equation as the macroscopic equation of an underlying mesoscopic stochastic option price variable. The dynamics of the latter is constructed and formulated in term…
DynBRO learns robustly from dynamic Byzantine workers.
Generative AI boosts productivity and improves customer service quality.
Deep neural networks improve sEMG-based hand gesture classification.
A method to robustly federate learning with non-i.i.d. data and Byzantine workers.
Cost-efficient distributed learning via combinatorial bandits.
Crowdsourcing is a relatively economic and efficient solution to collect annotations from the crowd through online platforms. Answers collected from workers with different expertise may be noisy and unreliable, and the quality of annotated data needs to be further maintained. Various solutions have been attempted to ob…
We establish basic geometric and topological properties of Thurston's Master Teapot and the Thurston set for superattracting unimodal self-maps of intervals. In particular, the Master Teapot is connected, contains the unit cylinder, and its intersection with a set grows monotonically with .…
Optimizes master faces for 2D and 3D face verification using evolutionary algorithms and neural networks.
To accelerate the training of machine learning models, distributed stochastic gradient descent (SGD) and its variants have been widely adopted, which apply multiple workers in parallel to speed up training. Among them, Local SGD has gained much attention due to its lower communication cost. Nevertheless, when the data …
We consider the problem faced by a service platform that needs to match limited supply with demand but also to learn the attributes of new users in order to match them better in the future. We introduce a benchmark model with heterogeneous "workers" (demand) and a limited supply of "jobs" that arrive over time. Job typ…