Partial model averaging improves Federated Learning performance.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Local averaging accurately distills manifold structure from noisy data.
Study the averaging principle for non-autonomous slow-fast systems and apply it to financial local stochastic volatility models.
This study optimizes model averaging for personalized collaborative learning.
New averaging technique speeds up Newton method convergence.
Reducing communication in training large-scale machine learning applications on distributed platform is still a big challenge. To address this issue, we propose a distributed hierarchical averaging stochastic gradient descent (Hier-AVG) algorithm with infrequent global reduction by introducing local reduction. As a gen…
The study finds the minimum average area ratio on hyperbolic manifolds and its relation to scalar curvature.
An approach to distributed machine learning is to train models on local datasets and aggregate these models into a single, stronger model. A popular instance of this form of parallelization is federated learning, where the nodes periodically send their local models to a coordinator that aggregates them and redistribute…
Averaged SGD optimizes a smoothed objective, leading to better generalization.
We present sufficient conditions for topological stability of continuous functions having finitely many local extrema with respect to averagings by discrete measures with finite supports.
Quantum walks blend patterns into splines when averaged.
New methods optimize functions faster with less gradient accuracy needed.
This dissertation advances the theoretical foundation of local optimization methods in Federated Learning.
A new method averages neural network parameters to rank features robustly.
Distributed statistical learning problems arise commonly when dealing with large datasets. In this setup, datasets are partitioned over machines, which compute locally, and communicate short messages. Communication is often the bottleneck. In this paper, we study one-step and iterative weighted parameter averaging in s…
We describe an averaging procedure on a Dirac manifold, with respect to a class of compatible actions of a compact Lie group. Some averaging theorems on the existence of invariant realizations of Poisson structures around (singular) symplectic leaves are derived. We show that the construction of coupling Dirac structur…
In distributed optimization and distributed numerical linear algebra, we often encounter an inversion bias: if we want to compute a quantity that depends on the inverse of a sum of distributed matrices, then the sum of the inverses does not equal the inverse of the sum. An example of this occurs in distributed Newton's…
Study on predicting graph labels at nodes using local averaging and distance estimation.
Bayesian framework mixes imperfect models for improved predictions.
The paper proves local laws for non-separable sample covariance matrices.
Deep neural network approximates flow averages for rough walls in multiscale simulations.
New perspective on federated learning as posterior inference, improving optimization.
LCMQR improves prediction intervals by adapting to local heteroscedasticity.
Federated learning is a distributed framework according to which a model is trained over a set of devices, while keeping data localized. This framework faces several systems-oriented challenges which include (i) communication bottleneck since a large number of devices upload their local updates to a parameter server, a…
Given a Finsler space (M,F) on a manifold M, the averaging method associates to Finslerian geometric objects affine geometric objects} living on . In particular, a Riemannian metric is associated to the fundamental tensor and an affine, torsion free connection is associated to the Chern-Rund connection. As an il…
Despite the success of single-agent reinforcement learning, multi-agent reinforcement learning (MARL) remains challenging due to complex interactions between agents. Motivated by decentralized applications such as sensor networks, swarm robotics, and power grids, we study policy evaluation in MARL, where agents with jo…
A2SGD reduces distributed SGD communication to O(1) per worker.
New method uses LP to achieve optimal sample complexity in multi-agent reinforcement learning.
This paper examines ADMM for network averaging, revealing its convergence rates and network topology impacts.
New method removes interference bias in causal models.
Sharp bounds established for Federated Averaging (FedAvg), improving convergence rates.
We derive identities for general flows of Riemannian metrics that may be regarded as local mean-value, monotonicity, or Lyapunov formulae. These generalize previous work of the first author for mean curvature flow and other nonlinear diffusions. Our results apply in particular to Ricci flow, where they yield a local mo…
Proposes a new random forest weighted local Fréchet regression method.
Large-scale machine learning training, in particular distributed stochastic gradient descent, needs to be robust to inherent system variability such as node straggling and random communication delays. This work considers a distributed training framework where each worker node is allowed to perform local model updates a…
This paper studies the correlations of the average winnings of agents and the volatilities of systems based on mix-game model which is an extension of minority game (MG). In mix-game, there are two groups of agents; group1 plays the majority game, but the group2 plays the minority game. The results show that the correl…
Federated Q-learning achieves linear speedup with heterogeneity, improving sample complexity.
Two-Tailed Averaging improves generalization by optimizing the number of leading iterates to ignore.
This paper proposes recurrent neuron networks (RNNs) for a fingerprinting indoor localization using WiFi. Instead of locating user's position one at a time as in the cases of conventional algorithms, our RNN solution aims at trajectory positioning and takes into account the relation among the received signal strength i…
Study compares local and global models for hierarchical forecasting accuracy.
Communication overhead is one of the key challenges that hinders the scalability of distributed optimization algorithms. In this paper, we study local distributed SGD, where data is partitioned among computation nodes, and the computation nodes perform local updates with periodically exchanging the model among the work…
We consider a random sparse graph with bounded average degree, in which a subset of vertices has higher connectivity than the background. In particular, the average degree inside this subset of vertices is larger than outside (but still bounded). Given a realization of such graph, we aim at identifying the hidden subse…
This paper proposes a simple but effective graph-based agglomerative algorithm, for clustering high-dimensional data. We explore the different roles of two fundamental concepts in graph theory, indegree and outdegree, in the context of clustering. The average indegree reflects the density near a sample, and the average…
Study shows Bergman kernels match averages on quotient spaces, proving non-vanishing of Poincaré series.
Communication-efficient SGD algorithms, which allow nodes to perform local updates and periodically synchronize local models, are highly effective in improving the speed and scalability of distributed SGD. However, a rigorous convergence analysis and comparative study of different communication-reduction strategies rem…
This study examined how the correlation and network structure of 30 global indices and 145 local Korean indices belonging to the KOSPI 200 have changed during the 13-year period, 2000-2012. The correlations among the indices were calculated. The results showed that although the average correlations of the global indice…
The paper rethinks the use of exponential averaging in machine learning optimization.
In this paper, we derive a new model of synaptic plasticity, based on recent algorithms for reinforcement learning (in which an agent attempts to learn appropriate actions to maximize its long-term average reward). We show that these direct reinforcement learning algorithms also give locally optimal performance for the…
The effects of weather on agriculture in recent years have become a major global concern. Hence, the need for an effective weather risk management tool (i.e., weather derivatives) that can hedge crop yields against weather uncertainties. However, most smallholder farmers and agricultural stakeholders are unwilling to p…