DSVGD improves federated learning with fewer communication rounds.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This work investigates how multi-round reasoning improves LLM performance.
In machine learning the best performance on a certain task is achieved by fully supervised methods when perfect ground truth labels are available. However, labels are often noisy, especially in remote sensing where manually curated public datasets are rare. We study the multi-modal cadaster map alignment problem for wh…
Dynamic SBI improves SBI efficiency without rounds, reducing simulation and training costs.
We introduce a collaborative learning framework allowing multiple parties having different sets of attributes about the same user to jointly build models without exposing their raw data or model parameters. In particular, we propose a Federated Stochastic Block Coordinate Descent (FedBCD) algorithm, in which each party…
A new decentralized algorithm DESTINY solves optimization over Stiefel manifold with single communication round.
The AdaBoost algorithm was designed to combine many "weak" hypotheses that perform slightly better than random guessing into a "strong" hypothesis that has very low error. We study the rate at which AdaBoost iteratively converges to the minimum of the "exponential loss." Unlike previous work, our proofs do not require …
A new algorithm finds minimizers in dueling optimization with a monotone adversary.
We address the M-best-arm identification problem in multi-armed bandits. A player has a limited budget to explore K arms (M<K), and once pulled, each arm yields a reward drawn (independently) from a fixed, unknown distribution. The goal is to find the top M arms in the sense of expected reward. We develop an algorithm …
This paper introduces a new metric, ULI, for RL that ensures both cumulative and instantaneous performance.
Knowledge distillation introduced in the deep learning context is a method to transfer knowledge from one architecture to another. In particular, when the architectures are identical, this is called self-distillation. The idea is to feed in predictions of the trained model as new target values for retraining (and itera…
In this paper, we present a scalable distributed implementation of the Sampled Limited-memory Symmetric Rank-1 (S-LSR1) algorithm. First, we show that a naive distributed implementation of S-LSR1 requires multiple rounds of expensive communications at every iteration and thus is inefficient. We then propose DS-LSR1, a …
FedAc accelerates Federated Averaging for distributed optimization.
This study models FOMC policy decisions using debate-based LLMs.
The paper develops a theory for iterative self-improvement of models, proving conditions for better performance with easy-to-hard curricula.
Support vector machines (SVMs) are invaluable tools for many practical applications in artificial intelligence, e.g., classification and event recognition. However, popular SVM solvers are not sufficiently efficient for applications with a great deal of samples as well as a large number of features. In this paper, thus…
New method averages SGD iterates to achieve adjustable regularization.
AI beats 95% of humans in Rock-Paper-Scissors.
A distributed bootstrap method for high-dimensional data reduces communication rounds efficiently.
New inequalities help optimize first-order algorithms for statistical risk analysis.
We consider the problem of strongly-convex online optimization in presence of adversarial delays; in a T-iteration online game, the feedback of the player's query at time t is arbitrarily delayed by an adversary for d_t rounds and delivered before the game ends, at iteration t+d_t-1. Specifically for \algo{online-gradi…
LocalNewton reduces communication in distributed learning.
We derive the optimal differential privacy (DP) parameters of a mechanism that satisfies a given level of Rényi differential privacy (RDP). Our result is based on the joint range of two -divergences that underlie the approximate and the Rényi variations of differential privacy. We apply our result to the moments acc…
We present a distributed proximal-gradient method for optimizing the average of convex functions, each of which is the private local objective of an agent in a network with time-varying topology. The local objectives have distinct differentiable components, but they share a common nondifferentiable component, which has…
Paper proposes AggITD for efficient federated hypergradient computation.
Gradient descent with biased rounding errors converges faster under certain conditions.
Flexible algorithms for maximizing rewards in structured bandits.
Around 2007, A. Chang, J. Qing, and P. Yang proved a conformal gap theorem for Bach-flat metrics with round sphere as the model case. In this article, we extend this result to prove conformally invariant gap theorems for Bach-flat -manifolds with and $(\mathbb{S}^2\times\mathbb{S}^2,g_{prod…
We study distributed algorithms for expected loss minimization where the datasets are large and have to be stored on different machines. Often we deal with minimizing the average of a set of convex functions where each function is the empirical risk of the corresponding part of the data. In the distributed setting wher…
Federated learning opens a number of research opportunities due to its high communication efficiency in distributed training problems within a star network. In this paper, we focus on improving the communication efficiency for fully decentralized federated learning over a graph, where the algorithm performs local updat…
Round surgery diagrams represent 3-manifolds in .
Critical volatility triggers log-normal to power-law transitions in interconnected systems.
In recent years, content recommendation systems in large websites (or \emph{content providers}) capture an increased focus. While the type of content varies, e.g.\ movies, articles, music, advertisements, etc., the high level problem remains the same. Based on knowledge obtained so far on the user, recommend the most d…
New methods improve Byzantine robustness in distributed learning.
Thompson sampling for multi-armed bandit problems is known to enjoy favorable performance in both theory and practice. However, it suffers from a significant limitation computationally, arising from the need for samples from posterior distributions at every iteration. We propose two Markov Chain Monte Carlo (MCMC) meth…
We consider the best-arm identification problem in multi-armed bandits, which focuses purely on exploration. A player is given a fixed budget to explore a finite set of arms, and the rewards of each arm are drawn independently from a fixed, unknown distribution. The player aims to identify the arm with the largest expe…
Contact round surgeries on help in constructing and understanding contact 3-manifolds.
In this article, we extend Huisken's theorem that convex surfaces flow to round points by mean curvature flow. We construct certain classes of mean convex and non-mean convex hypersurfaces that shrink to round points and use these constructions to create pathological examples of flows. We find a sequence of flows that …
Smooth Freund-Rubin backgrounds of eleven-dimensional supergravity of the form AdS_4 x X^7 and preserving at least half of the supersymmetry have been recently classified. Requiring that amount of supersymmetry forces X to be a spherical space form, whence isometric to the quotient of the round 7-sphere by a freely-act…
Optimizes sample and round complexity in adaptive sampling from multiple distributions.
The paper develops efficient estimators for semi-parametric binary models in distributed computing.
Consider an analytic map of a neighborhood of 0 in a vector space to a Euclidean space. Suppose that this map takes all germs of lines passing through 0 to germs of circles. Such a map is called rounding. We introduce a natural equivalence relation on roundings and prove that any rounding, whose differential at 0 has r…
New findings show infinitely many knots cannot be smoothly round handle slices.
New method for online low-rank matrix completion with improved regret.
We discuss the integrability of orthogonal almost complex structures on Riemannian products of even-dimensional round spheres and give a partial answer to the question raised by E. Calabi concerning the existence of complex structures on a product manifold of a round 2-sphere and a round 4-sphere.
Paper develops a robust PP distributed quasi-Newton estimation for Byzantine machines.
New findings on -solutions with round cylinder as asymptotic shrinker.
FedRD improves risk difference estimation in federated learning for clinical outcomes.