Local Gradient Descent with local steps converges to the centralized model in the interpolation regime.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
AutoStep MCMC adapts step size locally for better sampling efficiency.
GradSkip reduces local training steps for better communication efficiency.
LocalKMeans parallelizes Lloyd's algorithm for distributed data.
New -step policy gradient method avoids local optima in restricted policy classes.
Proposes a graph dynamics prior for more accurate relational inference.
This is a final step in a local classification of pseudo-Riemannian manifolds with parallel Weyl tensor that are not conformally flat or locally symmetric.
Efficient regularization mitigates catastrophic overfitting in single-step adversarial training.
New analysis shows Local SGD can achieve error scaling with only fixed number of communications.
A new method automatically and dynamically sets learning rates in deep learning.
Proposes a method combining CNFs and rejection-resampling for sampling from unnormalized densities.
Gradient descent converges linearly in finite-width networks with positive NTK and compatible conditions.
We investigate finite-time decoupled convergence in nonlinear two-time-scale stochastic approximation.
In this paper, we propose a simple, fast and easy to implement algorithm LOSSGRAD (locally optimal step-size in gradient descent), which automatically modifies the step-size in gradient descent during neural networks training. Given a function , a point , and the gradient of , we aim to find the s…
Distributed optimization often consists of two updating phases: local optimization and inter-node communication. Conventional approaches require working nodes to communicate with the server every one or few iterations to guarantee convergence. In this paper, we establish a completely different conclusion that each node…
Paper proposes faster method to find local minima in nonconvex optimization.
Stochastic (sub)gradient methods require step size schedule tuning to perform well in practice. Classical tuning strategies decay the step size polynomially and lead to optimal sublinear rates on (strongly) convex problems. An alternative schedule, popular in nonconvex optimization, is called \emph{geometric step decay…
Paper develops Gaussian approximations and bootstrap for federated LSA with trade-off bounds.
Inverted file and asymmetric distance computation (IVFADC) have been successfully applied to approximate nearest neighbor search and subsequently maximum inner product search. In such a framework, vector quantization is used for coarse partitioning while product quantization is used for quantizing residuals. In the ori…
GD converges in unstable regimes, even with oscillatory behavior.
Multiple gossip steps improve decentralized optimization convergence.
funLOCI identifies clusters in functional data.
New optimization for federated learning with local models.
New techniques improve distributed training with compressed gradients.
Proposes an exponentially increasing step-size for faster parameter estimation in statistical models.
We focus our attention on the notion of intrinsic Lipschitz graphs, inside a special class of metric spaces i.e. the Carnot groups. More precisely, we provide a characterization of locally intrinsic Lipschitz functions in Carnot groups of step 2 in terms of their intrinsic distributional gradients.
Data privacy is an important concern in learning, when datasets contain sensitive information about individuals. This paper considers consensus-based distributed optimization under data privacy constraints. Consensus-based optimization consists of a set of computational nodes arranged in a graph, each having a local ob…
FedSARSA converges with heterogeneous agents, achieving linear speed-up.
ISALT uses inference to simulate SDEs with large time-steps, improving efficiency.
Boosting Variational Inference improves posterior approximations with adaptive step-sizes.
Asynchronous decentralized SGD with quantized and local updates converges in gossip model.
New adaptive step-size method for convex optimization without tuning.
This paper is devoted to a priori estimates for strictly locally convex radial graphs with prescribed Weingarten curvature and boundary in space forms. By constructing two-step continuity process and applying degree theory arguments, existence results in space forms are established for prescribed Gauss curvature …
We prove that a compact toric locally conformally Kähler manifold which is not Kähler admits a toric Vaisman structure, a fact which was conjectured in \cite{mmp}. This is the final step leading to the classification of compact toric locally conformally Kähler manifolds started in \cite{p} and \cite{mmp}. We also show,…
We take initial steps in studying PAC-MDP algorithms with limited adaptivity, that is, algorithms that change its exploration policy as infrequently as possible during regret minimization. This is motivated by the difficulty of running fully adaptive algorithms in real-world applications (such as medical domains), and …
Conditional gradients constitute a class of projection-free first-order algorithms for smooth convex optimization. As such, they are frequently used in solving smooth convex optimization problems over polytopes, for which the computational cost of orthogonal projections would be prohibitive. However, they do not enjoy …
This paper is the second part of a series of papers on noncommutative geometry and conformal geometry. In this paper, we compute explicitly the Connes-Chern character of an equivariant Dirac spectral triple. The formula that we obtain for which was used in the first paper of the series. The computation has two main ste…
Gradient method converges locally linearly for overparameterized Gaussian mixtures.
Gradient descent converges linearly for overparameterized linear networks.
ATLAS adapts HMC step size and trajectory length for complex geometries.
CW-EDMD improves prediction accuracy by learning local Koopman models for different state-space regions.
New method calibrates local volatility models to marginal distributions.
We consider the non-parametric regression problem under Huber's -contamination model, in which an fraction of observations are subject to arbitrary adversarial noise. We first show that a simple local binning median step can effectively remove the adversary noise and this median estimator is minimax optimal up t…
In this paper, we obtain a localization formula in differential K-theory for -action. Then by combining an extension of Goette's result on the comparison of two types of equivariant -invariants, we establish a version of localization formula for equivariant -invariants. An important step of our approach is t…
Community detection in hypergraphs is explored. Under a generative hypergraph model called "d-wise hypergraph stochastic block model" (d-hSBM) which naturally extends the Stochastic Block Model from graphs to d-uniform hypergraphs, the asymptotic minimax mismatch ratio is characterized. For proving the achievability, w…
The family of Expectation-Maximization (EM) algorithms provides a general approach to fitting flexible models for large and complex data. The expectation (E) step of EM-type algorithms is time-consuming in massive data applications because it requires multiple passes through the full data. We address this problem by pr…
Deep learning model predicts traffic flows across entire network for multiple steps ahead.
Recent innovations in Information and Communication Technologies (ICT) provide new opportunities and challenges for integration of distributed energy resources (DERs) into the energy supply system as active market players. By increasing integration of DERs, novel market platform should be designed for these new market …