The paper analyzes the variance of different shuffling methods in stochastic gradient descent.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
RR algorithm improves convergence rate without strong convexity assumptions.
New convergence rates for shuffling gradient methods without strong convexity.
Unified analysis of asynchronous-SGD algorithms for distributed learning.
Sampling without replacement speeds up optimization in minimax problems.
Improved convergence for VIPs with SEG-RR, a variant of SEG with random reshuffling.
We study the performance of stochastic gradient descent (SGD) on smooth and strongly-convex finite-sum optimization problems. In contrast to the majority of existing theoretical works, which assume that individual functions are sampled with replacement, we focus here on popular but poorly-understood heuristics, which i…