EF21-Muon optimizes deep learning with error feedback, improving efficiency and accuracy.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Optimal scaling found to depend on operator norm across large models and datasets.
Gluon optimizes LMO-based methods for large-scale tasks, improving performance and theory-practice gap.
New optimization method combines gradient clipping and non-Euclidean smoothness.
New analysis reveals batch size effects on stochastic conditional gradient methods.
Muon optimizes Transformer training with heavy-tailed data, achieving optimal sample complexity.
Non-Euclidean BPM extends optimization theory to non-Euclidean norms.
New optimizer designs respect symmetry, improving deep learning models.