Optimal learning rates decay to zero in easy tasks and maintain a warmup phase in hard tasks.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
GradPower speeds up language model training with minimal code changes.
SignSGD outperforms SGD in linear regression with optimal scaling laws under PLRF model.
WSD schedule improves model training efficiency by adapting learning rates dynamically.
ScheduleFree+ improves large language model training without schedules or learning rates.
WSqD extends learning rate schedules for large model training without fixed horizons.
New method SF-AdamW trains large models without decay phases or memory overhead.
The paper presents a multi-power law for predicting loss curves across different learning rate schedules.
We discover scaling laws for kernel regression loss under various learning rate schedules.
We find optimal learning rate schedules for a random feature model.