UCB-Advantage learns MDPs with regret.
problem Model-free reinforcement learning in finite-horizon MDPs.
method Reference-Advantage decomposition for low regret.
result Achieves regret, matching best known bounds.
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
UCB-Advantage learns MDPs with regret.
This paper improves Q-learning bounds using reference-advantage decomposition.