UCB-V algorithm improves on UCB for MAB problems with variance estimates.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Unified proof for various bandit algorithms with logarithmic regret.
RAVEN-UCB addresses non-stationary MAB problems with tighter regret bounds.
The paper analyzes the sliding regret of stochastic bandit algorithms.
PCTS optimizes noisy, delayed, multi-fidelity feedbacks in black-box optimization.
The exploration/exploitation (E/E) dilemma arises naturally in many subfields of Science. Multi-armed bandit problems formalize this dilemma in its canonical form. Most current research in this field focuses on generic solutions that can be applied to a wide range of problems. However, in practice, it is often the case…