PHE adds pseudo-rewards to history to minimize regret in stochastic bandits.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We propose a bandit algorithm that explores by randomizing its history of rewards. Specifically, it pulls the arm with the highest mean reward in a non-parametric bootstrap sample of its history with pseudo rewards. We design the pseudo rewards such that the bootstrap mean is optimistic with a sufficiently high probabi…
Successor Options discovers reusable skills using landmark states.
New algorithm minimizes regret in stochastic linear bandits with perturbed history.
Paper shows RL's policy gradient methods can be viewed as supervised learning problems.
New algorithm for RL with horizon-free reward-free exploration for linear MDPs.