We consider model-based reinforcement learning in finite Markov De- cision Processes (MDPs), focussing on so-called optimistic strategies. In MDPs, optimism can be implemented by carrying out extended value it- erations under a constraint of consistency with the estimated model tran- sition probabilities. The UCRL2 alg…
On-device research index
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
169,181 papers · 148 categories
Trend · papers per month
2 results for “KL-UCRL”
Improved regret bound for reinforcement learning in MDPs with variance consideration.
problem Minimax lower bound for reinforcement learning in unknown MDPs.
method Novel analysis of KL-UCRL algorithm with variance-aware regret bound.
result Regret bound scaling as for ergodic MDPs.