Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

0111 · Apr 201019922001200920182026
2 results for KL-UCRL

We consider model-based reinforcement learning in finite Markov De- cision Processes (MDPs), focussing on so-called optimistic strategies. In MDPs, optimism can be implemented by carrying out extended value it- erations under a constraint of consistency with the estimated model tran- sition probabilities. The UCRL2 alg…

2010-04-29abs ↗pdf ↗

Improved regret bound for reinforcement learning in MDPs with variance consideration.

problem Minimax lower bound for reinforcement learning in unknown MDPs.
method Novel analysis of KL-UCRL algorithm with variance-aware regret bound.
result Regret bound scaling as O~(extstyleSs,aVs,aT)\widetilde {\mathcal O}\Bigl({ extstyle \sqrt{S\sum_{s,a}{\bf V}^\star_{s,a}T}}\Big) for ergodic MDPs.