Improved analysis of UCBVI algorithm with better empirical performance.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Federated UCBVI reduces communication costs while minimizing regret in multi-agent settings.
UCBVI-γ algorithm minimizes regret in discounted MDPs.
New formalism for decision making combines causal structures with MDPs, improving reinforcement learning performance.
Bayes-UCBVI tackles reinforcement learning with a new upper confidence bound method.
Kernel-UCBVI algorithm balances exploration and exploitation in metric state-action spaces.
Unified framework for distributional regret in bandits and reinforcement learning.
The paper introduces a method to learn and apply value envelopes for faster online reinforcement learning.
Develops a new model for RLHF accounting for partially observed states and intermediate feedback.