New algorithms achieve uniform-PAC guarantees for RL with bounded eluder dimension.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New algorithm FLUTE achieves uniform-PAC convergence in RL with linear approx.
Statistical performance bounds for reinforcement learning (RL) algorithms can be critical for high-stakes applications like healthcare. This paper introduces a new framework for theoretically measuring the performance of such algorithms called Uniform-PAC, which is a strengthening of the classical Probably Approximatel…
Unified framework for anytime-valid PAC-Bayes bounds.
Tests for overfitting in machine learning models.
Unified CS for GLMs improves bandit regret bounds.
Survey of reinforcement learning guarantees with data constraints.
This paper introduces a new metric, ULI, for RL that ensures both cumulative and instantaneous performance.
Efficient RL algorithm for MDPs with linear realizability, achieving optimal regret bound.