Improved model-free reinforcement learning with decision-estimation coefficient.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New bounds show complexity of adversarial decision making.
New bounds for -regret using modified Decision-Estimation Coefficient.
New DEC variant improves sample complexity bounds in decision making.
New complexity measure for interactive learning reduces regret to near-optimal levels.
Unified algorithm tackles various RL goals like reward-free and preference-based learning.
OE2D framework reduces contextual bandits to offline regression for near-optimal regret.
Unified framework for lower bounds in interactive decision making.
We characterize learnability for stochastic noisy bandits, identifying optimal query complexities.
New algorithms reduce sample complexity for multiclass contextual bandits.
Framework for robust decision making in changing environments with privacy constraints.
Framework reduces contextual bandit learning to offline regression with near-optimal regret.
Improves decision complexity in hybrid environments.
Study on multi-agent decision making complexity, showing sample efficiency gaps.
Study shows offline RL under -approximation and partial coverage is harder than previously thought.