Improved gap-dependent bounds for reinforcement learning with linear approximations.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Improved privacy in RL with near-optimal regret bounds.
Efficient RL algorithms for linear function approximation with limited adaptivity constraints.
New RL algorithm tackles nonstationary MDPs with linear approximations and varying rewards.
New RL algorithm achieves sublinear regret and constraint violation without simulators.
Logarithmic regret achieved in RL with linear function approximation.
A new RL approach optimizes reserve prices in multi-phase auctions, reducing revenue regret.
New framework for reinforcement learning with sporadic state observations.
We prove a lower bound for feature dimension in linear MDPs and propose a novel dynamics aggregation framework.
New algorithm tackles heavy-tailed rewards in RL with instance-dependent regret bounds.