New algorithms learn from fixed data without exploration.
problem Learning from fixed data without additional exploration.
method Introduces batch-constrained reinforcement learning.
result First continuous control RL algorithm for fixed data.
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
New algorithms learn from fixed data without exploration.
This paper benchmarks batch RL algorithms on Atari, finding DQN and partially-trained policies perform best.
CoDA augments data with counterfactuals from local causal structures.
State-constrained offline RL expands RL's learning scope.