Study non-asymptotic BPI guarantees for online RL.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Proposes proactive bed requests to reduce ED boarding and patient wait times.
NeuralChaos efficiently approximates complex stochastic processes.
DAGR improves navigation by refining goal representations conditioned on the current state.
Develops a new reinforcement learning framework for complex control problems.
DQN outperforms static policies in a dynamic fee environment for automated market makers.
This paper improves reinforcement learning by estimating return distributions using quantiles.
This paper improves reinforcement learning by estimating return distributions using quantiles.
This paper identifies drift Lipschitz budget K as key to diffusion policy expressivity and statistical trade-offs.
Survey of mathematical foundations for reinforcement learning.
EHR-MPC optimizes sepsis treatment using digital twins and inference-time control.
FORE evaluates occupancy ratios without requiring Bellman completeness.
TREK uses distillation to help students solve hard problems.
LF-IBIS learns optimal policies online without explicit likelihood.
Active-GRPO improves molecular optimization by actively deciding when to imitate or self-improve.
Transforming sparse outcomes into dense process rewards for efficient reinforcement learning.