GUM tackles MARL by avoiding overestimation through state-marginal restriction.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Exploration is critical to a reinforcement learning agent's performance in its given environment. Prior exploration methods are often based on using heuristic auxiliary predictions to guide policy behavior, lacking a mathematically-grounded objective with clear properties. In contrast, we recast exploration as a proble…
Unified probabilistic perspective on imitation learning methods using divergence minimization.
New CTRL algorithm adapts to varying problem difficulty.
AIF reformulated as convex MDP for adaptive behavior.
Improved exploration in RL with latent state marginalization.
A new method for learning policies from demonstrations without reinforcement.
Motivated by the many real-world applications of reinforcement learning (RL) that require safe-policy iterations, we consider the problem of off-policy evaluation (OPE) -- the problem of evaluating a new policy using the historical data obtained by different behavior policies -- under the model of nonstationary episodi…
GDT improves reinforcement learning by matching future state information efficiently.