Paper tackles multi-action policy learning from observational data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Neural Index Policy for multi-action bandits with heterogeneous budgets.
Paper tackles optimal policy learning with observational data in multi-action scenarios.
Lower bounds for PI on multi-action MDPs are established, showing complexity grows with action count.
Optimal policy for multi-armed multi-action bandits with unknown parameters.
Study proves duality in exotic option pricing under uncertain model and delayed information.
In this work we describe a novel deep reinforcement learning architecture that allows multiple actions to be selected at every time-step in an efficient manner. Multi-action policies allow complex behaviours to be learnt that would otherwise be hard to achieve when using single action selection techniques. We use both …
ARL and Hawkes processes improve market-making strategies with variable volatility.
Retargeting improves policy learning from observational data.
Paper derives policy rules from observational data for hepatitis C treatment.