Two algorithms achieve optimal regret with limited adaptivity in multinomial logistic bandits.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New RL algorithm achieves nearly optimal performance for linear MDPs.
Excessively changing policies in many real world scenarios is difficult, unethical, or expensive. After all, doctor guidelines, tax codes, and price lists can only be reprinted so often. We may thus want to only change a policy when it is probable that the change is beneficial. In cases that a policy is a threshold on …
Efficient algorithms for contextual slate bandits with limited adaptivity.
New actor-critic algorithm achieves optimal sample efficiency in RL.
We develop a robust RL algorithm for off-dynamics environments with improved suboptimality bounds and computational efficiency.