New UCB algorithms resist contamination in bandit problems.
problem Stochastic bandit problems with ε-contaminated rewards.
method Robust mean estimators and crUCB algorithms.
result Achieves O(√KTlogT) regret for small contamination proportions.
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
New UCB algorithms resist contamination in bandit problems.
CRB tackles rising rewards in combinatorial online learning.