Simplified proof for Tsallis-INF algorithm without conjugate functions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Improved regret bounds for Tsallis-INF in adversarial bandits and corruptions.
Develops a Best-of-Both-Worlds algorithm for linear contextual bandits with Tsallis entropy.
Algorithm for bandits with switching costs achieves optimal regret bounds.
ABoB optimizes online configuration tuning by clustering parameters and accelerating learning.
We derive an algorithm that achieves the optimal (within constants) pseudo-regret in both adversarial and stochastic multi-armed bandits without prior knowledge of the regime and time horizon. The algorithm is based on online mirror descent (OMD) with Tsallis entropy regularization with power and reduced-varian…