A new bandit algorithm observes arm rewards before playing, reducing regret.
problem Balancing exploration and exploitation with pre-observation costs.
method Design of OBP-UCB for single-player and C-MP-OBP for multi-player settings.
result Proved regret bounds for both single-player and multi-player settings.