Improved FPL algorithms for adversarial MDPs with better regret bounds.
problem Adversarial rewards and unknown transitions in MDPs.
method Refined analysis of FPL algorithms, matching current best regret bounds.
result Improved regret bounds for FPL algorithms in adversarial MDPs.