TUCRL efficiently explores and exploits in non-communicating MDPs without prior knowledge.
problem Efficient exploration-exploitation in non-communicating Markov Decision Processes (MDPs).
method Introduces TUCRL, the first algorithm for efficient exploration-exploitation in any finite MDP without prior knowledge.
result Derives a regret bound for weakly-communicating MDPs.