Enhances POLITEX for exploration in reinforcement learning with no-reward learning.
problem Learning to explore in reinforcement learning problems with no-reward.
method Modifies POLITEX to incorporate a pre-existing exploration policy.
result Achieves sublinear regret guarantees similar to POLITEX but without requiring all policies to explore.