UC-SSP algorithm tackles exploration in stochastic shortest path problems without loop-free assumption.
problem Exploration in goal-oriented reinforcement learning problems under stochastic shortest path formulation.
method UC-SSP algorithm with a novel stopping rule to interrupt and switch policies.
result Regret bound of after episodes.