Guiding reinforcement learning with suboptimal controllers speeds up training.
problem Sparse rewards in reinforcement learning make exploration inefficient.
method Use a suboptimal controller to guide exploration, applying a Q-filter loss conditionally.
result The approach leads to faster policy refinement and better performance.