A new method improves reinforcement learning by directing exploration towards new knowledge.
problem Efficient exploration in reinforcement learning, especially in complex environments.
method Proposed -values, a generalization of visit-counters, for directed exploration in model-free reinforcement learning.
result Improves learning and performance in continuous Markov Decision Processes (MDPs) compared to traditional methods.