Paper tackles online off-policy prediction challenges.
problem Learning online with behavior-contingent predictions off-policy.
method Developed new objective function for TD learning, enabling stochastic gradient descent.
result Empirical study of off-policy prediction methods in microworlds.