Upside-Down Reinforcement Learning Can Diverge in Stochastic Environments With Episodic Resetsstat.ML
UDRL fails to converge in stochastic environments with episodic resets.
problem UDRL's convergence in stochastic environments with resets is questioned.
method UDRL is a supervised learning approach that does not use value functions.
result UDRL diverges in a simple stochastic environment with resets.