Bayesian approach improves learning in RMDPs with faster adaptation.
problem Learning robust policies in RMDPs with changing or adversarial dynamics.
method Introduce Uncertainty Robust Bellman Equation (URBE) and DQN-URBE algorithm.
result URBE-based strategy leads to better trade-off between robustness and exploration.