Lyapunov-based analysis shows polynomial sample complexity for WCMDPs and RBs.
problem Learning in WCMDPs and RBs under a generative model.
method Lyapunov-based analysis framework.
result Near-optimal policies can be learned with polynomial complexity.