Paper unifies off-policy learning algorithms and introduces C-trace for better trade-offs.
problem Improving efficiency and scalability in off-policy learning.
method Unified view of off-policy algorithms, considering update variance, fixed-point bias, and contraction rate trade-offs.
result C-trace algorithm demonstrates better trade-offs and state-of-the-art performance.