Study off-policy evaluation in partially observable environments, reducing bias and errors.
problem Bias and large errors in off-policy evaluation for partially observable environments.
method Defined and solved off-policy evaluation for POMDPs, introduced Decoupled POMDP model.
result Demonstrated and compared off-policy evaluation methods, showing benefits of new approach.
Adapts GRPO for off-policy RL, improving reward.
problem Improving training stability and efficiency in RL.
method Adapts GRPO to off-policy setting, uses clipped surrogate objectives.
result Off-policy GRPO outperforms on-policy GRPO in empirical tests.
New method improves off-policy critic evaluation in reinforcement learning.
problem High variance and instability in off-policy policy evaluation.
method Doubly robust estimators applied to actor-critic algorithms.
result Doubly robust estimation significantly improves performance in continuous control tasks.
The paper introduces a new framework for off-policy reinforcement learning.
problem Improving off-policy reinforcement learning algorithms.
method Conceptual framework based on conditional importance sampling.
result Theoretical analysis and concrete algorithm investigation.
Evolutionary Strategies optimize hyper-parameters for off-policy learning.
problem Hyper-parameter sensitivity in off-policy learning.
method Application of Evolutionary Strategies for online hyper-parameter tuning.
result Our method outperforms state-of-the-art baselines.
New method reduces state distribution mismatch in off-policy RL.
problem State distribution mismatch in off-policy RL algorithms.
method Develops a novel constrained off-policy gradient objective to minimize state distribution shift.
result Minimizing state distribution shift improves performance in off-policy RL algorithms.
Memory-efficient algorithm reduces variance in off-policy RL.
problem High variance in off-policy policy optimization.
method Memory-efficient, stochastically variance-reduced algorithm using off-policy samples.
result Empirically validated effectiveness of the proposed algorithm.
Paper improves bootstrapping for off-policy reinforcement learning inference.
problem Improving bootstrapping for off-policy reinforcement learning inference.
method Proposes a bootstrapping FQE method for off-policy statistical inference and a subsampling procedure to improve runtime.
result Asymptotically efficient and distributionally consistent bootstrapping FQE method for off-policy inference.
New method corrects state distribution mismatch for off-policy policy optimization.
problem Mismatch between behavior and evaluation policy state distributions.
method Off-policy policy gradient with state distribution correction.
result Significantly improved policy quality in simulations.
We propose and analyze an alternate approach to off-policy multi-step temporal difference learning, in which off-policy returns are corrected with the current Q-function in terms of rewards, rather than with the target policy in terms of transition probabilities. We prove that such approximate corrections are sufficien…
A new estimator improves off-policy evaluation in RL, outperforming existing methods.
problem Estimating performance of a new policy using historical data from a different policy.
method Doubly-robust estimator based on Targeted Maximum Likelihood Estimation, with variance reduction techniques.
result Our estimator uniformly outperforms existing methods across various RL environments and levels of model misspecification.
New algorithm BEAR reduces instability in off-policy Q-learning.
problem High sensitivity of off-policy Q-learning methods to data distribution.
method Identified and mitigated bootstrapping error through constrained action selection.
result BEAR algorithm learns robustly from various off-policy distributions.
Paper develops a new multi-agent reinforcement learning algorithm.
problem Improving policies in a network of communicating agents.
method Develops a multi-agent off-policy actor-critic algorithm using emphatic temporal difference learning.
result Proves convergence of the algorithm under linear function approximation.
The paper introduces a novel method for stable off-policy learning using value function chaining.
problem Stability issues in off-policy reinforcement learning.
method The approach involves learning on-policy first, then chaining off-policy value estimates.
result The method guarantees convergence and can approximate off-policy TD solutions.
Improves off-policy evaluation weights for balanced state-action pair distribution.
problem Imbalance in importance sampling weights for off-policy evaluation of contextual bandits.
method Balanced Off-Policy Evaluation (B-OPE) method that minimizes imbalance to desired counterfactual distribution of state-action pairs.
result Experimental evidence shows B-OPE improves offline policy evaluation in both discrete and continuous action spaces.
Paper tackles online off-policy prediction challenges.
problem Learning online with behavior-contingent predictions off-policy.
method Developed new objective function for TD learning, enabling stochastic gradient descent.
result Empirical study of off-policy prediction methods in microworlds.
Temporal difference learning and Residual Gradient methods are the most widely used temporal difference based learning algorithms; however, it has been shown that none of their objective functions is optimal w.r.t approximating the true value function V. Two novel algorithms are proposed to approximate the true value…
Estimates policy value from off-policy data in contextual bandits.
problem Estimating policy value from limited off-policy data in contextual bandits.
method Empirical likelihood techniques for optimization.
result Improves over previous methods in finite sample regimes.
RO-TD learns sparse value functions efficiently.
problem Learning sparse value functions efficiently.
method RO-TD integrates off-policy convergent gradient TD methods and online convex regularization.
result RO-TD learns sparse value functions with low computational complexity.
Stabilizes policy optimization with off-policy data using divergence augmentation.
problem Premature convergence and instability in policy optimization with off-policy data.
method Incorporates Bregman divergence between behavior and current policies to ensure safe policy updates.
result Empirically shows better performance in data-scarce scenarios compared to other algorithms.
Paper improves off-policy evaluation by estimating behavior policy.
problem Evaluating policies with data from a different behavior policy.
method Importance sampling with an estimated behavior policy.
result Estimating behavior policy reduces mean squared error.
Proposes a convergent TD algorithm for off-policy RL.
problem Learning value function from different policies in RL.
method Convergent on-policy TD algorithm with linear function approximation.
result Proposes a convergent TD algorithm for off-policy RL.
Monotonic policy improvement and off-policy learning are two main desirable properties for reinforcement learning algorithms. In this paper, by lower bounding the performance difference of two policies, we show that the monotonic policy improvement is guaranteed from on- and off-policy mixture samples. An optimization …
New algorithm for sequential off-policy learning improves performance over batch methods.
problem Training policies from logged interaction data in a sequential setting.
method Combines Logarithmic Smoothing with online PAC-Bayesian tools.
result Improves performance and accelerates convergence in sequential off-policy learning.
Paper proposes a framework for reliable off-policy evaluation in reinforcement learning.
problem Quantifying uncertainty in off-policy estimates for safe deployment of target policies.
method Distributionally robust optimization for creating confidence bounds.
result Non-asymptotic and asymptotic guarantees for robust cumulative reward estimates.
Novel LSE estimator improves off-policy learning and evaluation.
problem High variance and poor performance with low-quality propensity scores and heavy-tailed reward distributions.
method Introduces a novel estimator based on the log-sum-exponential (LSE) operator.
result Achieves convergence rate of O(n−ε/(1+ε)) for regret bounds. New algorithms improve policy evaluation in reinforcement learning.
problem Off-policy stability and on-policy efficiency issues in policy evaluation.
method Introduced novel algorithms using oblique projection method.
result Demonstrated both off-policy stability and on-policy efficiency.
Method selects best estimator for off-policy evaluation.
problem Choosing the best estimator for off-policy evaluation.
method Generic data-driven method for estimator selection.
result Method is competitive with oracle estimator, up to a constant factor.
OSIRIS reduces variance in off-policy evaluation by omitting irrelevant states.
problem High variance in importance sampling-based OPE estimators.
method OSIRIS reduces variance by omitting likelihood ratios associated with states irrelevant to return.
result OSIRIS is unbiased and has lower variance than ordinary importance sampling.
Develops confidence bounds for off-policy evaluation in contextual bandits.
problem Evaluating policies that were not used to collect data.
method Martingale analysis for non-asymptotic, non-parametric, and valid confidence sequences.
result Empirically tight bounds on failure probability and width.
A new theorem and algorithm solve off-policy policy gradient problems.
problem Solving the theoretical gap in off-policy policy gradient methods.
method Introduced an off-policy policy gradient theorem using emphatic weightings and developed the ACE algorithm.
result Demonstrated ACE finds the optimal solution in off-policy learning, unlike previous methods.
Improves RL algorithms with two techniques.
problem Enhance off-policy RL performance.
method Formulates RL as proximal point iteration; uses value functions for improved action value estimate.
result Significant performance improvement on RL benchmarks.
New estimator uses clustering to improve off-policy evaluation accuracy.
problem Improving off-policy evaluation accuracy when logging and evaluation policies differ.
method Proposes an estimator that shares information across similar contexts using clustering.
result Clustering contexts improves estimation accuracy, especially in deficient information settings.
Paper unifies off-policy learning algorithms and introduces C-trace for better trade-offs.
problem Improving efficiency and scalability in off-policy learning.
method Unified view of off-policy algorithms, considering update variance, fixed-point bias, and contraction rate trade-offs.
result C-trace algorithm demonstrates better trade-offs and state-of-the-art performance.
SUNRISE improves off-policy RL algorithms by integrating ensemble methods.
problem Stability and exploration issues in off-policy RL algorithms.
method SUNRISE combines ensemble-based weighted Bellman backups and upper-confidence bounds for efficient exploration.
result SUNRISE improves the performance of off-policy RL algorithms across various domains.
New method estimates state-action stationary distribution for better off-policy policy evaluation.
problem Accurately estimating state-action stationary distribution for off-policy policy evaluation.
method Estimated Mixture Policy (EMP) for state and state-action stationary distribution corrections.
result Empirical validation shows improved accuracy over state-of-the-art methods.
New algorithm reduces bias in off-policy reinforcement learning.
problem Challenges in designing off-policy reinforcement learning algorithms.
method Doubly robust off-policy actor-critic (DR-Off-PAC) with a single timescale structure.
result Establishes the first overall sample complexity analysis for a single time-scale off-policy AC algorithm.
HO2 learns options from data efficiently, improving robot manipulation tasks.
problem Learning options from raw pixel inputs in 3D robot manipulation tasks.
method HO2 infers likely option choices and trains all policy components off-policy.
result HO2 outperforms existing methods on 3D robot manipulation tasks.
First differentially private off-policy evaluation method.
problem Privacy leakage in reinforcement learning using sensitive data.
method Differential privacy approach for off-policy evaluation.
result Empirically outperforms previous methods in speed of convergence.
Boosting for off-policy learning reduces empirical risk.
problem Learning from logged bandit feedback without labeled data.
method A boosting algorithm optimizing policy's expected reward.
result Excess empirical risk decreases with each round of boosting.
Paper addresses off-policy evaluation and learning with covariate shift.
problem Evaluating and training a new policy using historical data with a covariate shift.
method Derives efficiency bounds and proposes doubly robust estimators for OPE and OPL under covariate shift.
result Proposes estimators for off-policy evaluation and learning under covariate shift.
Improved off-policy reinforcement learning by discounting and soft normalization.
problem Divergence issues in off-policy reinforcement learning.
method Introducing a discount factor and a soft normalization penalty into COP-TD.
result Discounted COP-TD is better behaved both theoretically and empirically.
Paper proposes a method to create more reliable confidence intervals for off-policy evaluations.
problem Creating reliable confidence intervals for off-policy evaluations.
method Proposes a deeply-debiasing procedure to construct efficient, robust, and flexible confidence intervals.
result Validated by theoretical results and numerical experiments, the method improves the reliability of off-policy evaluations.
The paper analyzes off-policy TD-learning using generalized Bellman operators and provides finite-sample bounds.
problem High variance in off-policy TD-learning due to importance sampling.
method Derives finite-sample bounds for off-policy TD-like algorithms using generalized Bellman operators.
result First-known finite-sample guarantees for several off-policy TD algorithms.
This paper reviews off-policy evaluation methods in reinforcement learning.
problem Efficiency and accuracy of off-policy evaluation methods in reinforcement learning.
method Discussion of existing OPE methods, their statistical properties, and related research directions.
result Efficiency bounds and state-of-the-art OPE methods in reinforcement learning.
P3O merges on-policy and off-policy updates without extra hyper-parameters.
problem Combining on-policy and off-policy RL algorithms to reduce sample complexity.
method Interleaves off-policy updates with on-policy updates using effective sample size.
result P3O reduces sample complexity of state-of-the-art algorithms.
This work analyzes the gap between off-policy and on-policy policy gradient methods and provides conditions to reduce this gap.
problem The gap between off-policy and on-policy policy gradient methods and conditions to reduce it.
method Theoretical analysis and empirical evidence of conditions to reduce the on-off gap.
result Conditions to reduce the on-off gap between off-policy and on-policy policy gradient methods.
Simplifies RL training with fewer techniques, reducing bias and instability.
problem Training instabilities and high sample complexity in RL.
method Introduced a simple deterministic policy gradient, used propensity estimation, and delayed policy updates.
result Improved performance and reduced sample complexity through these techniques.