Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

6.3%12.5%18.8%25.0% · Mar 199319922001200920182026
48 results for Off-Policy Policy Evaluation

Study off-policy evaluation in partially observable environments, reducing bias and errors.

problem Bias and large errors in off-policy evaluation for partially observable environments.
method Defined and solved off-policy evaluation for POMDPs, introduced Decoupled POMDP model.
result Demonstrated and compared off-policy evaluation methods, showing benefits of new approach.

Paper improves bootstrapping for off-policy reinforcement learning inference.

problem Improving bootstrapping for off-policy reinforcement learning inference.
method Proposes a bootstrapping FQE method for off-policy statistical inference and a subsampling procedure to improve runtime.
result Asymptotically efficient and distributionally consistent bootstrapping FQE method for off-policy inference.

A new estimator improves off-policy evaluation in RL, outperforming existing methods.

problem Estimating performance of a new policy using historical data from a different policy.
method Doubly-robust estimator based on Targeted Maximum Likelihood Estimation, with variance reduction techniques.
result Our estimator uniformly outperforms existing methods across various RL environments and levels of model misspecification.

New method estimates state-action stationary distribution for better off-policy policy evaluation.

problem Accurately estimating state-action stationary distribution for off-policy policy evaluation.
method Estimated Mixture Policy (EMP) for state and state-action stationary distribution corrections.
result Empirical validation shows improved accuracy over state-of-the-art methods.

Paper addresses off-policy evaluation and learning with covariate shift.

problem Evaluating and training a new policy using historical data with a covariate shift.
method Derives efficiency bounds and proposes doubly robust estimators for OPE and OPL under covariate shift.
result Proposes estimators for off-policy evaluation and learning under covariate shift.

Improves off-policy evaluation weights for balanced state-action pair distribution.

problem Imbalance in importance sampling weights for off-policy evaluation of contextual bandits.
method Balanced Off-Policy Evaluation (B-OPE) method that minimizes imbalance to desired counterfactual distribution of state-action pairs.
result Experimental evidence shows B-OPE improves offline policy evaluation in both discrete and continuous action spaces.

Paper develops a new multi-agent reinforcement learning algorithm.

problem Improving policies in a network of communicating agents.
method Develops a multi-agent off-policy actor-critic algorithm using emphatic temporal difference learning.
result Proves convergence of the algorithm under linear function approximation.

New method improves off-policy policy evaluation by balancing policy values.

problem Accurate off-policy policy evaluation in reinforcement learning.
method Developed an MDP model with balanced representation to estimate both individual and average policy values.
result Substantially lower Mean Squared Error (MSE) in various benchmarks and a real-world HIV treatment simulation.

New estimator uses clustering to improve off-policy evaluation accuracy.

problem Improving off-policy evaluation accuracy when logging and evaluation policies differ.
method Proposes an estimator that shares information across similar contexts using clustering.
result Clustering contexts improves estimation accuracy, especially in deficient information settings.

Paper proposes a framework for reliable off-policy evaluation in reinforcement learning.

problem Quantifying uncertainty in off-policy estimates for safe deployment of target policies.
method Distributionally robust optimization for creating confidence bounds.
result Non-asymptotic and asymptotic guarantees for robust cumulative reward estimates.

Novel LSE estimator improves off-policy learning and evaluation.

problem High variance and poor performance with low-quality propensity scores and heavy-tailed reward distributions.
method Introduces a novel estimator based on the log-sum-exponential (LSE) operator.
result Achieves convergence rate of O(nε/(1+ε))O(n^{-ε/(1+ ε)}) for regret bounds.

Paper proposes a method to create more reliable confidence intervals for off-policy evaluations.

problem Creating reliable confidence intervals for off-policy evaluations.
method Proposes a deeply-debiasing procedure to construct efficient, robust, and flexible confidence intervals.
result Validated by theoretical results and numerical experiments, the method improves the reliability of off-policy evaluations.

DOLCE improves off-policy evaluation and learning by decomposing effects.

problem Bias in off-policy evaluation and learning due to policy mismatch.
method Uses lagged contexts and a moment-based training procedure to decompose and cancel bias.
result DOLCE achieves substantial improvements in off-policy evaluation and learning.

This paper optimizes off-policy evaluation in reinforcement learning with function approximation.

problem Estimating cumulative value of a new policy from logged data generated by an unknown policy.
method Regression-based fitted Q iteration method, equivalent to estimating conditional mean embedding of transition operator.
result The method is minimax-optimal, with nearly minimal estimation error.

This paper reviews off-policy evaluation methods in reinforcement learning.

problem Efficiency and accuracy of off-policy evaluation methods in reinforcement learning.
method Discussion of existing OPE methods, their statistical properties, and related research directions.
result Efficiency bounds and state-of-the-art OPE methods in reinforcement learning.

A new method for evaluating and selecting policies in contextual bandits improves confidence intervals and policy quality.

problem Evaluating and selecting policies in contextual bandits with logged data.
method Self-normalized Importance Weighting (SN) estimator with Efron-Stein tail inequality and multiplicative bias control.
result The method provides tighter confidence intervals and better policy selection compared to competitors.

Proposes a method for obtaining interval bounds in off-policy evaluation.

problem Provides provably correct upper and lower bounds for off-policy evaluation.
method Searches for the maximum and minimum values of the expected reward among Lipschitz Q-functions.
result Introduces a Lipschitz value iteration method to monotonically tighten interval bounds.

Paper reduces variance in infinite horizon off-policy evaluation with bias reduction.

problem High variance in infinite horizon off-policy evaluation.
method Doubly robust augmentation of Liu et al. (2018a) method using learned value function.
result Significant reduction in bias with higher accuracy when either density ratio or value function is accurate.

STAR framework reduces OPE variance by distilling complex problems into discrete ARPs.

problem High variance and bias in off-policy evaluation methods.
method STAR framework that includes various OPE estimators and leverages state abstraction.
result Predictions from ARPs estimated from off-policy data are asymptotically correct.

We propose and analyze an alternate approach to off-policy multi-step temporal difference learning, in which off-policy returns are corrected with the current Q-function in terms of rewards, rather than with the target policy in terms of transition probabilities. We prove that such approximate corrections are sufficien…

2016-02-16abs ↗pdf ↗

Study improves off-policy evaluation from non-i.i.d. bandit samples.

problem Improving off-policy evaluation from non-independent bandit samples.
method Constructing an estimator from a standardized martingale difference sequence.
result Proposed estimator performs better than existing methods.

Monotonic policy improvement and off-policy learning are two main desirable properties for reinforcement learning algorithms. In this paper, by lower bounding the performance difference of two policies, we show that the monotonic policy improvement is guaranteed from on- and off-policy mixture samples. An optimization …

2017-10-10abs ↗pdf ↗

Combines parametric and nonparametric models for better off-policy evaluation.

problem Improving off-policy evaluation in reinforcement learning.
method Mixture-of-experts approach combining parametric and nonparametric models.
result Mixture-based approach outperforms individual models and state-of-the-art estimators.

Study shows accurate OPE depends on calibrated behaviour policy models.

problem Estimating a behaviour policy for OPE when true policy is unknown.
method Empirical studies comparing parametric vs non-parametric models.
result Simple non-parametric k-nearest neighbors model produces better calibrated behaviour policy estimates.

New methods improve robust decision-making under uncertainty in off-policy evaluation.

problem Statistical uncertainty and causal considerations in off-policy evaluation.
method Marginal Ratio (MR) estimator, Conformal Off-Policy Prediction (COPP), causal bounds.
result Improved robustness and uncertainty quantification in off-policy decision-making.

Unified minimax value interval for off-policy evaluation and optimization.

problem Overcoming the exponential variance in off-policy evaluation and policy optimization.
method Unified minimax value interval using marginalized importance weights.
result Unified value interval with double robustness, valid when either value-function or importance-weight class is well specified.

Concept-driven OPE reduces variance in off-policy decision evaluation.

problem High variance in off-policy decision evaluation due to limited sample sizes.
method Integrating human-explainable concepts into OPE to reduce variance.
result Concept-based OPE estimators remain unbiased and reduce variance when concepts are known and predefined.

New metric for off-policy evaluation without models or importance sampling.

problem Evaluating deep RL models in real-world environments is costly and impractical.
method Framed as a PU classification problem, using Q-function as decision function.
result Metric reliably predicts policy performance in various scenarios.

Improved off-policy evaluation for MDPs with weak distributional overlap.

problem Evaluation of policies when target and data-collection distributions are not strongly overlapping.
method Truncated Doubly Robust (TDR) estimators for off-policy evaluation in MDPs under weak distributional overlap.
result TDR estimators can recover large-sample behavior and are consistent even when distribution ratios are not square-integrable.

New methods improve off-policy evaluation for survival outcomes with censoring.

problem Systematic underestimation of policy performance due to censoring bias in survival outcomes.
method Proposes IPCW-IPS and IPCW-DR to handle censoring bias in survival outcomes.
result The proposed methods are unbiased and achieve double robustness.

Paper tackles efficient off-policy evaluation in long-horizon settings.

problem Efficient off-policy evaluation in long-horizon settings with diminishing overlap.
method Derives efficiency bounds for OPE under Markovian and time-invariant structures, develops a new DRL estimator.
result DRL estimator provides efficient OPE even with just one dependent trajectory in time-invariant Markov decision processes.