Policy-gradient method controls multiple non-cohesive targets.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Framework improves policy generalizability under biased training data.
Study designs logging policies to minimize off-policy evaluation error.
We observe that several existing policy gradient methods (such as vanilla policy gradient, PPO, A2C) may suffer from overly large gradients when the current policy is close to deterministic (even in some very simple environments), leading to an unstable training process. To address this issue, we propose a new method, …
In this paper, we present a new approach to Transfer Learning (TL) in Reinforcement Learning (RL) for cross-domain tasks. Many of the available techniques approach the transfer architecture as a method of speeding up the target task learning. We propose to adapt and reuse the mapped source task optimal-policy directly …
The paper targets optimal interventions for long-term outcomes using imputed data and policy learning.
Off-policy learning exhibits greater instability when compared to on-policy learning in reinforcement learning (RL). The difference in probability distribution between the target policy () and the behavior policy (b) is a major cause of instability. High variance also originates from distributional mismatch. The var…
RL policy tracks dynamic targets in partially known environments robustly.
Proposes a method to improve treatment policies in data-scarce clinical settings.
Estimates personalized policies robust to shifts in target populations.
New model-based methods adapt pre-trained policies to unseen environments efficiently.
Transfer Learning (TL) has shown great potential to accelerate Reinforcement Learning (RL) by leveraging prior knowledge from past learned policies of relevant tasks. Existing transfer approaches either explicitly computes the similarity between tasks or select appropriate source policies to provide guided explorations…
Develops methods to estimate and quantify uncertainty in off-policy evaluation.
We propose and analyze an alternate approach to off-policy multi-step temporal difference learning, in which off-policy returns are corrected with the current Q-function in terms of rewards, rather than with the target policy in terms of transition probabilities. We prove that such approximate corrections are sufficien…
Paper proposes a framework for reliable off-policy evaluation in reinforcement learning.
Combines trial and observational data to improve policy evaluation.
Paper proposes a method to create more reliable confidence intervals for off-policy evaluations.
New method stabilizes FQE by reweighting Bellman targets.
RL approach for target tracking with unknown dynamics and sensor control.
Presents SPEED, an algorithm for optimal policy evaluation in linear bandits with heteroscedastic noise.
The paper tackles fair policy targeting by optimizing allocation rules to minimize unfairness.
Scarce data is a major challenge to scaling robot learning to truly complex tasks, as we need to generalize locally learned policies over different task contexts. Contextual policy search offers data-efficient learning and generalization by explicitly conditioning the policy on a parametric context space. In this paper…
Scarce data is a major challenge to scaling robot learning to truly complex tasks, as we need to generalize locally learned policies over different "contexts". Bayesian optimization approaches to contextual policy search (CPS) offer data-efficient policy learning that generalize over a context space. We propose to impr…
Domain randomization (DR) is a successful technique for learning robust policies for robot systems, when the dynamics of the target robot system are unknown. The success of policies trained with domain randomization however, is highly dependent on the correct selection of the randomization distribution. The majority of…
New split rules improve subpopulation targeting in policy-making.
Data augmentation (DA) has been widely utilized to improve generalization in training deep neural networks. Recently, human-designed data augmentation has been gradually replaced by automatically learned augmentation policy. Through finding the best policy in well-designed search space of data augmentation, AutoAugment…
A drone catches another agile drone using competitive reinforcement learning.
Generically learns movement control policies from exploration data.
New method for estimating value of optimal policies in uncertain scenarios.
Learning the value function of a given policy (target policy) from the data samples obtained from a different policy (behavior policy) is an important problem in Reinforcement Learning (RL). This problem is studied under the setting of off-policy prediction. Temporal Difference (TD) learning algorithms are a popular cl…
Study shows how macroprudential policies affect credit growth in Israel, especially in housing and business sectors.
We study the problem of off-policy evaluation (OPE) in Reinforcement Learning (RL), where the aim is to estimate the performance of a new policy given historical data that may have been generated by a different policy, or policies. In particular, we introduce a novel doubly-robust estimator for the OPE problem in RL, b…
This paper bridges the gap between theoretical and practical OPE for bandit problems.
Paper improves policy updates in reinforcement learning to speed up learning.
Recent policy optimization approaches (Schulman et al., 2015a; 2017) have achieved substantial empirical successes by constructing new proxy optimization objectives. These proxy objectives allow stable and low variance policy learning, but require small policy updates to ensure that the proxy objective remains an accur…
DDO-RM improves reward-based policies by converting reward scores into a target distribution.
Transfer reinforcement learning (RL) aims at improving the learning efficiency of an agent by exploiting knowledge from other source agents trained on relevant tasks. However, it remains challenging to transfer knowledge between different environmental dynamics without having access to the source environments. In this …
The paper discusses the role of monetary policy when potential output depends on the inflation rate. If the intention of the central bank is to maximize actual output growth, then it has to be credibly committed to a strict inflation targeting rule, and to take the MOGIR (the Maximizing Output Growth Inflation Rate) as…
A new method for evaluating and selecting policies in contextual bandits improves confidence intervals and policy quality.
The paper analyzes the sample complexities for policy evaluation with linear function approximation.
VA-OPE improves OPE by incorporating variance information, achieving tighter error bounds.
Attackers can poison environments to force RL agents to follow target policies.
This paper studies the problem of optimally allocating treatments in the presence of spillover effects, using information from a (quasi-)experiment. I introduce a method that maximizes the sample analog of average social welfare when spillovers occur. I construct semi-parametric welfare estimators with known and unknow…
Computer simulation provides an automatic and safe way for training robotic control policies to achieve complex tasks such as locomotion. However, a policy trained in simulation usually does not transfer directly to the real hardware due to the differences between the two environments. Transfer learning using domain ra…
Paper develops IV method for consistent OPE in confounded MDPs.
This paper studies the statistical theory of batch data reinforcement learning with function approximation. Consider the off-policy evaluation problem, which is to estimate the cumulative value of a new target policy from logged history generated by unknown behavioral policies. We study a regression-based fitted Q iter…
WAPPO optimizes feature distributions for better visual transfer in RL.
The problem of on-line off-policy evaluation (OPE) has been actively studied in the last decade due to its importance both as a stand-alone problem and as a module in a policy improvement scheme. However, most Temporal Difference (TD) based solutions ignore the discrepancy between the stationary distribution of the beh…