Paper enhances RL policies using trust region optimization for offline data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Softmax policy gradient methods converge at rate with constants depending on problem and initialization.
Adaptive optimal control of nonlinear dynamic systems with deterministic and known dynamics under a known undiscounted infinite-horizon cost function is investigated. Policy iteration scheme initiated using a stabilizing initial control is analyzed in solving the problem. The convergence of the iterations and the optim…
Study how untrained policies explore in RL environments.
This paper studies how gradient descent in control systems can perform well on unseen data.
Study finds dividend policy has no significant effect on IPO stock prices.
In quadruped gait learning, policy search methods that scale high dimensional continuous action spaces are commonly used. In most approaches, it is necessary to introduce prior knowledge on the gaits to limit the highly non-convex search space of the policies. In this work, we propose a new approach to encode the symme…
Adaptive optimal control using value iteration initiated from a stabilizing control policy is theoretically analyzed in terms of stability of the system during the learning stage without ignoring the effects of approximation errors. This analysis includes the system operated using any single/constant resulting control …
Adaptive optimal control using value iteration (VI) initiated from a stabilizing policy is theoretically analyzed in various aspects including the continuity of the result, the stability of the system operated using any single/constant resulting control policy, the stability of the system operated using the evolving/ti…
Meta-reinforcement learning improves fault-adaptive control efficiency.
We consider the generic approach of using an experience memory to help exploration by adapting a restart distribution. That is, given the capacity to reset the state with those corresponding to the agent's past observations, we help exploration by promoting faster state-space coverage via restarting the agent from a mo…
We propose an algorithm for deterministic continuous Markov Decision Processes with sparse rewards that computes the optimal policy exactly with no dependency on the size of the state space. The algorithm has time complexity of and memory complexity of , where is the…
We present a method for a certain class of Markov Decision Processes (MDPs) that can relate the optimal policy back to one or more reward sources in the environment. For a given initial state, without fully computing the value function, q-value function, or the optimal policy the algorithm can determine which rewards w…
We study the problem of learning a good search policy for combinatorial search spaces. We propose retrospective imitation learning, which, after initial training by an expert, improves itself by learning from \textit{retrospective inspections} of its own roll-outs. That is, when the policy eventually reaches a feasible…
Softmax PG methods can take extremely long to converge, even with exact gradients.
Softmax policy gradient achieves global optimality in wide neural networks with entropy regularization.
Model-based reinforcement learning (MBRL) with model-predictive control or online planning has shown great potential for locomotion control tasks in terms of both sample efficiency and asymptotic performance. Despite their initial successes, the existing planning methods search from candidate sequences randomly generat…
Perry uses auxiliary data to estimate RL policy values with confidence intervals.
A new estimator for evaluating policies in unknown environments.
PBVFs generalize across policies using learned value functions.
Develops methods to estimate and quantify uncertainty in off-policy evaluation.
The policy gradient theorem is defined based on an objective with respect to the initial distribution over states. In the discounted case, this results in policies that are optimal for one distribution over initial states, but may not be uniformly optimal for others, no matter where the agent starts from. Furthermore, …
New strategies for identifying the best arm in bandits with decreasing variances.
Policy gradient converges to globally optimal policy in nearly linear-quadratic systems.
We tackle the Multi-task Batch Reinforcement Learning problem. Given multiple datasets collected from different tasks, we train a multi-task policy to perform well in unseen tasks sampled from the same distribution. The task identities of the unseen tasks are not provided. To perform well, the policy must infer the tas…
This paper addresses the question of how a previously available control policy can be used as a supervisor to more quickly and safely train a new learned control policy for a robot. A weighted average of the supervisor and learned policies is used during trials, with a heavier weight initially on the superv…
Gradient-based methods are often used for policy optimization in deep reinforcement learning, despite being vulnerable to local optima and saddle points. Although gradient-free methods (e.g., genetic algorithms or evolution strategies) help mitigate these issues, poor initialization and local optima are still concerns …
Efficient RL algorithm for MDPs with linear realizability, achieving optimal regret bound.
Scarce data is a major challenge to scaling robot learning to truly complex tasks, as we need to generalize locally learned policies over different task contexts. Contextual policy search offers data-efficient learning and generalization by explicitly conditioning the policy on a parametric context space. In this paper…
Imitation learning, followed by reinforcement learning algorithms, is a promising paradigm to solve complex control tasks sample-efficiently. However, learning from demonstrations often suffers from the covariate shift problem, which results in cascading errors of the learned policy. We introduce a notion of conservati…
Scarce data is a major challenge to scaling robot learning to truly complex tasks, as we need to generalize locally learned policies over different "contexts". Bayesian optimization approaches to contextual policy search (CPS) offer data-efficient policy learning that generalize over a context space. We propose to impr…
Policy gradient methods with actor-critic schemes demonstrate tremendous empirical successes, especially when the actors and critics are parameterized by neural networks. However, it remains less clear whether such "neural" policy gradient methods converge to globally optimal policies and whether they even converge at …
We study a policy gradient method with L2 regularization for MAB problems.
Improving the sample efficiency in reinforcement learning has been a long-standing research problem. In this work, we aim to reduce the sample complexity of existing policy gradient methods. We propose a novel policy gradient algorithm called SRVR-PG, which only requires episodes to find an -approxima…
A new approach to fine-tuning LLMs with human feedback.
Deep reinforcement learning has been successful in a variety of tasks, such as game playing and robotic manipulation. However, attempting to learn \textit{tabula rasa} disregards the logical structure of many domains as well as the wealth of readily available knowledge from domain experts that could help "warm start" t…
VLBM learns MDP transitions from limited data, improving OPE performance.
This paper offers a financial economic perspective on the optimal time (and age) at which the owner of a Variable Annuity (VA) policy with a Guaranteed Living Withdrawal Benefit (GLWB) rider should initiate guaranteed lifetime income payments. We abstract from utility, bequest and consumption preference issues by treat…
The task of multi-step ahead prediction in language models is challenging considering the discrepancy between training and testing. At test time, a language model is required to make predictions given past predictions as input, instead of the past targets that are provided during training. This difference, known as exp…
Policy gradient converges linearly with Hadamard parameterization in tabular settings.
A novel approach learns goal-conditioned policies for locomotion using batch RL.
This study uses OPE methods to quickly assess auction policies.
Lapse-supported life insurance exacerbates adverse selection risks.
We introduce a novel apprenticeship learning algorithm to learn an expert's underlying reward structure in off-policy model-free \emph{batch} settings. Unlike existing methods that require a dynamics model or additional data acquisition for on-policy evaluation, our algorithm requires only the batch data of observed ex…
Study on adversarial training's impact on deep neural reinforcement learning policies.
AER dynamically adjusts entropy regularization for better LLM reinforcement learning.
Three approaches learn personalized treatment policies for UTI patients.
Optimizes dividend policies in a Brownian model with controlled rates.