Large deviations theory applied to policy gradient methods.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Bayesian Parametric Portfolio Policies corrects overestimation of utility and risk in traditional PPP.
We consider off-policy evaluation and optimization with continuous action spaces. We focus on observational data where the data collection policy is unknown and needs to be estimated. We take a semi-parametric approach where the value function takes a known parametric form in the treatment, but we are agnostic on how i…
Study evaluates policies in partially observable environments without full model specification.
Optimistic actor-critic tackles linear MDPs with parametric policies.
In this work, we consider the problem of estimating a behaviour policy for use in Off-Policy Policy Evaluation (OPE) when the true behaviour policy is unknown. Via a series of empirical studies, we demonstrate how accurate OPE is strongly dependent on the calibration of estimated behaviour policy models: how precisely …
Paper proposes a policy-search algorithm to learn entropy-maximizing exploration policies in reward-free environments.
Recent policy optimization approaches have achieved substantial empirical success by constructing surrogate optimization objectives. The Approximate Policy Iteration objective (Schulman et al., 2015a; Kakade and Langford, 2002) has become a standard optimization target for reinforcement learning problems. Using this ob…
The paper tackles robust policy learning in MDPs using statistical methods.
We present an off-policy actor-critic algorithm for Reinforcement Learning (RL) that combines ideas from gradient-free optimization via stochastic search with learned action-value function. The result is a simple procedure consisting of three steps: i) policy evaluation by estimating a parametric action-value function;…
The paper sets bounds on how much regret is unavoidable in adaptive LQR with unknown B-matrix.
The paper explains why estimating a history-dependent policy can reduce MSE in reinforcement learning.
New findings show optimization is crucial for OPL in large action spaces.
We study a policy gradient method with L2 regularization for MAB problems.
We consider a model-based approach to perform batch off-policy evaluation in reinforcement learning. Our method takes a mixture-of-experts approach to combine parametric and non-parametric models of the environment such that the final value estimate has the least expected error. We do so by first estimating the local a…
Optimal learning for parametric prophet inequalities with exponential-type distributions
Policy gradient methods are among the most effective methods in challenging reinforcement learning problems with large state and/or action spaces. However, little is known about even their most basic theoretical convergence properties, including: if and how fast they converge to a globally optimal solution or how they …
MoMA improves model-based RL by using unrestricted policy classes.
Reinforcement learning with sparse rewards is challenging because an agent can rarely obtain non-zero rewards and hence, gradient-based optimization of parameterized policies can be incremental and slow. Recent work demonstrated that using a memory buffer of previous successful trajectories can result in more effective…
Quantum algorithms speed up reinforcement learning policies in large state-action spaces.
We consider a dynamic pricing problem for repeated contextual second-price auctions with multiple strategic buyers who aim to maximize their long-term time discounted utility. The seller has limited information on buyers' overall demand curves which depends on a non-parametric market-noise distribution, and buyers may …
Policy gradient based reinforcement learning algorithms coupled with neural networks have shown success in learning complex policies in the model free continuous action space control setting. However, explicitly parameterized policies are limited by the scope of the chosen parametric probability distribution. We show t…
Learning policies that generalize across multiple tasks is an important and challenging research topic in reinforcement learning and robotics. Training individual policies for every single potential task is often impractical, especially for continuous task variations, requiring more principled approaches to share and t…
Study dynamic pricing with semi-parametric models to minimize regret.
New algorithm reduces dynamic regret for noisy gradient feedback with piecewise polynomial comparators.
Policy evaluation or value function or Q-function approximation is a key procedure in reinforcement learning (RL). It is a necessary component of policy iteration and can be used for variance reduction in policy gradient methods. Therefore its quality has a significant impact on most RL algorithms. Motivated by manifol…
Develops ODRPO to improve RL algorithms with better performance and stability.
Paper tackles efficient policy gradient estimation from off-policy data.
Develops neural network framework for risk-reward optimization problems.
The paper tackles performative policy learning with strategic agents, improving scalability and generalizability.
Dynamic pricing policy converges to Nash equilibrium with low regret.
Develops confidence bounds for off-policy evaluation in contextual bandits.
We study the problem of identifying the policy space of a learning agent, having access to a set of demonstrations generated by its optimal policy. We introduce an approach based on statistical testing to identify the set of policy parameters the agent can control, within a larger parametric policy space. After present…
We identify a fundamental problem in policy gradient-based methods in continuous control. As policy gradient methods require the agent's underlying probability distribution, they limit policy representation to parametric distribution classes. We show that optimizing over such sets results in local movement in the actio…
Develops CLTs for Markov chain transition probabilities and policies.
New method for estimating counterfactual means in adaptive experiments.
Deep Reinforcement Learning (DRL) algorithms for continuous action spaces are known to be brittle toward hyperparameters as well as \cut{being}sample inefficient. Soft Actor Critic (SAC) proposes an off-policy deep actor critic algorithm within the maximum entropy RL framework which offers greater stability and empiric…
Study optimal adjustment sets for causal policies with hidden variables.
New framework for PMD convergence in non-tabular environments.
Efficient policy learning from observational data using weighted classification reductions.
Framework improves policy generalizability under biased training data.
Proximal policy optimization and trust region policy optimization (PPO and TRPO) with actor and critic parametrized by neural networks achieve significant empirical success in deep reinforcement learning. However, due to nonconvexity, the global convergence of PPO and TRPO remains less understood, which separates theor…
We accelerate PMD algorithms for reinforcement learning using functional methods.
Scarce data is a major challenge to scaling robot learning to truly complex tasks, as we need to generalize locally learned policies over different task contexts. Contextual policy search offers data-efficient learning and generalization by explicitly conditioning the policy on a parametric context space. In this paper…
In reinforcement learning, temporal difference (TD) is the most direct algorithm to learn the value function of a policy. For large or infinite state spaces, exact representations of the value function are usually not available, and it must be approximated by a function in some parametric family. However, with \emph{no…
FOCOPS optimizes agent's behavior while adhering to constraints.
We propose a novel framework for multi-task reinforcement learning (MTRL). Using a variational inference formulation, we learn policies that generalize across both changing dynamics and goals. The resulting policies are parametrized by shared parameters that allow for transfer between different dynamics and goal condit…
KL-regularized RL from expert demos can lead to slow, unstable learning.