Adapts model-based advice to stabilize black-box policies for nonlinear control.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Policy gradient converges to globally optimal policy in nearly linear-quadratic systems.
The paper provides a non-asymptotic error bound for linear system identification under nonlinear policies.
The paper automates policy learning for nonlinear welfare criteria using machine learning and debiasing techniques.
KCRL learns stable policies for nonlinear systems with formal guarantees.
This paper analyzes a simplified strategy for nonlinear control using local linear models and iLQR updates.
This paper addresses the model-free nonlinear optimal problem with generalized cost functional, and a data-based reinforcement learning technique is developed. It is known that the nonlinear optimal control problem relies on the solution of the Hamilton-Jacobi-Bellman (HJB) equation, which is a nonlinear partial differ…
This paper studies GAIL's global convergence for general MDP and nonlinear rewards.
We present a data-efficient reinforcement learning algorithm resistant to observation noise. Our method extends the highly data-efficient PILCO algorithm (Deisenroth & Rasmussen, 2011) into partially observed Markov decision processes (POMDPs) by considering the filtering process during policy evaluation. PILCO conduct…
Breaks down complex nonlinear dynamics into simpler components.
New method learns optimal policies in presence of unmeasured confounders.
Bayesian sOED uses PG reinforcement learning for efficient experiment design.
The control design problem is considered for nonlinear systems with unknown internal system model. It is known that the nonlinear control problem can be transformed into solving the so-called Hamilton-Jacobi-Isaacs (HJI) equation, which is a nonlinear partial differential equation that is genera…
This work explains RL policies using causal models, revealing important patterns and failures.
This paper uses NLDT to find interpretable control rules from complex DRL policies.
New method learns nonlinear systems from single finite trajectory samples.
Optimizes exploration for nonlinear systems to learn controllers efficiently.
RichID learns optimal control policies from nonlinear observations.
Mixed RL improves RL efficiency with dual representations.
Learning policies that generalize across multiple tasks is an important and challenging research topic in reinforcement learning and robotics. Training individual policies for every single potential task is often impractical, especially for continuous task variations, requiring more principled approaches to share and t…
This paper analyzes the sample complexity of two timescale reinforcement learning algorithms.
This work overcomes bias in concave multi-objective reinforcement learning.
Softmax policy gradient achieves global optimality in wide neural networks with entropy regularization.
A key problem in reinforcement learning for control with general function approximators (such as deep neural networks and other nonlinear functions) is that, for many algorithms employed in practice, updates to the policy or -function may fail to improve performance---or worse, actually cause the policy performance …
There is by now a large consensus in modern monetary policy. This consensus has been built upon a dynamic general equilibrium model of optimal monetary policy as developed by, e.g., Goodfriend and King (1997), Clarida et al. (1999), Svensson (1999) and Woodford (2003). In this paper we extend the standard optimal monet…
This paper tackles data-efficient nonlinear control in Hamiltonian systems using symplectic geometry.
Adaptive optimal control of nonlinear dynamic systems with deterministic and known dynamics under a known undiscounted infinite-horizon cost function is investigated. Policy iteration scheme initiated using a stabilizing initial control is analyzed in solving the problem. The convergence of the iterations and the optim…
We present GradientDICE for estimating the density ratio between the state distribution of the target policy and the sampling distribution in off-policy reinforcement learning. GradientDICE fixes several problems of GenDICE (Zhang et al., 2020), the state-of-the-art for estimating such density ratios. Namely, the optim…
Bayesian neural networks improve macroeconomic forecasting and model nonlinearities.
Reinforcement learning is a promising approach to synthesizing policies for challenging robotics tasks. A key problem is how to ensure safety of the learned policy---e.g., that a walking robot does not fall over or that an autonomous car does not run into an obstacle. We focus on the setting where the dynamics are know…
DFPV improves PCL for confounded bandit policy evaluation.
We propose a new algorithm for solving parabolic partial differential equations (PDEs) and backward stochastic differential equations (BSDEs) in high dimension, by making an analogy between the BSDE and reinforcement learning with the gradient of the solution playing the role of the policy function, and the loss functi…
Unified analysis for nonlinear parametric models in Bayesian optimization.
The question of how to explore, i.e., take actions with uncertain outcomes to learn about possible future rewards, is a key question in reinforcement learning (RL). Here, we show a surprising result: We show that Q-learning with nonlinear Q-function and no explicit exploration (i.e., a purely greedy policy) can learn s…
NK bandits improve performance on nonlinear tasks.
We study the safe reinforcement learning problem with nonlinear function approximation, where policy optimization is formulated as a constrained optimization problem with both the objective and the constraint being nonconvex functions. For such a problem, we construct a sequence of surrogate convex constrained optimiza…
The goal of this paper is to present a method for simultaneous trajectory and local stabilizing policy optimization to generate local policies for trajectory-centric model-based reinforcement learning (MBRL). This is motivated by the fact that global policy optimization for non-linear systems could be a very challengin…
In reinforcement learning, temporal difference (TD) is the most direct algorithm to learn the value function of a policy. For large or infinite state spaces, exact representations of the value function are usually not available, and it must be approximated by a function in some parametric family. However, with \emph{no…
Robust model predictive control (MPC) is a well-known control technique for model-based control with constraints and uncertainties. In classic robust tube-based MPC approaches, an open-loop control sequence is computed via periodically solving an online nominal MPC problem, which requires prior model information and fr…
This paper proposes a method for estimating the effect of a policy intervention on an outcome over time. We train recurrent neural networks (RNNs) on the history of control unit outcomes to learn a useful representation for predicting future outcomes. The learned representation of control units is then applied to the t…
Payments data and machine learning improve nowcasting accuracy for macroeconomic indicators.
Compressed imitation learning uses simplicity priors for efficient expert behavior copying.
Framework for sensitivity analysis in biomanufacturing processes.
In this paper we discuss policy iteration methods for approximate solution of a finite-state discounted Markov decision problem, with a focus on feature-based aggregation methods and their connection with deep reinforcement learning schemes. We introduce features of the states of the original problem, and we formulate …
We present the first provably convergent two-timescale off-policy actor-critic algorithm (COF-PAC) with function approximation. Key to COF-PAC is the introduction of a new critic, the emphasis critic, which is trained via Gradient Emphasis Learning (GEM), a novel combination of the key ideas of Gradient Temporal Differ…
We extend Bayesian multi-armed bandit (MAB) algorithms beyond their original setting by making use of sequential Monte Carlo (SMC) methods. A MAB is a sequential decision making problem where the goal is to learn a policy that maximizes long term payoff, where only the reward of the executed action is observed. In the …
DFIV uses deep neural nets to learn nonlinear features in IV regression.
We introduce a methodology for efficiently computing a lower bound to empowerment, allowing it to be used as an unsupervised cost function for policy learning in real-time control. Empowerment, being the channel capacity between actions and states, maximises the influence of an agent on its near future. It has been sho…