Safe policy optimization using Gaussian process models.
problem Optimizing safe policies for task performance.
method Training a Gaussian process model to capture dynamics, ensuring safe policies only.
result Closed-form computation of error gradients and constraint violation probability.
New method learns policies without limiting to Gaussian distributions.
problem Limitations of Gaussian parameterization in policy learning.
method Advantage weighted quantile regression for implicit policy modeling.
result Comparable or superior performance on MuJoCo benchmarks.
EPG unifies SPG and DPG for reinforcement learning.
problem Improving reinforcement learning algorithms for policy optimization.
method EPG integrates across actions for gradient estimation, using practical results for Gaussian policies and extending to broader classes of policies.
result EPG reduces gradient variance without deterministic policies and with minimal overhead.
Bayesian approach uses Gaussian process for reinforcement learning.
problem Robotic locomotion environments
method Bayesian actor-critic, model-free reinforcement learning with Gaussian process for exploration and policy optimization.
result Gaussian process method outperforms current algorithms in robotic locomotion environments.
Proposes a Riemannian optimization for policy improvement in MDPs.
problem Optimizing policy functions in Markov decision processes (MDPs).
method Riemannian proximal optimization algorithm with Gaussian mixture model (GMM).
result Guaranteed convergence and efficacy demonstrated in preliminary experiments.
PFPN uses particle filtering to improve character control in physics-based simulations.
problem Premature commitment to suboptimal actions in high-dimensional continuous control problems for articulated characters.
method Proposes a particle-based action policy using particle filtering to dynamically explore and discretize the action space.
result Demonstrates better imitation performance and robustness to external perturbations compared to Gaussian policies.
RL solves discrete LQ control with Gaussian optimal policy.
problem Discrete-time linear-quadratic control problem.
method Entropy-based RL to find Gaussian optimal policy.
result RL algorithm solves mean-variance asset-liability management problem.
New algorithm learns Gaussian policies from corrective human feedback, outperforming current methods.
problem Learning from corrective human feedback for complex systems.
method Gaussian Process Coach (GPC) that uses Gaussian Processes and policy uncertainty for optimal feedback selection and learning rate adaptation.
result Demonstrated superior performance in OpenAI Gym benchmarks compared to COACH.
Proposes DGCN with trajectory sampling for data-efficient policy search in MBRL.
problem Improving data efficiency in model-based reinforcement learning.
method Combines trajectory sampling and DGCN for uncertainty propagation in probabilistic world models.
result Improves sample-efficiency over other uncertainty propagation methods and probabilistic models.
Study on variance of policy gradient in simple RL environments.
problem Understanding variance of policy gradient estimators in continuous RL.
method Analyzes REINFORCE estimator in linear-quadratic environments with Gaussian noise.
result Derives and validates bounds on estimator variance empirically.
RL approach to continuous-time MV portfolio selection with optimal policy being Gaussian.
problem Achieving optimal tradeoff between exploration and exploitation in continuous-time MV portfolio selection.
method Entropy-regularized, relaxed stochastic control problem; policy improvement theorem; RL algorithm.
result RL algorithm outperforms adaptive control and deep neural networks methods.
EPG unifies SPG and DPG, reducing variance and improving exploration.
problem Improving policy gradient methods for reinforcement learning.
method EPG integrates across actions for gradient estimation, reducing variance.
result EPG reduces variance without requiring deterministic policies.
A new method optimizes in nonstationary environments with many arms efficiently.
problem Optimizing in nonstationary environments with a large number of arms.
method Gaussian interpolation to learn continuous Lipschitz reward functions in nonstationary environments.
result Efficiently learns continuous Lipschitz reward functions with O ∗ ( T ) \mathcal{O}^*(\sqrt{T}) O ∗ ( T ) cumulative regret. Improved exploration in SAC using Normalizing Flows policies.
problem Brittleness and inefficiency of DRL algorithms in continuous action spaces.
method Introducing Normalizing Flow policies within the SAC framework to learn more expressive policies.
result Increased stability and better exploration in sparse reward settings.
Improved exploration in cooperative multi-agent reinforcement learning.
problem Limited expressiveness of Gaussian policies in DecSPG hinders effective exploration.
method Proposes decentralized diffusion policy learning (DDPL) with denoising diffusion probabilistic models.
result Consistently improved performance on various MARL benchmarks.
Improved statistical efficiency of Thompson Sampling for combinatorial semi-bandits.
problem Efficiency of policies in stochastic combinatorial multi-armed bandits with semi-bandit feedback.
method Analysis of Combinatorial Thompson Sampling (CTS) using Beta and Gaussian priors for mutually independent and multivariate sub-Gaussian outcomes.
result CTS provides an efficient policy with optimal asymptotic regret for both mutually independent and multivariate sub-Gaussian outcomes.
ARPs improve exploration and sample efficiency in continuous control tasks.
problem Limited exploration in continuous control tasks leading to low sample efficiency.
method Introduce autoregressive policies (ARPs) with temporally coherent standard normal distributions.
result ARPs enhance exploration and sample efficiency in both simulated and real-world domains.
Paper formulates mutual information optimal control for discrete-time systems.
problem Optimal control of discrete-time linear systems with mutual information.
method Formulates MIOCP as an extension of MEOCP, derives optimal policy and prior, proposes alternating minimization algorithm.
result Proposes an alternating minimization algorithm for MIOCP.
The paper analyzes the reward improvement of aligned policies in large language models.
problem Optimizing policies in large language models while staying close to a reference policy.
method Information-theoretic analysis and reduction to exponential order statistics.
result Information-theoretic upper bounds on reward improvement are derived.
Paper optimizes Bayesian optimization for complex functions with macro-actions.
problem Optimizing complex, highly uncertain functions efficiently.
method Generalized GP upper confidence bound with macro-actions for scalable lookahead.
result Asymptotically optimal anytime variant of epsilon-Macro-GPO policy.
The paper proves the convergence of Q-value for Gaussian rewards.
problem Existing proofs cannot guarantee convergence of the Q-function for Gaussian rewards.
method Using the central limit theorem and relaxing the condition to E [ r ( s , a ) 2 ] < ∞ E[r(s,a)^2]<\infty E [ r ( s , a ) 2 ] < ∞ . result Proves the convergence of the Q-function under the condition of E [ r ( s , a ) 2 ] < ∞ E[r(s,a)^2]<\infty E [ r ( s , a ) 2 ] < ∞ . Develops a new RL algorithm for medical treatment regimes.
problem Optimal dose determination in continuous action environments.
method Quasi-optimal learning algorithm for near-optimal actions.
result Guaranteed convergence and effectiveness in real applications.
Framework for sensitivity analysis in biomanufacturing processes.
problem High complexity and uncertainty in biomanufacturing processes.
method Shapley value estimation for linear and nonlinear pKG models, using quasi-Monte Carlo and antithetic sampling.
result Improved efficiency and accuracy in sensitivity analysis for biomanufacturing processes.
Bayesian framework for policy learning in decision problems.
problem Maximizing expected welfare in decision-making problems.
method Loss-based Bayesian updating and squared-loss surrogate for welfare maximization.
result General Bayes posterior over decision rules with Gaussian pseudo-likelihood interpretation.
Bayesian sOED uses PG reinforcement learning for efficient experiment design.
problem Optimizing sequential experiments for nonlinear models with limited data.
method Formulated as POMDP, solved via PG methods with neural network parameterization.
result Demonstrated advantages over batch and greedy designs in contaminant source inversion.
The paper interprets policy-gradient algorithms using continuation theory.
problem Optimizing nonconvex functions in reinforcement learning.
method Formulates policy optimization as optimization by continuation, interprets policy-gradient algorithms as implicitly optimizing deterministic policies.
result Exploration in policy-gradient algorithms is seen as computing a continuation of the return of the policy.
Develops methods for eliciting multiple, continuously valued treatment policies using causal inverse classification.
problem Tackles the problem of eliciting multiple, continuously valued treatment policies.
method Adopts a causal approach to inverse classification, developing the inverse classification potential outcomes framework (ICPOF) and approximate propensity score (APS).
result Demonstrates the viability of the methods on student performance.
Optimal policy found for observing noisy time series.
problem Minimizing posterior variance plus observation costs in discrete-time Gaussian random walks.
method Developed a simple threshold-based policy and proved its optimality.
result Simple threshold policy is optimal for observing noisy time series.
Data-driven RL solves Merton's expected utility problem via policy randomization.
problem Maximizing expected utility in an incomplete market with unknown primitives.
method Policy randomization in continuous-time reinforcement learning.
result RL algorithms solve Merton's problem without estimating model primitives.
New reinforcement learning algorithms improve policy optimization with entropy regularization.
problem Improving policy optimization in reinforcement learning.
method Soft policy gradient theorem (SPGT) and new policy optimization algorithms.
result New algorithms outperform prior works on various benchmark tasks.
New approach to portfolio optimization shows entropy regularization is ineffective.
problem Entropy regularization in mean-variance portfolio optimization under drift uncertainty.
method Combining Bayesian filtering and stochastic policy optimization.
result Entropy regularization does not accelerate learning about unknown drift.
Revisits PPO design choices, exposing failure modes and proposing alternatives.
problem Failure modes of standard PPO in new environments.
method Revisits standard PPO design choices, exposes failure modes, and proposes alternative approaches.
result Alternative design choices prevent failure modes in new environments.
The paper tackles estimating optimal policy value in linear bandits with general context distributions.
problem Estimating the optimal policy value in linear bandits with general context distributions.
method The paper provides lower bounds and an algorithm for sublinear estimation of V ∗ V^* V ∗ under stronger assumptions. result A practical algorithm that estimates a problem-dependent upper bound on V ∗ V^* V ∗ with O ~ ( d ) \widetilde{\mathcal{O}}(\sqrt{d}) O ( d ) samples. Bayesian Markowitz portfolio problem shows entropy regularization is ineffective.
problem Entropy regularization in Bayesian Markowitz portfolio optimization.
method Combines continuous-time Bayesian filtering with stochastic policy optimization.
result Entropy regularization does not accelerate learning of unknown drift.
UK's rapid vaccine rollout linked to reduced COVID-19 mortality.
problem Assessing the impact of accelerated vaccine rollout on public health outcomes.
method Flexible probabilistic models combining interrupted time series analysis and synthetic control methods with multi-output Gaussian processes.
result Substantial reduction in COVID-19 mortality with little effect on transmission rates.
This work optimizes RL algorithms using entropy regularisation for continuous-time LQ problems.
problem Designing RL algorithms to balance exploration and exploitation in noisy environments.
method Entropy regularisation in two formulations: exploratory control and proximal policy update.
result Regret of O ( N ) \mathcal{O}(\sqrt{N}) O ( N ) for both learning algorithms over N N N episodes. The study analyzes how neural reward models learn features for policy optimization in a Gaussian single-index model.
problem Reward modeling in policy optimization and its impact on downstream value.
method Two-stage neural reward model: first learns hidden direction, then fits readout layer.
result For any feature-learning temperature above a dimension-free threshold, a constant fraction of neurons recover the hidden direction.
New approach learns walk and trot gaits from simulated quadruped using strategic exploration.
problem Learning symmetric gaits (walk and trot) from high-dimensional action spaces.
method Introduced symmetry properties into initial covariance of Gaussian search distribution for strategic exploration. Used episode-based likelihood ratio policy gradient and relative entropy policy search.
result Significant performance enhancement in learning walk and trot gaits compared to random gaits.
New algorithm reduces worst-case regret for heavy-tailed bandits.
problem Stochastic Multi-Armed Bandit problem with heavy-tailed rewards.
method Modified minimax policy MOSS with saturated empirical mean.
result Worst-case regret matching lower bound for heavy-tailed distributions.
A novel approach uses an ensemble of Gaussian processes for robust and adaptive reinforcement learning.
problem Adaptive reinforcement learning in large or continuous state spaces.
method Online scalable (OS) approach with a weighted ensemble of Gaussian processes.
result The ensemble approach improves performance in adversarial settings.
A method for accurate pricing of multidimensional derivatives under uncertain volatility.
problem High-dimensional stochastic control problem in uncertain volatility model.
method Backward actor-critic stochastic policy gradient scheme combining DP, PPO, and neural networks.
result Accurate and efficient pricing of multidimensional derivatives compared to benchmarks.
The paper analyzes how employers can efficiently screen candidates using multiple tests, considering both skill estimation and fairness.
problem How to efficiently screen candidates using multiple noisy signals without violating fairness.
method The paper extends traditional screening models to a multi-test setting, analyzing optimal employer policies for both fixed and dynamic test assignments.
result A fundamental impossibility emerges when noise levels vary across groups, making it impossible to administer the same number of tests and maintain the same outcomes.
This paper presents a novel nonmyopic adaptive Gaussian process planning (GPP) framework endowed with a general class of Lipschitz continuous reward functions that can unify some active learning/sensing and Bayesian optimization criteria and offer practitioners some flexibility to specify their desired choices for defi…
Survey explores geometric aspects of policy optimization in control systems.
problem Understanding the geometric relationships between control design and optimization.
method Geometric perspective on policy optimization, focusing on parameterization and topology.
result Implications of policy geometry on stability and performance of local search algorithms.
The paper develops a safe policy gradient algorithm for reinforcement learning.
problem Safety issues in reinforcement learning for real-world control tasks.
method Stochastic optimization perspective, meta-parameter schedules, adaptive selection of step size and batch size.
result Monotonic improvement guarantees for a wide class of parametric policies.
The Knowledge Gradient (KG) policy was originally proposed for online ranking and selection problems but has recently been adapted for use in online decision making in general and multi-armed bandit problems (MABs) in particular. We study its use in a class of exponential family MABs and identify weaknesses, including …
A new method improves policy evaluation in RL by tracking value uncertainties.
problem Limitations in existing policy evaluation methods for deep RL tasks.
method KOVA (Kalman Optimization for Value Approximation) based on extended Kalman filter.
result KOVA minimizes a regularized objective function that considers parameter and noisy return uncertainties.
RBI improves RL policies by reducing evaluation errors and avoiding performance degradation.
problem Evaluation errors in Q-function learning cause policy improvement penalties.
method RBI attenuates low-probability actions to minimize improvement penalties.
result RBI reduces regret and improves data efficiency in RL tasks.