Adapts model-based advice to stabilize black-box policies for nonlinear control.
problem Stabilizing machine-learned policies for nonlinear control with limited model information.
method Proposes an adaptive λ-confident policy to combine black-box and model-based advice. result Proves the stability of the adaptive λ-confident policy and its competitive ratio. In this work, we explore how probabilistic programs can be used to represent policies in sequential decision problems. In this formulation, a probabilistic program is a black-box stochastic simulator for both the problem domain and the agent. We relate classic policy gradient techniques to recently introduced black-box…
Paper introduces CageBO for optimizing complex public policy problems.
problem Complex decision-making and implicit constraints in public policy.
method CageBO framework using conditional variational autoencoder.
result CageBO outperforms baselines in optimizing large-scale police redistricting.
Bringing transparency to black-box decision making systems (DMS) has been a topic of increasing research interest in recent years. Traditional active and passive approaches to make these systems transparent are often limited by scalability and/or feasibility issues. In this paper, we propose a new notion of black-box D…
This paper uses NLDT to find interpretable control rules from complex DRL policies.
problem Complex, non-interpretable policies from black-box AI methods.
method Evolutionary optimization of NLDT for hierarchical control rules.
result Interpretable control rules with similar performance to black-box DRL.
MAPS and MAPS-SE learn from multiple suboptimal experts to improve policies efficiently.
problem Learning from multiple suboptimal experts in reinforcement learning.
method Active policy improvement from multiple black-box oracles.
result MAPS and MAPS-SE achieve sample efficiency and state-wise imitation learning.
We derive an optimal policy for adaptively restarting a randomized algorithm, based on observed features of the run-so-far, so as to minimize the expected time required for the algorithm to successfully terminate. Given a suitable Bayesian prior, this result can be used to select the optimal black-box optimization algo…
The most data-efficient algorithms for reinforcement learning in robotics are model-based policy search algorithms, which alternate between learning a dynamical model of the robot and optimizing a policy to maximize the expected return given the model and its uncertainties. Among the few proposed approaches, the recent…
System interprets complex treatment effects for personalized policies.
problem Complex, hard-to-understand treatment effect models.
method Scalable, interpretable personalized experimentation system.
result Learned explanations and generated interpretable policies.
Paper proposes interpretable RL policies from a mixture of experts.
problem Making RL policies transparent and understandable in real-world applications.
method Policy iteration scheme with interpretable experts and prototypical states.
result Proposed algorithm learns policies comparable to neural networks but more interpretable.
Minimalistic attacks reveal deep RL policies' vulnerabilities with little perturbation.
problem Tackling the vulnerability of deep reinforcement learning policies to minimal perturbations.
method Three key settings: black-box policy access, fractional-state adversary, and tactically-chanced attack. Formulated adversarial attacks on six Atari games.
result Deep RL policies can be significantly fooled by minimal perturbations, even in 0.01% of the input state.
This paper develops explainable treatment policies for RPM using clinical knowledge.
problem Barriers to adoption of DHIs and lack of interpretability in purely black-box algorithms.
method Developed a pipeline for learning explainable treatment policies using clinician-informed representations.
result Policies learned from clinician-informed representations are more efficacious and efficient than black-box policies.
Embedding RL policies in RKHS for robustness and theoretical guarantees.
problem Stability and theoretical guarantees in RL policy representation.
method Low-dimensional embedding of RL policies in RKHS.
result Embedded policies maintain high return with strong theoretical guarantees.
Novel approach simplifies VI problems with faster performance.
problem Black-box VI optimization problems.
method Sample Average Approximation (SAA) combined with quasi-Newton methods and line search.
result Achieves faster performance than existing methods.
Genetic algorithms have been widely used in many practical optimization problems. Inspired by natural selection, operators, including mutation, crossover and selection, provide effective heuristics for search and black-box optimization. However, they have not been shown useful for deep reinforcement learning, possibly …
Improved MORE algorithm reduces regret in black-box optimization and RL tasks.
problem Noisy fitness evaluations and poor sample quality in black-box optimization.
method Decouples mean and covariance updates, uses entropy scheduling, and simplifies model learning.
result Significantly reduces regret in black-box optimization and RL tasks.
New symmetries improve meta-reinforcement learning's generalization.
problem Black-box meta reinforcement learning struggles with generalization.
method Introduce symmetries to a black-box meta RL system.
result Incorporating symmetries improves generalization to new environments.
Evolutionary Strategies optimize hyper-parameters for off-policy learning.
problem Hyper-parameter sensitivity in off-policy learning.
method Application of Evolutionary Strategies for online hyper-parameter tuning.
result Our method outperforms state-of-the-art baselines.
Deep reinforcement learning has led to several recent breakthroughs, though the learned policies are often based on black-box neural networks. This makes them difficult to interpret and to impose desired specification constraints during learning. We present an iterative framework, MORL, for improving the learned polici…
Deep RL policies are vulnerable to adversarial perturbations, but vanilla training yields more robust policies.
problem Vulnerability of deep reinforcement learning policies to adversarial perturbations.
method Analysis of deep reinforcement learning policy landscape and comparison of vanilla vs. adversarial training.
result Vanilla training yields more robust policies compared to adversarial training.
Bayesian approach for policy search in stochastic domains.
problem Policy search in stochastic domains.
method Nested probabilistic programs, Lightweight Metropolis-Hastings (LMH) adaptation.
result Similar quality policies learned with simpler algorithm.
New method decomposes local projections to reveal historical drivers of estimates.
problem Uncertainty in interpreting local projections due to black-box nature.
method Decomposes LP estimates into contributions of historical events, interpreting weights as shocks and proximity scores.
result Dominant historical events drive impulse response estimates, revealing underlying mechanisms.
New method estimates off-policy data without needing known behavior policy.
problem Long-horizon reinforcement learning with limited on-policy data.
method Formulates problem as fixed point of operator, uses RKHS for estimation.
result Asymptotic consistency and finite-sample generalization proven.
Simplifies complex RL policies by ranking important decisions.
problem Complexity in RL policies makes them hard to analyze and interpret.
method Statistical fault localisation to rank states and prune unimportant decisions.
result Pruned policies can perform similarly to original policies, improving interpretability.
This paper explains CART random forests using stochastic control theory.
problem Understanding the inner workings of CART random forests.
method Developed a stochastic-control perspective on CART random forests, interpreting feature subsampling as a random feasible action set and the split rule as a policy.
result Established that the CART policy is locally stabilizing but globally suboptimal for the forest objective.
Machine learning classifiers are known to be vulnerable to inputs maliciously constructed by adversaries to force misclassification. Such adversarial examples have been extensively studied in the context of computer vision applications. In this work, we show adversarial attacks are also effective when targeting neural …
This paper investigates a class of attacks targeting the confidentiality aspect of security in Deep Reinforcement Learning (DRL) policies. Recent research have established the vulnerability of supervised machine learning models (e.g., classifiers) to model extraction attacks. Such attacks leverage the loosely-restricte…
OPAL optimizes labeling strategy for precise inference from uncertain models.
problem Inference from uncertain machine learning models is brittle.
method OPAL learns a smooth policy to adaptively label data points based on model uncertainty.
result OPAL yields estimators with the lowest variance and achieves nominal coverage in finite samples.
ESOP uses Bayesian optimization to find optimal lock-down schedules.
problem Finding optimal lock-down schedules balancing health and economy.
method Bayesian optimization interacting with epidemiological models.
result ESOP schedules balance public health and economic impacts.
This work explores adaptations of successful multi-armed bandits policies to the online contextual bandits scenario with binary rewards using binary classification algorithms such as logistic regression as black-box oracles. Some of these adaptations are achieved through bootstrapping or approximate bootstrapping, whil…
Overlap between treatment groups is required for non-parametric estimation of causal effects. If a subgroup of subjects always receives the same intervention, we cannot estimate the effect of intervention changes on that subgroup without further assumptions. When overlap does not hold globally, characterizing local reg…
We present an off-policy actor-critic algorithm for Reinforcement Learning (RL) that combines ideas from gradient-free optimization via stochastic search with learned action-value function. The result is a simple procedure consisting of three steps: i) policy evaluation by estimating a parametric action-value function;…
Proximal Policy Optimization (PPO) is a highly popular model-free reinforcement learning (RL) approach. However, we observe that in a continuous action space, PPO can prematurely shrink the exploration variance, which leads to slow progress and may make the algorithm prone to getting stuck in local optima. Drawing insp…
New method estimates and optimizes policy differences using orthogonal learning.
problem Offline reinforcement learning with safety concerns and cost limitations.
method Dynamic R-learner for estimating and optimizing Qπ(s,1)−Qπ(s,0), leveraging orthogonal estimation. result Consistent policy optimization with improved convergence rates.
GACEM optimizes complex multi-modal problems using neural networks.
problem Black-box optimization and constraint satisfaction in multi-modal environments.
method Modified Cross-Entropy Method with masked auto-regressive neural network.
result GACEM outperforms traditional CEM in diverse solutions, mode discovery, and sample efficiency.
Bayesian optimization improves policy search in reinforcement learning.
problem Finding optimal policies with high variance estimates from random samples.
method Develops an algorithm combining Bayesian optimization and policy gradients.
result Improves sample complexity and reduces variance in empirical evaluations.
A new reinforcement learning method uses model derivatives to improve policy optimization.
problem Improving sample efficiency and performance in model-based reinforcement learning.
method Constructs an actor-critic algorithm that uses the pathwise derivative of the learned model and policy.
result Consistently more sample efficient and matches model-free algorithms' asymptotic performance.
CPR models complex decision processes by breaking them into context-specific policies, improving interpretability and accuracy.
problem Interpreting dynamic human decision-making processes in medical contexts.
method Develops Contextualized Policy Recovery (CPR) framework for multi-task learning, modeling each context-specific policy as a linear map.
result Achieves state-of-the-art performance in predicting medical decisions, closing the gap between interpretable and black-box methods.
Understanding black-box machine learning models is crucial for their widespread adoption. Learning globally interpretable models is one approach, but achieving high performance with them is challenging. An alternative approach is to explain individual predictions using locally interpretable models. For locally interpre…
RPI combines imitation and reinforcement learning to improve policies efficiently.
problem High sample complexity in reinforcement learning.
method Active interleaving between imitation and reinforcement learning, using oracle queries for exploration.
result RPI outperforms existing methods across various domains.
Control policies, trained using the Deep Reinforcement Learning, have been recently shown to be vulnerable to adversarial attacks introducing even very small perturbations to the policy input. The attacks proposed so far have been designed using heuristics, and build on existing adversarial example crafting techniques …
Approach to optimize bidding policies offline using reinforcement learning.
problem Optimizing spending in online advertising under budget constraints.
method Offline reinforcement learning for optimizing differentiable base policies.
result Statistically significant performance gains in production bidding environments.
Modern vision-based reinforcement learning techniques often use convolutional neural networks (CNN) as universal function approximators to choose which action to take for a given visual input. Until recently, CNNs have been treated like black-box functions, but this mindset is especially dangerous when used for control…
Efficient neural architecture search by sampling structure and operations.
problem Efficiently searching for optimal neural architectures.
method Decouples structure and operation search, using reinforcement learning with policy vectors.
result Significantly improved efficiency compared to traditional methods.
Simplifies complex pricing models for better interpretability and revenue.
problem Complex pricing models are hard to interpret and not widely adopted.
method Model distillation to create interpretable pricing policies.
result Maximizes revenue while maintaining interpretability.
POETS optimizes LLMs by combining policy ensembles and KL regularization.
problem Balancing exploration and exploitation in decision-making and optimization.
method POETS uses policy ensembles and KL regularization to optimize LLMs efficiently.
result POETS achieves state-of-the-art sample efficiency across various domains.
Managers of US National Forests must decide what policy to apply for dealing with lightning-caused wildfires. Conflicts among stakeholders (e.g., timber companies, home owners, and wildlife biologists) have often led to spirited political debates and even violent eco-terrorism. One way to transform these conflicts into…
EI-MTD defends edge intelligence against adversarial attacks with dynamic scheduling.
problem Adversarial attacks on edge intelligence models.
method EI-MTD uses differential knowledge distillation to create robust member models and a dynamic scheduling policy based on a Bayesian Stackelberg game.
result EI-MTD effectively protects edge intelligence from black-box adversarial attacks.