This paper calibrates deep reinforcement learning models to improve planning and performance.
problem Inaccurate predictive uncertainties from deep learning systems hinder model-based reinforcement learning.
method Describes a simple method to calibrate uncertainties in model-based reinforcement learning agents.
result Calibrated model-based reinforcement learning agents achieve state-of-the-art performance with fewer samples.
Simple model-based reinforcement learning outperforms model-free methods in complex tasks.
problem Lagging performance of model-based reinforcement learning agents in non-trivial environments.
method Combining soft value estimates with stochastic value gradients.
result Simple model-based agents achieve state-of-the-art results in a high-dimensional humanoid control task.
A new multi-step model improves model-based reinforcement learning efficiency.
problem Expensive environmental interaction in reinforcement learning.
method Proposes a multi-step model for predicting action sequences with variable length.
result Multi-step model outperforms one-step model in preliminary tests.
Paper tackles action delays in reinforcement learning, proposing a delay-aware framework.
problem Action delays degrade reinforcement learning performance in real-world systems.
method Formal definition of delay-aware MDP, transformation into standard MDP with augmented states, delay-aware model-based reinforcement learning framework.
result Proposed framework is more efficient in training and transferable between systems with various delay durations.
A framework schedules hyperparameters for model-based reinforcement learning, improving performance.
problem Inadequate scheduling of hyperparameters in model-based reinforcement learning.
method Theoretical analysis and AutoMBPO framework to automatically schedule real data ratio and other hyperparameters.
result Training with hyperparameters scheduled by AutoMBPO significantly improves performance.
Survey of integrating planning and learning in model-based reinforcement learning.
problem Sequential decision making in AI, formalized as MDP optimization.
method Systematic coverage of dynamics model learning and planning-learning integration.
result Broad conceptual overview of model-based reinforcement learning.
Wasserstein and value-aware loss are shown to be equivalent in model-based RL.
problem Challenges in learning useful models in approximate settings.
method Equivalence between Wasserstein metric and VAML objective.
result Minimizing VAML objective is equivalent to minimizing Wasserstein metric.
Clarifies model-based RL's theoretical issues and counterexamples for popular losses.
problem Model-based reinforcement learning's empirical performance vs. theoretical properties and popular loss functions.
method Analyzes empirical and theoretical aspects of model-based RL and constructs counterexamples for losses.
result MuZero loss fails in stochastic and deterministic environments, leading to exponential sample complexity.
FOCUS improves offline RL by incorporating causal structure into world-models.
problem Learning effective policies from historical data without interaction.
method FOCUS proposes a practical algorithm that learns and leverages causal structure in offline RL.
result FOCUS outperforms plain model-based offline RL algorithms and other causal model-based RL algorithms.
A new method reduces compounding errors in model-based reinforcement learning.
problem Compounding errors in long horizon predictions from model-based reinforcement learning.
method Maximum Entropy Model Rollouts (MEMR) with non-uniform sampling and prioritized experience replay.
result Significantly reduces computation requirements compared to other model-based methods.
I2As learn to use model predictions to create flexible plans in reinforcement learning.
problem Improving data efficiency and robustness in reinforcement learning models.
method Imagination-Augmented Agents (I2As) combine model-free and model-based reinforcement learning, learning to interpret model predictions to construct flexible plans.
result I2As outperform baselines in data efficiency, performance, and robustness to model misspecification.
Residual algorithms improve reinforcement learning performance.
problem Distribution mismatch in model-based planning.
method Bidirectional target network technique for residual algorithms.
result Residual reinforcement learning significantly outperforms vanilla methods.
New hybrid RL algorithm outperforms model-free and model-based methods.
problem Improving reinforcement learning algorithms for MDPs.
method Combines model-free and model-based learning, with a PAC analysis.
result Outperforms both model-free and model-based methods in most cases.
The study compares reinforcement learning models and finds model-based approaches superior for complex MDPs.
problem Complexity of optimal Q-functions and policies in MDPs exceeds dynamics, hindering model-free methods.
method Theoretical analysis and empirical testing of neural network expressivity for policies, Q-functions, and dynamics.
result Model-based planning yields better policies for complex MDPs, improving performance on MuJoCo tasks.
POMBU improves model-based RL's asymptotic performance by estimating and using uncertainty.
problem Model-based reinforcement learning struggles with model errors, leading to suboptimal performance.
method POMBU uses estimated uncertainty to optimize policies conservatively, improving asymptotic performance.
result POMBU outperforms existing methods in sample efficiency and asymptotic performance.
A multi-step model reduces compounding errors in reinforcement learning.
problem Compounding errors in one-step models lead to inaccurate predictions in reinforcement learning.
method Introduced a multi-step model that directly outputs the outcome of a sequence of actions.
result The multi-step model yields better action selection and more accurate value-function estimation.
We present a new deep meta reinforcement learner, which we call Deep Episodic Value Iteration (DEVI). DEVI uses a deep neural network to learn a similarity metric for a non-parametric model-based reinforcement learning algorithm. Our model is trained end-to-end via back-propagation. Despite being trained using the mode…
We study Lipschitz models in reinforcement learning to bound prediction and value-function errors.
problem Bounding errors in reinforcement learning models with Lipschitz continuity constraints.
method We provide bounds on multi-step prediction error and value-function estimate using the Wasserstein metric for Lipschitz models.
result Lipschitz models lead to bounded errors in prediction and value-function estimates.
POME combines model-free and model-based methods to improve reinforcement learning performance.
problem Combining model-free and model-based reinforcement learning methods to improve performance.
method POME uses a model-free component for action value estimation and a model-based component for state value prediction, adding the error of these estimations as exploration value.
result POME outperforms PPO on 33 out of 49 Atari 2600 games.
Asynchronous framework speeds up model-based RL to real-time.
problem Real-time learning on real robots with model-based RL.
method Asynchronous framework for model-based reinforcement learning.
result Reduced run time to data collection time, improved sample complexity.
Unified reinforcement learning methods using hybrid inference.
problem Combining model-based and model-free reinforcement learning approaches.
method Control as Hybrid Inference (CHI) framework.
result CHI algorithm balances model-based and model-free learning.
Introduces LoCA regret to evaluate model-based RL methods.
problem Lack of consistent metrics to evaluate model-based RL methods.
method Inspired by neuroscience, introduces LoCA regret to measure model-based behavior.
result LoCA regret can identify model-based behavior and assess how close methods are to optimal model-based behavior.
Proposes H-UCRL for efficient model-based RL with sublinear regret.
problem Greedy policy exploration in model-based RL ignores epistemic uncertainty.
method Reparameterizes plausible models, hallucinates control, augments input space, solves with greedy planners.
result H-UCRL achieves provably sublinear regret for Gaussian Process models.
A new algorithm improves model-based reinforcement learning by guiding latent representations.
problem Improving model-based reinforcement learning through additional supervision.
method Proposed a novel asymmetric representation learning objective using latent guidance.
result Significantly improved performance over previous asymmetric approaches.
Study batch reinforcement learning methods for personalized medical treatments.
problem Batch reinforcement learning for personalized medical treatments.
method Direct policy learning and model-based learning approaches.
result Model-based learning is impossible with finite model classes but feasible with relaxed conditions.
Baconian simplifies MBRL experiments by providing a flexible framework.
problem Lack of reusable open-source frameworks for MBRL research.
method Developed a flexible and modularized framework, Baconian.
result Facilitates MBRL experiments by allowing customization and reuse of algorithms.
Hybrid controller combines model-based and policy-based reinforcement learning.
problem Combining model-based and policy-based reinforcement learning for stability and robustness.
method Designs a hybrid controller that interpolates a model-based linear controller and a differentiable policy.
result Proven to maintain stability and universal approximation properties.
New RL framework models continuous-time dynamics using neural ODEs.
problem Modeling continuous-time dynamics in semi-Markov decision processes.
method Model-based reinforcement learning with neural ODEs.
result High-performing policies developed with minimal data.
Model-based deep reinforcement learning improves Minecraft task performance.
problem Optimizing performance in Minecraft block-placing tasks.
method Combining DNN-based transition model with Monte Carlo tree search.
result Model-based approach achieves comparable performance to model-free methods but learns faster.
Survive method improves model-based RL by avoiding terminal states, reducing sample complexity.
problem High sample complexity in model-free RL methods limits real-world applications.
method Introduces 'survival' concept to model-based RL, focusing on avoiding terminal states instead of maximizing rewards.
result Survive method reduces training effort by focusing on terminal states, improving model-based RL performance.
Greedy policies in model-based RL achieve tight regret bounds without full planning.
problem Achieving efficient RL algorithms in MDP settings.
method Using greedy policies for 1-step planning in model-based RL.
result Greedy policies achieve i l d e O ( H S A T ) ilde{\mathcal{O}}(\sqrt{HSAT}) i l d e O ( H S A T ) regret bounds. The paper improves model-based reinforcement learning by using multi-timestep objectives.
problem Compounding errors in one-step dynamics models as trajectory length increases.
method Developed a multi-timestep objective as a weighted sum of losses at various future horizons.
result Exponentially decaying weights significantly improve long-horizon performance.
Policy Prediction Network improves continuous control problems with model-free and model-based learning.
problem Improving sample complexity and performance in continuous control problems.
method Integrates model-free and model-based reinforcement learning, introduces implicit model-based learning for continuous action space.
result First to introduce implicit model-based learning to Policy Gradient algorithms for continuous action space.
STEVE improves sample efficiency in reinforcement learning.
problem Combining model-free and model-based reinforcement learning with low sample complexity.
method Dynamic interpolation between model rollouts of various horizon lengths.
result STEVE achieves an order-of-magnitude increase in sample efficiency.
A method for learning a context latent vector to improve generalization in model-based RL.
problem Learning a global dynamics model that can generalize across different dynamics.
method Decomposes learning a global dynamics model into two stages: learning a context latent vector and predicting next states.
result Achieves superior generalization across various simulated robotics and control tasks.
Model-based RL with adversarial training for efficient recommendation policies.
problem Expensive model learning in model-free RL for recommender systems.
method Generative adversarial network for offline policy learning, discriminator for reward scaling and model bias reduction.
result Effective policy learning from offline and generated data, reducing bias in the learned model.
Introduces GAMPS for better model-based policy learning.
problem Misspecified model classes lead to poor policy estimates.
method Exploits current policy to learn approximate transition model, focusing on relevant parts of the environment.
result Empirically validated GAMPS on benchmark domains, demonstrating improved properties.
SOLAR learns efficient representations for RL in complex image domains.
problem Efficient model-based reinforcement learning in domains with complex observations like images.
method Optimizes structured representations for inferring simple dynamics and cost models from data.
result Substantially better final performance than other model-based RL methods, more efficient than model-free RL.
Curious Meta-Controller alternates between model-based and model-free control to improve sample efficiency.
problem Combining the benefits of model-based and model-free control to enhance sample efficiency.
method Adaptive alternation between model-based and model-free control using curiosity feedback.
result Significantly improved sample efficiency and near-optimal performance on robotic tasks.
A contraction analysis improves model-based RL's error recovery.
problem Theoretical understanding of model-based reinforcement learning.
method Contraction analysis applied to both stochastic and deterministic state transitions.
result Error reduction in cumulative reward using branched rollouts.
Curious Replay improves model-based reinforcement learning agents' adaptability.
problem Existing model-based reinforcement learning agents struggle to adapt quickly to changing environments.
method Curious Replay uses a curiosity-based priority signal for prioritized experience replay tailored to model-based agents.
result Agents using Curious Replay achieve improved performance in exploration and on benchmarks.
Paper analyzes model usage in policy optimization, improving sample efficiency and performance.
problem Balancing ease of data generation with model bias in reinforcement learning.
method Formulated and analyzed model-based reinforcement learning algorithm with empirical model generalization.
result Simple model-based approach outperforms existing methods in sample efficiency and asymptotic performance.
Improved model-based reinforcement learning for multi-agent Markov games.
problem Suboptimal sample complexity for model-based algorithms in multi-agent reinforcement learning.
method Optimistic Nash Value Iteration (Nash-VI) for two-player zero-sum Markov games.
result First model-based algorithm matching information-theoretic lower bound with improved sample complexity.
This study tackles adversarial corruption in model-based reinforcement learning.
problem Adversarial corruption in model-based reinforcement learning.
method Maximum likelihood estimation (MLE) approach for learning transition model in both online and offline settings.
result Proves a regret of i l d e O ( T + C ) ilde{\mathcal{O}}(\sqrt{T} + C) i l d e O ( T + C ) for CR-OMLE and a suboptimality of O ( C / n ) \mathcal{O}(C/n) O ( C / n ) for CR-PMLE. A new method, Count-MORL, improves offline reinforcement learning by using state-action frequency.
problem Improving offline reinforcement learning performance.
method Integrates count-based conservatism into model-based offline reinforcement learning.
result The learned policy is near-optimal and outperforms existing methods.
CVRL tackles complex visual observations in reinforcement learning.
problem Complex visual observations in natural environments.
method Contrastive Variational Reinforcement Learning (CVRL) learns a contrastive variational model by maximizing mutual information between latent states and observations.
result CVRL achieves comparable performance with state-of-the-art model-based DRL methods and significantly outperforms them on tasks with complex observations.
Deep reinforcement learning controls drones without model knowledge.
problem Real-time robot control without engineered models.
method Learnt probabilistic model of drone dynamics, model-based reinforcement learning.
result Controller and value function optimized through generated latent trajectories.
The paper introduces a new reinforcement learning model that improves sample efficiency and performance.
problem Improving sample efficiency and performance in reinforcement learning.
method The paper proposes a transcoder network that includes model learning in the DQN loss function, providing a richer training signal.
result The proposed method leads to superior results on Atari games compared to vanilla DQN.