Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

2.8%5.5%8.3%11.0% · Jun 201919922001200920182026
48 results for model-based reinforcement

This paper calibrates deep reinforcement learning models to improve planning and performance.

problem Inaccurate predictive uncertainties from deep learning systems hinder model-based reinforcement learning.
method Describes a simple method to calibrate uncertainties in model-based reinforcement learning agents.
result Calibrated model-based reinforcement learning agents achieve state-of-the-art performance with fewer samples.

Simple model-based reinforcement learning outperforms model-free methods in complex tasks.

problem Lagging performance of model-based reinforcement learning agents in non-trivial environments.
method Combining soft value estimates with stochastic value gradients.
result Simple model-based agents achieve state-of-the-art results in a high-dimensional humanoid control task.

Paper tackles action delays in reinforcement learning, proposing a delay-aware framework.

problem Action delays degrade reinforcement learning performance in real-world systems.
method Formal definition of delay-aware MDP, transformation into standard MDP with augmented states, delay-aware model-based reinforcement learning framework.
result Proposed framework is more efficient in training and transferable between systems with various delay durations.

A framework schedules hyperparameters for model-based reinforcement learning, improving performance.

problem Inadequate scheduling of hyperparameters in model-based reinforcement learning.
method Theoretical analysis and AutoMBPO framework to automatically schedule real data ratio and other hyperparameters.
result Training with hyperparameters scheduled by AutoMBPO significantly improves performance.

Wasserstein and value-aware loss are shown to be equivalent in model-based RL.

problem Challenges in learning useful models in approximate settings.
method Equivalence between Wasserstein metric and VAML objective.
result Minimizing VAML objective is equivalent to minimizing Wasserstein metric.

Clarifies model-based RL's theoretical issues and counterexamples for popular losses.

problem Model-based reinforcement learning's empirical performance vs. theoretical properties and popular loss functions.
method Analyzes empirical and theoretical aspects of model-based RL and constructs counterexamples for losses.
result MuZero loss fails in stochastic and deterministic environments, leading to exponential sample complexity.

FOCUS improves offline RL by incorporating causal structure into world-models.

problem Learning effective policies from historical data without interaction.
method FOCUS proposes a practical algorithm that learns and leverages causal structure in offline RL.
result FOCUS outperforms plain model-based offline RL algorithms and other causal model-based RL algorithms.

A new method reduces compounding errors in model-based reinforcement learning.

problem Compounding errors in long horizon predictions from model-based reinforcement learning.
method Maximum Entropy Model Rollouts (MEMR) with non-uniform sampling and prioritized experience replay.
result Significantly reduces computation requirements compared to other model-based methods.

I2As learn to use model predictions to create flexible plans in reinforcement learning.

problem Improving data efficiency and robustness in reinforcement learning models.
method Imagination-Augmented Agents (I2As) combine model-free and model-based reinforcement learning, learning to interpret model predictions to construct flexible plans.
result I2As outperform baselines in data efficiency, performance, and robustness to model misspecification.

The study compares reinforcement learning models and finds model-based approaches superior for complex MDPs.

problem Complexity of optimal Q-functions and policies in MDPs exceeds dynamics, hindering model-free methods.
method Theoretical analysis and empirical testing of neural network expressivity for policies, Q-functions, and dynamics.
result Model-based planning yields better policies for complex MDPs, improving performance on MuJoCo tasks.

POMBU improves model-based RL's asymptotic performance by estimating and using uncertainty.

problem Model-based reinforcement learning struggles with model errors, leading to suboptimal performance.
method POMBU uses estimated uncertainty to optimize policies conservatively, improving asymptotic performance.
result POMBU outperforms existing methods in sample efficiency and asymptotic performance.

A multi-step model reduces compounding errors in reinforcement learning.

problem Compounding errors in one-step models lead to inaccurate predictions in reinforcement learning.
method Introduced a multi-step model that directly outputs the outcome of a sequence of actions.
result The multi-step model yields better action selection and more accurate value-function estimation.

We study Lipschitz models in reinforcement learning to bound prediction and value-function errors.

problem Bounding errors in reinforcement learning models with Lipschitz continuity constraints.
method We provide bounds on multi-step prediction error and value-function estimate using the Wasserstein metric for Lipschitz models.
result Lipschitz models lead to bounded errors in prediction and value-function estimates.

POME combines model-free and model-based methods to improve reinforcement learning performance.

problem Combining model-free and model-based reinforcement learning methods to improve performance.
method POME uses a model-free component for action value estimation and a model-based component for state value prediction, adding the error of these estimations as exploration value.
result POME outperforms PPO on 33 out of 49 Atari 2600 games.

Introduces LoCA regret to evaluate model-based RL methods.

problem Lack of consistent metrics to evaluate model-based RL methods.
method Inspired by neuroscience, introduces LoCA regret to measure model-based behavior.
result LoCA regret can identify model-based behavior and assess how close methods are to optimal model-based behavior.

Proposes H-UCRL for efficient model-based RL with sublinear regret.

problem Greedy policy exploration in model-based RL ignores epistemic uncertainty.
method Reparameterizes plausible models, hallucinates control, augments input space, solves with greedy planners.
result H-UCRL achieves provably sublinear regret for Gaussian Process models.

A new algorithm improves model-based reinforcement learning by guiding latent representations.

problem Improving model-based reinforcement learning through additional supervision.
method Proposed a novel asymmetric representation learning objective using latent guidance.
result Significantly improved performance over previous asymmetric approaches.

Study batch reinforcement learning methods for personalized medical treatments.

problem Batch reinforcement learning for personalized medical treatments.
method Direct policy learning and model-based learning approaches.
result Model-based learning is impossible with finite model classes but feasible with relaxed conditions.

Hybrid controller combines model-based and policy-based reinforcement learning.

problem Combining model-based and policy-based reinforcement learning for stability and robustness.
method Designs a hybrid controller that interpolates a model-based linear controller and a differentiable policy.
result Proven to maintain stability and universal approximation properties.

Model-based deep reinforcement learning improves Minecraft task performance.

problem Optimizing performance in Minecraft block-placing tasks.
method Combining DNN-based transition model with Monte Carlo tree search.
result Model-based approach achieves comparable performance to model-free methods but learns faster.

Survive method improves model-based RL by avoiding terminal states, reducing sample complexity.

problem High sample complexity in model-free RL methods limits real-world applications.
method Introduces 'survival' concept to model-based RL, focusing on avoiding terminal states instead of maximizing rewards.
result Survive method reduces training effort by focusing on terminal states, improving model-based RL performance.

The paper improves model-based reinforcement learning by using multi-timestep objectives.

problem Compounding errors in one-step dynamics models as trajectory length increases.
method Developed a multi-timestep objective as a weighted sum of losses at various future horizons.
result Exponentially decaying weights significantly improve long-horizon performance.

Policy Prediction Network improves continuous control problems with model-free and model-based learning.

problem Improving sample complexity and performance in continuous control problems.
method Integrates model-free and model-based reinforcement learning, introduces implicit model-based learning for continuous action space.
result First to introduce implicit model-based learning to Policy Gradient algorithms for continuous action space.

A method for learning a context latent vector to improve generalization in model-based RL.

problem Learning a global dynamics model that can generalize across different dynamics.
method Decomposes learning a global dynamics model into two stages: learning a context latent vector and predicting next states.
result Achieves superior generalization across various simulated robotics and control tasks.

Model-based RL with adversarial training for efficient recommendation policies.

problem Expensive model learning in model-free RL for recommender systems.
method Generative adversarial network for offline policy learning, discriminator for reward scaling and model bias reduction.
result Effective policy learning from offline and generated data, reducing bias in the learned model.

Introduces GAMPS for better model-based policy learning.

problem Misspecified model classes lead to poor policy estimates.
method Exploits current policy to learn approximate transition model, focusing on relevant parts of the environment.
result Empirically validated GAMPS on benchmark domains, demonstrating improved properties.

SOLAR learns efficient representations for RL in complex image domains.

problem Efficient model-based reinforcement learning in domains with complex observations like images.
method Optimizes structured representations for inferring simple dynamics and cost models from data.
result Substantially better final performance than other model-based RL methods, more efficient than model-free RL.

Curious Meta-Controller alternates between model-based and model-free control to improve sample efficiency.

problem Combining the benefits of model-based and model-free control to enhance sample efficiency.
method Adaptive alternation between model-based and model-free control using curiosity feedback.
result Significantly improved sample efficiency and near-optimal performance on robotic tasks.

Curious Replay improves model-based reinforcement learning agents' adaptability.

problem Existing model-based reinforcement learning agents struggle to adapt quickly to changing environments.
method Curious Replay uses a curiosity-based priority signal for prioritized experience replay tailored to model-based agents.
result Agents using Curious Replay achieve improved performance in exploration and on benchmarks.

Paper analyzes model usage in policy optimization, improving sample efficiency and performance.

problem Balancing ease of data generation with model bias in reinforcement learning.
method Formulated and analyzed model-based reinforcement learning algorithm with empirical model generalization.
result Simple model-based approach outperforms existing methods in sample efficiency and asymptotic performance.

Improved model-based reinforcement learning for multi-agent Markov games.

problem Suboptimal sample complexity for model-based algorithms in multi-agent reinforcement learning.
method Optimistic Nash Value Iteration (Nash-VI) for two-player zero-sum Markov games.
result First model-based algorithm matching information-theoretic lower bound with improved sample complexity.

This study tackles adversarial corruption in model-based reinforcement learning.

problem Adversarial corruption in model-based reinforcement learning.
method Maximum likelihood estimation (MLE) approach for learning transition model in both online and offline settings.
result Proves a regret of ildeO(T+C) ilde{\mathcal{O}}(\sqrt{T} + C) for CR-OMLE and a suboptimality of O(C/n)\mathcal{O}(C/n) for CR-PMLE.

A new method, Count-MORL, improves offline reinforcement learning by using state-action frequency.

problem Improving offline reinforcement learning performance.
method Integrates count-based conservatism into model-based offline reinforcement learning.
result The learned policy is near-optimal and outperforms existing methods.

CVRL tackles complex visual observations in reinforcement learning.

problem Complex visual observations in natural environments.
method Contrastive Variational Reinforcement Learning (CVRL) learns a contrastive variational model by maximizing mutual information between latent states and observations.
result CVRL achieves comparable performance with state-of-the-art model-based DRL methods and significantly outperforms them on tasks with complex observations.

The paper introduces a new reinforcement learning model that improves sample efficiency and performance.

problem Improving sample efficiency and performance in reinforcement learning.
method The paper proposes a transcoder network that includes model learning in the DQN loss function, providing a richer training signal.
result The proposed method leads to superior results on Atari games compared to vanilla DQN.