A new framework estimates expert policy support to create a reward function for imitation learning.
problem Imitation learning from expert trajectories without reinforcement signals.
method Estimating the support of the expert policy to compute a fixed reward function.
result Comparable or better performance than state-of-the-art methods on discrete and continuous domains.
Paper introduces a novel reward function for noisy financial markets using imitation learning.
problem Noisy reward function in financial markets hinders RL agent performance.
method Integrates imitation learning feedback with reinforcement learning to improve reward function design.
result Improves financial performance metrics compared to traditional benchmarks and RL agents.
Improves imitation learning in RL by learning reward function efficiently.
problem Lack of effective reward function approximation in AIRL for imitation tasks.
method Proposes Off-Policy AIRL that combines adversarial learning with efficient reward function approximation.
result Shows superior imitation performance and efficiency compared to state-of-the-art AIL algorithms.
New reward function improves GAIL performance in task-based environments.
problem Reward bias in adversarial imitation learning.
method Proposed a new reward function to overcome existing biases.
result New reward function outperforms existing methods in task-based environments.
A new method uses optimal transport for imitation learning, improving efficiency and performance.
problem Recovering expert policies from demonstrations with efficient and general reward functions.
method Proposes Wasserstein Adversarial Imitation Learning, using Kantorovich potentials as reward functions and regularized optimal transport.
result Significantly improves sample-efficiency and average cumulative rewards in robotic experiments.
D-REX learns reward functions from ranked demonstrations to beat the demonstrator's performance.
problem Limited performance of imitation learning compared to the demonstrator.
method Disturbance-based Reward Extrapolation (D-REX) that generates ranked demonstrations from behavioral cloning.
result D-REX significantly outperforms standard imitation learning approaches and beats the demonstrator's performance.
Researchers improved Minecraft game performance using imitation learning.
problem Achieving state-of-the-art performance in immersive environments like Minecraft.
method Applied imitation learning to Minecraft, optimizing network architecture, loss function, and data augmentation.
result Reported stronger results than previous experiments, reaching second place in a competition.
Paper proposes Imitative Models combining IL and planning for flexible goal achievement.
problem Difficult to direct IL to arbitrary goals and specify reward functions for goal-directed planning.
method Imitative Models are probabilistic predictive models that plan interpretable expert-like trajectories.
result Imitative Models outperform IL and planning approaches in a dynamic autonomous driving task.
New metric solves correspondence problem for robotic arm imitation learning.
problem Establishing corresponding states and actions between different robotic arms.
method Introducing a distance measure between dissimilar robotic arms and using it as a loss function.
result The distance measure effectively learns imitation policies by minimizing distance between robotic arms.
GASIL encourages agents to imitate past good trajectories in reinforcement learning.
problem Long-term credit assignment in sparse and delayed reward environments.
method Generative Adversarial Imitation Learning (GASIL) framework.
result GASIL improves performance in reinforcement learning tasks with delayed rewards.
Paper tackles imitation learning with sparse rewards and heterogeneous actions.
problem Challenges of imitation learning with sparse rewards and different actions.
method Proposes a method that balances imitation and reinforcement learning objectives.
result Agent efficiently leverages sparse rewards and learns from different actions.
Study optimal investment under imitation of decision-changing rates.
problem Optimal investment under imitation of decision-changing rates.
method Proposed integral disparity to quantify imitation, derived general solution using variational method, analyzed asymptotic properties, validated with real data.
result Investor's optimal decisions under imitation of decision-changing rates.
Bayesian Robust Optimization for Imitation Learning (BROIL) balances risk and reward.
problem Learning robust policies for new states in imitation learning.
method Bayesian reward function inference and user-specific risk tolerance.
result BROIL outperforms risk-sensitive and risk-neutral algorithms.
This work improves imitation learning and goal-conditioned RL by estimating value densities.
problem Effective solutions for imitation and goal-conditioned reinforcement learning require reliably reaching specified states or demonstrations.
method The approach uses recent advances in density estimation to learn value functions efficiently and without hindsight bias.
result The method achieves state-of-the-art demonstration sample-efficiency in imitation learning and is both efficient and bias-free in goal-conditioned reinforcement learning.
WDAIL uses Wasserstein distance for more effective reward shaping in IL.
problem Fixed reward functions in GAIL limit performance on complex tasks.
method Introduces Wasserstein distance and PPO for improved reward shaping and stability.
result Significant performance improvement in complex MuJoCo tasks.
New method minimizes f-divergence for better imitation learning.
problem Learning from multi-modal demonstrations, especially interpolating between modes.
method Minimizing reverse KL divergence or I-projection for any f-divergence.
result Our method reliably imitates multi-modal behaviors better than existing methods.
BAIL learns high-performing policies from batch data without additional interactions.
problem Learning high-performing policies from batch data without additional interactions.
method BAIL learns a V function, selects high-performing actions, and trains a policy network using those actions.
result BAIL outperforms other batch Q-learning and imitation-learning schemes in MuJoCo benchmark.
New approach uses inverse reinforcement learning to improve language model training.
problem Training large language models using imitation learning methods.
method Developed a new method of inverse reinforcement learning to optimize sequences directly.
result IRL-based fine-tuning leads to better performance and diversity in language generation.
Study shows a linear quadratic regulator's imitation learning converges globally.
problem Global convergence of imitation learning for linear quadratic regulators.
method Analyzed alternating gradient algorithm and established Q-linear rate of convergence.
result Established a unique saddle point for globally optimal policy and reward function.
SQIL uses a simple reward strategy to encourage long-horizon imitation of expert demonstrations.
problem Challenges in imitation learning with high-dimensional, continuous observations and unknown dynamics.
method Imitates expert demonstrations by providing a constant reward of +1 for matching actions in demonstrated states, and 0 for all others.
result Empirically outperforms behavioral cloning and achieves competitive results compared to GAIL.
Paper improves AIRL by enhancing policy imitation and addressing reward recovery issues.
problem Inadequate policy imitation and limited transferable reward recovery in AIRL.
method Substituted built-in algorithm with SAC for policy updating and proposed PPO-AIRL + SAC hybrid framework.
result SAC improves policy imitation but hinders reward recovery; PPO-AIRL + SAC achieves satisfactory transfer effect.
New state-only IL algorithm tackles MDP transition mismatch.
problem Transition dynamics mismatch between expert and imitator MDPs.
method Adversarial state-only IL with two subproblems solved iteratively.
result Effective performance improvement in transition dynamics mismatch scenarios.
QFIL improves offline RL by filtering data to reduce bias and variance.
problem Improving offline reinforcement learning policies with limited data.
method QFIL uses a filtered dataset to improve policies, trading off bias and variance through quantile selection.
result QFIL provides a safe policy improvement step with function approximation and effectively balances bias and variance.
Compressed imitation learning uses simplicity priors for efficient expert behavior copying.
problem Efficiently learn expert behaviors with minimal data.
method Utilizes policy simplicity as a prior for sample-efficient imitation learning.
result Significantly higher scores achieved with limited expert demonstrations.
Dual-structured method improves cross-domain imitation learning.
problem Learning policies in a target domain from source domain expert demonstrations.
method Extracts domain and policy expertness features, synthesizes new features.
result Superior performance compared to other algorithms in MuJoCo tasks.
PWIL learns agent behavior from expert using Wasserstein distance.
problem Learning agent behavior from expert demonstrations efficiently.
method Primal Wasserstein Imitation Learning (PWIL) with offline reward function.
result PWIL achieves expert behavior on MuJoCo tasks with minimal interactions.
Continuous control imitation learning fails if expert actions are smooth.
problem Continuous control imitation learning fails if expert actions are smooth.
method Study of imitation learning in discrete-time, continuous state-and-action control systems.
result Any smooth, deterministic imitator policy suffers exponentially larger error than the expert.
New method aligns states for better imitation learning.
problem Different dynamics models between imitator and expert.
method State alignment-based reinforcement learning.
result Superior performance on various imitation learning settings.
Algorithm corrects mistakes in imitation learning tasks.
problem Covariate shift problem in learning from demonstrations.
method Value Iteration with Negative Sampling (VINS) for conservative extrapolation.
result VINS corrects mistakes of the behavioral cloning policy.
A method learns to imitate from a single video demonstration using contrastive training and Siamese networks.
problem Learning to imitate from video demonstrations without direct access to state or action information.
method Contrastive training with Siamese recurrent neural networks to learn rewards and an RL policy to minimize distance.
result The method significantly improves policy learning and outperforms current techniques in various environments.
We consider a heterogeneous agent-based economic model where economic agents have strictly bounded rationality and where income allocation strategies evolve through selective imitation. Income is calculated by a Cobb-Douglas type production function, and selection of strategies for imitation depends on the income growt…
Inspiration Learning expands imitation learning to different action spaces.
problem Current imitation learning techniques are limited to agents with the same action space.
method Designs Advantage Actor-Critic algorithms using Preferential based Reinforcement Learning.
result Method extends imitation learning to new action spaces.
DSAC improves robustness of imitation learning.
problem Learning from few expert data and distribution shift.
method DSAC uses adversarial reward function.
result DSAC performs well on PyBullet environments.
New algorithm reduces sample inefficiency and reward bias in AI learning.
problem Implicit reward bias and high sample inefficiency in AI learning.
method Discriminator-Actor-Critic using off-policy Reinforcement Learning.
result Average 10x reduction in policy-environment interaction samples.
Develops a new model for spatio-temporal data using imitation learning.
problem Capturing complex spatio-temporal dependence in discrete event data.
method NEST point process model with imitation learning for efficient model fitting.
result The imitation learning approach leads to more robust and interpretable results.
New method learns agent actions from raw video demonstrations.
problem Imitation learning for natural-looking movements.
method Generative Adversarial Networks (GANs) for joint state and reward learning.
result Adversarial imitation learning on raw videos produces similar performance to state-of-the-art methods.
ICIL learns policies invariant to multiple environments, improving generalization.
problem Learning policies from multiple environments leads to spurious correlations.
method ICIL learns invariant feature representations and a matching imitation policy.
result ICIL policies generalize better to unseen environments.
This work explains why online imitation learning improves faster than theory predicts.
problem Online imitation learning's empirical policy improvement speed exceeds theoretical predictions.
method The authors analyze online imitation learning with a convex, smooth, and non-negative loss function, proving policy improvement in expectation and high probability.
result Adopting a sufficiently expressive policy class in online IL increases both policy improvement speed and performance bias.
This paper improves understanding of GAIL's generalization and computational efficiency.
problem Understanding the theoretical properties of GAIL, especially its generalization and computational aspects.
method Investigates GAIL's theoretical properties, showing guarantees for generalization and computational efficiency.
result GAIL can be efficiently solved by stochastic first order optimization algorithms with sublinear convergence.
ADVISOR dynamically balances imitation and reinforcement learning to overcome the imitation gap.
problem The gap between imitation learning and reinforcement learning when teaching agents have privileged information.
method Adaptive Insubordination (ADVISOR) dynamically weights imitation and reward-based reinforcement learning losses.
result On-the-fly switching with ADVISOR outperforms pure imitation, pure reinforcement learning, and their combinations.
Survey of methods for learning from observation without requiring expert actions.
problem Lack of access to expert actions in imitation learning.
method Survey and classification of state-only imitation learning (SOIL) methods.
result Identification of open problems and future research directions.
HGAIL learns policies without expert demonstrations.
problem Lack of expert demonstrations in imitation learning.
method Combines hindsight and GAIL to learn policies.
result Comparable performance to current methods, with curriculum learning.
Bayesian REX learns Atari games from demonstrations efficiently.
problem Bayesian reward learning for complex control problems is computationally intractable.
method Bayesian Reward Extrapolation (Bayesian REX) pre-trains a low-dimensional feature encoding and uses preferences to perform fast Bayesian inference.
result Bayesian REX learns Atari games from demonstrations in 5 minutes, competitive with state-of-the-art methods.
Recommender systems aim to find an accurate and efficient mapping from historic data of user-preferred items to a new item that is to be liked by a user. Towards this goal, energy-based sequence generative adversarial nets (EB-SeqGANs) are adopted for recommendation by learning a generative model for the time series of…
B-REX efficiently learns Atari game policies from pixel inputs using Bayesian methods.
problem Learning reward functions from visual inputs with uncertainty and safety considerations.
method Bayesian Reward Extrapolation (B-REX) using successor features and preferences.
result B-REX generates posterior samples efficiently, enabling high-confidence performance bounds.
MAPS and MAPS-SE learn from multiple suboptimal experts to improve policies efficiently.
problem Learning from multiple suboptimal experts in reinforcement learning.
method Active policy improvement from multiple black-box oracles.
result MAPS and MAPS-SE achieve sample efficiency and state-wise imitation learning.
New algorithm reduces sample complexity for imitation learning.
problem High sample complexity limits deployment of adversarial imitation algorithms.
method Trajectory-centric reinforcement learning ideas.
result Improved learning rate and efficiency in imitation tasks.
Reinforcement learning mimics expert behavior.
problem Learning from expert demonstrations in reinforcement learning.
method Reduction to reinforcement learning with a stationary reward.
result Expert reward can be recovered and imitation learning is bounded.