State-only imitation learning improves dexterous manipulation learning from videos.
problem High sample complexity in complex domains like dexterous manipulation.
method Train an inverse dynamics model to predict actions from states and train the policy jointly.
result Performs on par with state-action approaches and outperforms RL alone.
New state-only IL algorithm tackles MDP transition mismatch.
problem Transition dynamics mismatch between expert and imitator MDPs.
method Adversarial state-only IL with two subproblems solved iteratively.
result Effective performance improvement in transition dynamics mismatch scenarios.
Imitation learning is an effective alternative approach to learn a policy when the reward function is sparse. In this paper, we consider a challenging setting where an agent and an expert use different actions from each other. We assume that the agent has access to a sparse reward function and state-only expert observa…
Generative model for hypergraphs captures complex interactions without pairwise reductions.
problem Challenges in generating realistic hypergraphs with pairwise reductions.
method Structured stochastic diffusion on relaxed incidence matrices.
result Generative model preserves structure-aware noising and yields explicit Gaussian law.
Survey of methods for learning from observation without requiring expert actions.
problem Lack of access to expert actions in imitation learning.
method Survey and classification of state-only imitation learning (SOIL) methods.
result Identification of open problems and future research directions.
Imitation from observation is the framework of learning tasks by observing demonstrated state-only trajectories. Recently, adversarial approaches have achieved significant performance improvements over other methods for imitating complex behaviors. However, these adversarial imitation algorithms often require many demo…
Imitation from observation (IfO) is the problem of learning directly from state-only demonstrations without having access to the demonstrator's actions. The lack of action information both distinguishes IfO from most of the literature in imitation learning, and also sets it apart as a method that may enable agents to l…
This paper studies Learning from Observations (LfO) for imitation learning with access to state-only demonstrations. In contrast to Learning from Demonstration (LfD) that involves both action and state supervision, LfO is more practical in leveraging previously inapplicable resources (e.g. videos), yet more challenging…
SPEDER extracts state-action abstraction from dynamics for reinforcement learning.
problem Curse of dimensionality and limited applicability of spectral methods.
method Spectral Decomposition Representation (SPEDER) that extracts state-action abstraction from dynamics without policy dependence.
result Theoretical analysis establishes sample efficiency in online and offline settings.
PQR estimates reward functions from actions and states without assuming state-only rewards.
problem Estimating reward functions from actions and states without state-only assumptions.
method Deep learning approach that sequentially estimates policy, Q-function, and reward.
result PQR uniquely recovers true reward with known transitions and bounds error with unknown transitions.
CFIL uses coupled flows to model state distributions for imitation learning.
problem Lack of explicit modeling of state distributions in reinforcement and imitation learning.
method Coupled normalizing flows for state and state-action distributions.
result CFIL achieves state-of-the-art performance on benchmark tasks.
Paper solves tracking control for (x,u)-flat systems using classical states.
problem Tracking control for (x,u)-flat systems. method Quasi-static feedback of classical states.
result Achieves linear, decoupled and asymptotically stable tracking error dynamics.
TGD improves conditional sampling by concentrating computation on promising trajectories.
problem Efficiently training-free conditional sampling with diffusion priors.
method Tempered Guided Diffusion (TGD) using annealed sequential Monte Carlo.
result TGD yields a consistent particle approximation to the posterior as the number of particles grows.
New RL method learns from state transitions without actions.
problem Offline RL with missing action labels.
method State policy discretisation and decQN algorithm.
result Improves convergence speed and performance in online RL.
Neural network models colloidal particle dynamics in non-equilibrium systems.
problem Analyzing non-equilibrium dynamics of many-body colloidal systems.
method Combining power functional theory and machine learning, training a neural network to predict internal force fields.
result The neural network accurately predicts dynamics in non-equilibrium systems, in good agreement with simulations.