Progressive reinforcement learning combines expert policies for multi-skilled motion control.
problem Integrating policies for multiple skills in continuous control problems.
method Policy distillation, input injection, transfer learning.
result Incremental policy augmentation with new skills.
Investigates conditions for Poincaré map existence and uniqueness in systems with impulse effects.
problem Existence and uniqueness of Poincaré maps for systems with impulse effects.
method Investigates sufficient conditions for the existence and uniqueness of Poincaré maps for dynamical systems with impulse effects evolving on a differentiable manifold.
result Shows sufficient conditions for the existence and uniqueness of Poincaré maps for systems with impulse effects.
EMG analysis quantifies bipedal standing quality in SCI patients.
problem Quantifying the quality of bipedal standing in spinal cord injury patients.
method Multi-channel surface EMG recordings during spinal stimulation therapy sessions.
result Multi-channel EMG recording can provide accurate, fast, and robust estimation for standing quality in SCI patients.
FORK improves model-free reinforcement learning performance.
problem Improving model-free reinforcement learning performance.
method Introducing a new forward-looking Actor (FORK) for Actor-Critic algorithms.
result FORK significantly improves performance in various environments.
Paper proposes a method to learn locomotion tasks for microrobots.
problem Challenges in designing gaits for compliant or micro-scale robots.
method Formalizes locomotion as a contextual policy search task, learns multi-objective locomotion primitives.
result Successfully learned locomotion primitives for microrobots without prior knowledge of their dynamics.
Method learns to map dynamics of different systems.
problem Mapping dynamics of different systems.
method Learned latent dynamical system for mapping.
result Learned correspondences enable imagined motions and bisimulation.
A method for predicting human locomotion using sensor correlations.
problem Predicting missing signals in human locomotion data.
method Coregionalised Locomotion Envelopes - multi-dimensional manifold regression.
result Developed a qualitative method for robust control of rehabilitation robots.
A novel approach learns goal-conditioned policies for locomotion using batch RL.
problem Training goal-conditioned policies for rotation invariant locomotion.
method Data augmentation and Siamese framework for invariance.
result Our approach outperforms existing RL algorithms on 3D locomotion agents.
A novel RL approach learns robotic manipulation without human demonstrations.
problem Learning robotic manipulation policies efficiently and effectively.
method Introducing simulated locomotion demonstration rewards (SLDRs) to enable RL learning.
result The approach achieves higher success rates and faster learning compared to alternatives.
The paper develops a learning framework for diverse legged robots.
problem General and autonomous learning of core skills in locomotion.
method Data-efficient, off-policy multi-task RL algorithm with semantically identical reward functions.
result The same algorithm can learn diverse and reusable locomotion skills across different legged robots.
New framework for task-independent legged locomotion.
problem Building stable legged locomotion systems in robotics.
method Task-independent spiking central pattern generator using learning methods.
result Robotic legged locomotion at different speeds and within the same gait cycle.
A decentralized deep RL controller improves hexapod locomotion learning.
problem Deep RL struggles with real-world legged robot control.
method Decentralized deep RL on a hexapod robot.
result Decentralized approach learns better and faster.
New friction model for geometric locomotion systems.
problem Modeling asymmetric friction in locomotion systems.
method Introducing asymmetric friction into geometric locomotion models using Finsler metrics.
result Generalized motility map for systems with asymmetric friction.
Robotics: Rolling robots on a moving platform can be controlled.
problem Controlling the motion of rolling robots atop a moving platform.
method Developed a mathematical model and demonstrated simulations.
result Platform acceleration can control robot's heading and motion.
RTN generates realistic character transitions from context.
problem Manual transition animation creation is tedious and time-consuming.
method Recurrent Neural Network (RNN) with LSTM architecture trained on past context.
result RTN produces realistic transitions that rival motion capture.
This paper examines linear embeddings for high-dimensional Bayesian optimization, identifying and addressing issues to improve performance.
problem Scaling Bayesian optimization to high-dimensional spaces while maintaining sample efficiency.
method Study and empirical evaluation of linear embeddings for BO, addressing design choices and their impact on performance.
result Properly addressing issues in linear embeddings significantly improves their efficacy in BO.
SAVO actor improves reinforcement learning by avoiding local optima in complex Q-functions.
problem Gradient ascent in complex Q-functions leads to suboptimal solutions.
method SAVO actor generates multiple action proposals and truncates poor local optima.
result SAVO actor finds optimal actions more frequently and outperforms other architectures.
PerCDL learns personalized dictionaries for physiological signals combining global and local structures.
problem Representing datasets with both global and local structures in human physiological signals.
method Personalized Convolutional Dictionary Learning (PerCDL) that combines a global and personalized local dictionary.
result PerCDL effectively learns interpretable representations for human locomotion data.
CARL controls a quadruped to move naturally in complex environments.
problem Motion synthesis in dynamic environments with complex constraints.
method CARL uses GANs to adapt high-level controls to action distributions and deep reinforcement learning for dynamic recovery.
result CARL can be controlled with high-level directives and react naturally to dynamic environments.
Paper proposes PRR network for better experience reuse in reinforcement learning.
problem Efficient experience reuse in reinforcement learning across multiple granularities.
method Proposes PRR network trained on multi-level architecture to extract and store experience.
result PRR network leads to better experience reuse and improved performance.
New method learns diverse solutions in reinforcement learning without gradient bias.
problem Lack of diverse solutions in reinforcement learning tasks.
method Maximizes state-action-based mutual information directly, using variational lower bound.
result Successfully learns an infinite set of diverse solutions.
Survey examines challenges and solutions in sim-to-real transfer for robotics.
problem Challenges in transferring robotic systems from simulation to real-world environments.
method Leveraging techniques like domain randomization, real-to-sim transfer, state and action abstractions, and sim-real co-training.
result Promising results in closing the reality gap across various robotic domains.
Unified policy controls diverse agents through modular neural networks.
problem Learning control policies for various agent morphologies.
method Shared Modular Policies (SMP) with decentralized control and message passing.
result A single modular policy controls multiple agent morphologies.
SMiRL learns to minimize surprise in unstable environments, improving agent performance.
problem Learning useful behaviors in unpredictable, unstable environments.
method Alternates between learning a density model and improving policy to seek more predictable stimuli.
result SMiRL agents can play games, control robots, and navigate mazes without task-specific rewards.
Reinforcement learning for embodied agents is a challenging problem. The accumulated reward to be optimized is often a very rugged function, and gradient methods are impaired by many local optimizers. We demonstrate, in an experimental setting, that incorporating an intrinsic reward can smoothen the optimization landsc…
Deep RL learns robot walking gaits in real-world environments.
problem Difficulty in applying deep RL to real-world robotic tasks due to poor sample complexity and hyperparameter sensitivity.
method Sample-efficient deep RL algorithm based on maximum entropy RL, requiring minimal per-task tuning and modest trials.
result Acquired stable walking gaits on a real-world Minitaur robot in about two hours.
PDERL improves evolutionary reinforcement learning by using learning-based variation operators.
problem Scalability issue in Genetic Algorithms when combined with Deep Neural Networks.
method Integrates evolutionary and reinforcement learning through a hierarchical approach with learning-based variation operators.
result PDERL outperforms traditional evolutionary and reinforcement learning methods in robot locomotion tasks.
Trajectory segmentation is the process of subdividing a trajectory into parts either by grouping points similar with respect to some measure of interest, or by minimizing a global objective function. Here we present a novel online algorithm for segmentation and summary, based on point density along the trajectory, and …
When the Poincaré map associated with a periodic orbit of a hybrid dynamical system has constant-rank iterates, we demonstrate the existence of a constant-dimensional invariant subsystem near the orbit which attracts all nearby trajectories in finite time. This result shows that the long-term behavior of a hybrid model…
Method identifies multi-scale behavioral motifs in rat locomotion.
problem Challenging to identify behavioral motifs across different time scales.
method Preprocessing and parallel HMM training for different scales.
result Rat locomotion is composed of distinct motifs from second to minute scale.
CP-DRL improves curriculum reinforcement learning by leveraging causal relationships.
problem Designing effective task sequences for reinforcement learning.
method Causal-Paced Deep Reinforcement Learning (CP-DRL) that approximates SCM differences based on interaction data.
result CP-DRL outperforms existing methods on benchmarks, achieving faster convergence and higher returns.
CoNES optimizes blackbox functions using convex optimization and information geometry.
problem Optimizing high-dimensional blackbox functions efficiently.
method Formulated as a convex program that adapts evolutionary strategies gradient estimates.
result Vastly outperforms conventional blackbox optimization methods on benchmarks and MuJoCo tasks.
Study shows not all ML models are uniquely identifiable from data.
problem Identifiability issues in machine learning models.
method Investigated through a case study on gait dynamics using a bipedal-spring mass model.
result Some parameters can be identified, but others remain unidentifiable.
Paper tackles skill transfer in RL for morphologically different agents.
problem Transfer skills between morphologically different reinforcement learning agents.
method Proposes a paired variational encoder-decoder model (PVED) for subspace learning.
result Demonstrates improved skill transfer efficiency compared to state-of-the-art methods.
We adapt the ideas underlying the success of Deep Q-Learning to the continuous action domain. We present an actor-critic, model-free algorithm based on the deterministic policy gradient that can operate over continuous action spaces. Using the same learning algorithm, network architecture and hyper-parameters, our algo…
Bayesian approach uses Gaussian process for reinforcement learning.
problem Robotic locomotion environments
method Bayesian actor-critic, model-free reinforcement learning with Gaussian process for exploration and policy optimization.
result Gaussian process method outperforms current algorithms in robotic locomotion environments.
Contextual policy search allows adapting robotic movement primitives to different situations. For instance, a locomotion primitive might be adapted to different terrain inclinations or desired walking speeds. Such an adaptation is often achievable by modifying a small number of hyperparameters. However, learning, when …
For sophisticated reinforcement learning (RL) systems to interact usefully with real-world environments, we need to communicate complex goals to these systems. In this work, we explore goals defined in terms of (non-expert) human preferences between pairs of trajectory segments. We show that this approach can effective…
D4PG combines distributional reinforcement learning with distributed learning for control tasks.
problem Continuous control tasks in reinforcement learning.
method Adapting distributional reinforcement learning to continuous control, using a distributed framework, N-step returns, and prioritized experience replay.
result D4PG achieves state-of-the-art performance across various control tasks.
New PAC-Bayesian approach stabilizes actor-critic learning.
problem Training instability in actor-critic algorithms.
method Employing PAC-Bayesian bound as the critic training objective.
result Significant improvement in online learning performance.
Symbolic regression constructs simple equations for complex systems.
problem Creating accurate yet simple models for dynamic systems.
method Employing symbolic regression with two genetic programming algorithms.
result Analytic models outperform neural networks and local regression.
Develops ODRPO to improve RL algorithms with better performance and stability.
problem RL algorithms converge to sub-optimal solutions due to limited policy representation.
method Integrates DRO approach to solve trust region constrained optimization problem without parameterizing policies.
result Achieves globally optimal policy update and higher sample efficiency.
EMI uses predictive signals to guide exploration in sparse reward settings.
problem Challenges of reinforcement learning with sparse reward signals.
method Constructs embedding representations of states and actions for forward prediction in the representation space.
result Competitive results on challenging tasks with continuous control and discrete actions.
Relational Mimic improves visual imitation learning from video demonstrations.
problem Improving robustness and sample efficiency in visual imitation learning.
method Combines generative adversarial networks and relational learning.
result Improves agent performance in challenging locomotion tasks.
KL-regularized RL from expert demos can lead to slow, unstable learning.
problem Pathological training dynamics in KL-regularized RL from expert demonstrations.
method Empirical analysis and non-parametric behavioral reference policies.
result KL-regularized RL can be significantly improved by using non-parametric behavioral policies.
GRAM enhances deep RL for reliable real-world deployment.
problem Generalizing deep RL across in-distribution and out-of-distribution scenarios.
method Introduces a robust adaptation module and a joint training pipeline.
result GRAM achieves strong generalization performance in simulations and hardware.
RANDPOL uses randomized networks for efficient reinforcement learning in continuous state and action MDPs.
problem Efficient reinforcement learning in environments with continuous state and action spaces.
method RANDPOL uses randomized function approximation to represent policy and value functions, providing finite time guarantees and improved numerical performance.
result RANDPOL achieves better numerical performance and provides finite time guarantees compared to deep neural network based algorithms.
The paper proposes a new method to estimate optimal policies using MCMC.
problem Estimating the optimal policy for systems with unknown dynamics and reward functions.
method Using Markov Chain Monte Carlo to generate samples from the posterior distribution of parameters conditioned on optimality.
result The method provably converges to the globally optimal stochastic policy with similar variance to policy gradient methods.