Using results on the topology of moduli space of polygons [Jaggi, 92; Kapovich and Millson, 94], it can be shown that for a planar robot arm with n n n segments there are some values of the base-length, z z z , at which the configuration space of the constrained arm (arm with its end effector fixed) has two disconnected com…
A robot assists a human in a bandit task to learn and improve performance.
problem Learning preferences in humans when they are also learning.
method Introduces assistive multi-armed bandit, where a robot helps a human maximize cumulative reward.
result Human performance can be better when effectively communicating observed rewards to the robot, not just by learning optimally.
New metric solves correspondence problem for robotic arm imitation learning.
problem Establishing corresponding states and actions between different robotic arms.
method Introducing a distance measure between dissimilar robotic arms and using it as a loss function.
result The distance measure effectively learns imitation policies by minimizing distance between robotic arms.
The configuration space of the mechanism of a planar robot is studied. We consider a robot which has n n n arms such that each arm is of length 1+1 and has a rotational joint in the middle, and that the endpoint of the k k k -th arm is fixed to R e 2 ( k − 1 ) π n i Re^{\frac{2(k-1)π}ni} R e n 2 ( k − 1 ) π i . Generically, the configuration space is diffeomorphic t…
New method selects best exploration strategies in uncertain environments.
problem Selecting optimal strategies in unknown, multi-strategy environments.
method Formulates Multi-Armed Bandits problem with diversity of effects as reward signal.
result Method outperforms fixed mixtures of strategies in diverse, challenging conditions.
The paper presents a method to reduce arm motion complexity for prosthetics and robotics.
problem Reducing the complexity of human arm motions for robotic and prosthetic control.
method Data-driven techniques including DTW, DBA, Ward's distance, batch-DTW, and fPCA.
result Representative motion clusters and averages for different arm DOF levels.
It is known that a closed polygon P is a critical point of the oriented area function if and only if P is a cyclic polygon, that is, P P P can be inscribed in a circle. Moreover, there is a short formula for the Morse index. Going further in this direction, we extend these results to the case of open polygonal chains, or…
Deep RL controls robotic arms efficiently.
problem Continuous control of robotic arms.
method Combination of two reinforcement learning methods and preprocessing techniques.
result The new combination learns more effectively than a baseline.
Robotic arm learns to manipulate a ball by choosing goals from learned experience.
problem Efficient discovery of skills for long-living autonomous agents without supervision.
method Intrinsically motivated goal exploration using learned goal spaces from deep representation learning.
result Recent results show applicability of learned goal spaces on real-world robotic tasks.
CRB tackles rising rewards in combinatorial online learning.
problem Rising rewards in combinatorial online learning.
method CRB framework and CRUCB algorithm.
result Empirical and theoretical validation of CRUCB's effectiveness.
For a m-tuple a=(a_1,...,a_m) of positive real numbers, the robot arm of type a in R^d is the map f^a:(S^{d-1})^m -> R^d defined by f^a(z_1,...,z_m) to be the sum of the a_jz_j's. Our aim is to attack the inverse problem via the horizontal liftings for the distribution Delta^a orthogonal to the fibers of f^a. One shows…
Reinforcement learning is a promising approach to developing hard-to-engineer adaptive solutions for complex and diverse robotic tasks. However, learning with real-world robots is often unreliable and difficult, which resulted in their low adoption in reinforcement learning research. This difficulty is worsened by the …
Robot learns user preferences from brain signals.
problem Decoding user preferences for robot motions from brain signals.
method Proposes a novel approach using electroencephalography to decode user preferences from brain signals.
result Brain signals can reliably infer user preferences for robot trajectories.
A graph bandit algorithm learns optimal paths on unknown graphs.
problem Optimal path selection on unknown graphs under uncertainty.
method G-UCB algorithm based on offline graph planning and optimism principle.
result Achieves tight regret bound of O ( ∣ S ∣ T log ( T ) + D ∣ S ∣ log T ) O(\sqrt{|S|T\log(T)}+D|S|\log T) O ( ∣ S ∣ T log ( T ) + D ∣ S ∣ log T ) . This work tackles force control for contact-rich manipulation tasks with rigid robots using RL.
problem Challenges in working with real robotic hardware, especially position-controlled robots.
method Combines RL with traditional force control techniques, implementing parallel position/force control and admittance control.
result Validated methods on both simulation and real robot (UR3 e-series) for force control.
Novel method decomposes configuration space for improved collision checking.
problem Improving collision checking in high-degree-of-freedom robot motion planning.
method Proposes a configuration space decomposition method to build a composite classifier.
result Composite classifier outperforms state-of-the-art single classifier methods.
This work evaluates task-agnostic exploration methods for fixed-batch learning.
problem Expensive real-world experience for robotics tasks.
method Fixed datasets for arbitrary task learning.
result Improved offline learning for robotics tasks.
Proposes DEXP3.M for unknown delay in multi-arm bandit with multiple play.
problem Unknown delays in adversarial multi-armed bandit with multiple play.
method DEXP3.M algorithm addressing the challenge of associating feedback losses to arms.
result Regret bound is only slightly worse than single play setting.
Robots learn intentions from multiple cues to reduce uncertainty.
problem Uncertainty in human-robot interaction for vulnerable users.
method Multimodal classifier fusion using Bayesian Independent Opinion Pool.
result Fused classifiers outperform individual modalities in accuracy and uncertainty reduction.
Paper presents derivative-free methods for online inverse dynamics modeling.
problem Online learning of inverse dynamics models without numerical differentiation.
method Derivative-free framework for rigid body dynamics, data-driven, and semiparametric models.
result Proposed `derivative-free' methods outperform existing methodologies in real data experiments.
TossingBot learns to throw objects accurately with residual physics.
problem Learning to throw arbitrary objects accurately and quickly.
method End-to-end formulation that learns control parameters from visual observations.
result TossingBot achieves 600+ grasps per hour with 85% throwing accuracy.
New algorithm reduces sample complexity for imitation learning.
problem High sample complexity limits deployment of adversarial imitation algorithms.
method Trajectory-centric reinforcement learning ideas.
result Improved learning rate and efficiency in imitation tasks.
TIDBD adapts step sizes online for better robotic predictions.
problem Choosing appropriate learning parameters for online prediction-learning.
method Temporal-Difference Incremental Delta-Bar-Delta (TIDBD) for step-size adaptation.
result TIDBD performs comparably to classic TD learning and detects sensor failures.
Robot-assisted dressing offers an opportunity to benefit the lives of many people with disabilities, such as some older adults. However, robots currently lack common sense about the physical implications of their actions on people. The physical implications of dressing are complicated by non-rigid garments, which can r…
Robot learns multiple tasks hierarchically by transferring knowledge.
problem Learning multiple complex tasks in open-ended environments.
method Task-oriented procedures, goal-babbling, imitation learning, active learning, intrinsic motivation.
result Robots can learn complex tasks more efficiently by transferring knowledge from simpler ones.
New model accounts for continuous human trajectories in robotics.
problem Inaccurate probabilistic models of human behavior in robotics.
method Developed a new probabilistic model that considers distances between continuous trajectories.
result The new model outperforms existing models in explaining human behavior and improving robot inference.
RIDM combines imitation and RL with a single demo, no action info needed.
problem Learning from a single observed demonstration without action information.
method Reinforced Inverse Dynamics Modeling (RIDM) that operates on raw state features.
result RIDM performs favorably compared to baseline on simulated and real tasks.
Paper introduces timing-based adversarial attacks on DRL-based navigation systems.
problem Vulnerability of DRL-based navigation systems to adversarial attacks.
method Timing-based adversarial strategies using physical noise patterns.
result Adversarial timing attacks significantly degrade DRL-based navigation performance.
Method trains vision and control policies on real robots quickly.
problem Training vision-based control policies on real robots efficiently.
method Multi-task Reinforcement Learning with auxiliary tasks.
result Significant learning speed-ups and task learning from-scratch.
New method learns time-invariant rewards from demonstrations.
problem Learning robust rewards for tasks with varying execution times.
method Model-based inverse reinforcement learning with time-invariant costs.
result Approach enables learning from misaligned demonstrations and generalizes spatially.
This work proposes a RL approach to learn versatile robotic manipulation tasks.
problem Challenging manipulation tasks in robotics and vision.
method Reinforcement learning (RL) to combine primitive skills, no intermediate rewards, few demonstrations, and efficient skill learning.
result Versatile robotic manipulation in challenging settings with temporary occlusions and dynamic scene changes.
Robots learn spatial perception from sensorimotor invariants.
problem Developing autonomous robots that perceive space without human intuition.
method Study how a robot's motor commands relate to changes in exteroceptive inputs to deduce its spatial configuration.
result Robots can learn the configuration space of their sensors, revealing a planar position and orientation.
This work analyzes how multi-agent reinforcement learning can bridge the gap to reality in distributed multi-robot systems.
problem Collaborative learning in distributed multi-robot systems with varying sensors and actuators.
method Simulation-based analysis using PPO and Bullet physics engine, considering different types of perturbations.
result PPO's robustness is affected by the presence of different types of perturbations and the number of agents experiencing them.
New RL approach speeds up training across tasks.
problem Training reinforcement learning models quickly and efficiently.
method Combines planning quasi-metric and task-specific aimers.
result Achieves multiple-fold speed-up on bit-flip and robotic arm tasks.
Robots learn new skills from demonstrations, using active learning to detect missing information.
problem Detecting missing information during skill generalization and transitioning to new tasks.
method Novel active learning algorithm based on deep generative models and metric learning in latent spaces.
result Smooth trajectories generated by asking for additional demonstrations when non-smooth transitions are detected.
Time-agnostic predictors predict frames without fixed time intervals.
problem Predicting events in the future or between waypoints is difficult.
method Decouple visual prediction from a rigid notion of time, discovering predictable 'bottleneck' frames.
result Predictions are of higher visual quality and correspond to coherent semantic subgoals.
Robots rely on sensors to provide them with information about their surroundings. However, high-quality sensors can be extremely expensive and cost-prohibitive. Thus many robotic systems must make due with lower-quality sensors. Here we demonstrate via a case study how modeling a sensor can improve its efficacy when em…
SOLAR learns efficient representations for RL in complex image domains.
problem Efficient model-based reinforcement learning in domains with complex observations like images.
method Optimizes structured representations for inferring simple dynamics and cost models from data.
result Substantially better final performance than other model-based RL methods, more efficient than model-free RL.
This paper presents a novel approach for incremental semiparametric inverse dynamics learning. In particular, we consider the mixture of two approaches: Parametric modeling based on rigid body dynamics equations and nonparametric modeling based on incremental kernel methods, with no prior information on the mechanical …
Fairness in AI decisions for users with varying performance.
problem Ensuring fairness in AI decisions for users with different performance levels.
method Contextual Multi-Armed Bandit algorithm with fairness constraints.
result Accounting for user contexts improves fairness in AI decisions.
New method uses simple sensor intentions to learn complex tasks.
problem Defining reward schemes for exploration in robotic systems.
method Introduce simple sensor intentions (SSIs) to define auxiliary tasks.
result Learning system can solve complex robotic tasks using only raw sensor streams.
Paper presents new dataset for disentanglement learning from physical objects.
problem Transfer of disentanglement models from synthetic to real-world data.
method Developed a dataset of physical objects with controlled variations, used a robotic arm for precise manipulation.
result Disentanglement models perform poorly on real data but selection of models and hyperparameters improves transfer.
Reinforcement learning optimizes robot trajectories for unknown dynamics.
problem Optimizing robot trajectories for systems with unknown dynamics.
method Curriculum learning with reinforcement learning to generate smooth trajectories.
result Reinforcement learning agent outperforms PID controllers in trajectory tracking.
Autonomous robots need to interact with unknown, unstructured and changing environments, constantly facing novel challenges. Therefore, continuous online adaptation for lifelong-learning and the need of sample-efficient mechanisms to adapt to changes in the environment, the constraints, the tasks, or the robot itself a…
One of the most interesting features of Bayesian optimization for direct policy search is that it can leverage priors (e.g., from simulation or from previous tasks) to accelerate learning on a robot. In this paper, we are interested in situations for which several priors exist but we do not know in advance which one fi…
The most data-efficient algorithms for reinforcement learning in robotics are model-based policy search algorithms, which alternate between learning a dynamical model of the robot and optimizing a policy to maximize the expected return given the model and its uncertainties. However, the current algorithms lack an effec…
New RL framework learns task completion without prior knowledge.
problem Learning task completion without linguistic or perceptual knowledge.
method Sequentially imagining visual goals and choosing actions.
result Framework outperforms flat and hierarchical architectures.
AWAC combines offline and online data to accelerate RL learning.
problem Challenges in applying RL to real-world robotic control due to exploration and sample complexity.
method Combines sample-efficient dynamic programming with maximum likelihood policy updates.
result AWAC enables rapid learning of robotic skills with prior data and online experience.