Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

2457 · Sep 201919922001200920182026
48 results for bipedal locomotion

Investigates conditions for Poincaré map existence and uniqueness in systems with impulse effects.

problem Existence and uniqueness of Poincaré maps for systems with impulse effects.
method Investigates sufficient conditions for the existence and uniqueness of Poincaré maps for dynamical systems with impulse effects evolving on a differentiable manifold.
result Shows sufficient conditions for the existence and uniqueness of Poincaré maps for systems with impulse effects.

EMG analysis quantifies bipedal standing quality in SCI patients.

problem Quantifying the quality of bipedal standing in spinal cord injury patients.
method Multi-channel surface EMG recordings during spinal stimulation therapy sessions.
result Multi-channel EMG recording can provide accurate, fast, and robust estimation for standing quality in SCI patients.

Paper proposes a method to learn locomotion tasks for microrobots.

problem Challenges in designing gaits for compliant or micro-scale robots.
method Formalizes locomotion as a contextual policy search task, learns multi-objective locomotion primitives.
result Successfully learned locomotion primitives for microrobots without prior knowledge of their dynamics.

A novel RL approach learns robotic manipulation without human demonstrations.

problem Learning robotic manipulation policies efficiently and effectively.
method Introducing simulated locomotion demonstration rewards (SLDRs) to enable RL learning.
result The approach achieves higher success rates and faster learning compared to alternatives.

The paper develops a learning framework for diverse legged robots.

problem General and autonomous learning of core skills in locomotion.
method Data-efficient, off-policy multi-task RL algorithm with semantically identical reward functions.
result The same algorithm can learn diverse and reusable locomotion skills across different legged robots.

This paper examines linear embeddings for high-dimensional Bayesian optimization, identifying and addressing issues to improve performance.

problem Scaling Bayesian optimization to high-dimensional spaces while maintaining sample efficiency.
method Study and empirical evaluation of linear embeddings for BO, addressing design choices and their impact on performance.
result Properly addressing issues in linear embeddings significantly improves their efficacy in BO.

SAVO actor improves reinforcement learning by avoiding local optima in complex Q-functions.

problem Gradient ascent in complex Q-functions leads to suboptimal solutions.
method SAVO actor generates multiple action proposals and truncates poor local optima.
result SAVO actor finds optimal actions more frequently and outperforms other architectures.

PerCDL learns personalized dictionaries for physiological signals combining global and local structures.

problem Representing datasets with both global and local structures in human physiological signals.
method Personalized Convolutional Dictionary Learning (PerCDL) that combines a global and personalized local dictionary.
result PerCDL effectively learns interpretable representations for human locomotion data.

CARL controls a quadruped to move naturally in complex environments.

problem Motion synthesis in dynamic environments with complex constraints.
method CARL uses GANs to adapt high-level controls to action distributions and deep reinforcement learning for dynamic recovery.
result CARL can be controlled with high-level directives and react naturally to dynamic environments.

Paper proposes PRR network for better experience reuse in reinforcement learning.

problem Efficient experience reuse in reinforcement learning across multiple granularities.
method Proposes PRR network trained on multi-level architecture to extract and store experience.
result PRR network leads to better experience reuse and improved performance.

New method learns diverse solutions in reinforcement learning without gradient bias.

problem Lack of diverse solutions in reinforcement learning tasks.
method Maximizes state-action-based mutual information directly, using variational lower bound.
result Successfully learns an infinite set of diverse solutions.

Survey examines challenges and solutions in sim-to-real transfer for robotics.

problem Challenges in transferring robotic systems from simulation to real-world environments.
method Leveraging techniques like domain randomization, real-to-sim transfer, state and action abstractions, and sim-real co-training.
result Promising results in closing the reality gap across various robotic domains.

SMiRL learns to minimize surprise in unstable environments, improving agent performance.

problem Learning useful behaviors in unpredictable, unstable environments.
method Alternates between learning a density model and improving policy to seek more predictable stimuli.
result SMiRL agents can play games, control robots, and navigate mazes without task-specific rewards.

Deep RL learns robot walking gaits in real-world environments.

problem Difficulty in applying deep RL to real-world robotic tasks due to poor sample complexity and hyperparameter sensitivity.
method Sample-efficient deep RL algorithm based on maximum entropy RL, requiring minimal per-task tuning and modest trials.
result Acquired stable walking gaits on a real-world Minitaur robot in about two hours.

PDERL improves evolutionary reinforcement learning by using learning-based variation operators.

problem Scalability issue in Genetic Algorithms when combined with Deep Neural Networks.
method Integrates evolutionary and reinforcement learning through a hierarchical approach with learning-based variation operators.
result PDERL outperforms traditional evolutionary and reinforcement learning methods in robot locomotion tasks.

When the Poincaré map associated with a periodic orbit of a hybrid dynamical system has constant-rank iterates, we demonstrate the existence of a constant-dimensional invariant subsystem near the orbit which attracts all nearby trajectories in finite time. This result shows that the long-term behavior of a hybrid model…

2011-09-08abs ↗pdf ↗

CP-DRL improves curriculum reinforcement learning by leveraging causal relationships.

problem Designing effective task sequences for reinforcement learning.
method Causal-Paced Deep Reinforcement Learning (CP-DRL) that approximates SCM differences based on interaction data.
result CP-DRL outperforms existing methods on benchmarks, achieving faster convergence and higher returns.

CoNES optimizes blackbox functions using convex optimization and information geometry.

problem Optimizing high-dimensional blackbox functions efficiently.
method Formulated as a convex program that adapts evolutionary strategies gradient estimates.
result Vastly outperforms conventional blackbox optimization methods on benchmarks and MuJoCo tasks.

Paper tackles skill transfer in RL for morphologically different agents.

problem Transfer skills between morphologically different reinforcement learning agents.
method Proposes a paired variational encoder-decoder model (PVED) for subspace learning.
result Demonstrates improved skill transfer efficiency compared to state-of-the-art methods.

We adapt the ideas underlying the success of Deep Q-Learning to the continuous action domain. We present an actor-critic, model-free algorithm based on the deterministic policy gradient that can operate over continuous action spaces. Using the same learning algorithm, network architecture and hyper-parameters, our algo…

2015-09-09abs ↗pdf ↗

Contextual policy search allows adapting robotic movement primitives to different situations. For instance, a locomotion primitive might be adapted to different terrain inclinations or desired walking speeds. Such an adaptation is often achievable by modifying a small number of hyperparameters. However, learning, when …

2015-11-13abs ↗pdf ↗

For sophisticated reinforcement learning (RL) systems to interact usefully with real-world environments, we need to communicate complex goals to these systems. In this work, we explore goals defined in terms of (non-expert) human preferences between pairs of trajectory segments. We show that this approach can effective…

2017-06-12abs ↗pdf ↗

D4PG combines distributional reinforcement learning with distributed learning for control tasks.

problem Continuous control tasks in reinforcement learning.
method Adapting distributional reinforcement learning to continuous control, using a distributed framework, N-step returns, and prioritized experience replay.
result D4PG achieves state-of-the-art performance across various control tasks.

EMI uses predictive signals to guide exploration in sparse reward settings.

problem Challenges of reinforcement learning with sparse reward signals.
method Constructs embedding representations of states and actions for forward prediction in the representation space.
result Competitive results on challenging tasks with continuous control and discrete actions.

KL-regularized RL from expert demos can lead to slow, unstable learning.

problem Pathological training dynamics in KL-regularized RL from expert demonstrations.
method Empirical analysis and non-parametric behavioral reference policies.
result KL-regularized RL can be significantly improved by using non-parametric behavioral policies.

RANDPOL uses randomized networks for efficient reinforcement learning in continuous state and action MDPs.

problem Efficient reinforcement learning in environments with continuous state and action spaces.
method RANDPOL uses randomized function approximation to represent policy and value functions, providing finite time guarantees and improved numerical performance.
result RANDPOL achieves better numerical performance and provides finite time guarantees compared to deep neural network based algorithms.

The paper proposes a new method to estimate optimal policies using MCMC.

problem Estimating the optimal policy for systems with unknown dynamics and reward functions.
method Using Markov Chain Monte Carlo to generate samples from the posterior distribution of parameters conditioned on optimality.
result The method provably converges to the globally optimal stochastic policy with similar variance to policy gradient methods.