Generative models improve pose transfer between people.
problem Transferring actions from one person to another.
method Used nearest neighbor and generative models (pix2pix) for pose transfer.
result Generative models outperform k-NN in generating corresponding frames and generalizing outside the action set.
New framework for RL transfer learning with state-action mismatch.
problem High sample complexity in RL from scratch.
method Embeddings to transfer knowledge between MDPs with different state- and action-spaces.
result Successful transfer learning in scenarios with state- and action-space mismatches.
New landmark states improve transfer learning in multi-task RL.
problem Improving sample complexity and regret in new RL tasks.
method Topological landmark covering, landmark value functions, action pruning.
result Theoretical bounds on Q values at state-action pairs.
Robot learns multiple tasks hierarchically by transferring knowledge.
problem Learning multiple complex tasks in open-ended environments.
method Task-oriented procedures, goal-babbling, imitation learning, active learning, intrinsic motivation.
result Robots can learn complex tasks more efficiently by transferring knowledge from simpler ones.
Size-independent neural transfer for RDDL planning.
problem Sample inefficiency and time-consuming training for neural planners of RDDL MDPs.
method Two key innovations: state encoder and parameter-tied action decoder.
result Powerful transfer across problem sizes with superior learning curves.
MULTIPOLAR aggregates diverse source policies for efficient transfer RL.
problem Efficiently transfer knowledge between different environmental dynamics.
method Adaptive action aggregation and residual prediction network.
result Demonstrated effectiveness across diverse simulated environments.
MO2 learns useful behaviours from past experience for new tasks.
problem Discovering useful behaviours from past experience and transferring them to new tasks.
method Model-Based Offline Options (MO2) framework supporting sample-efficient bottleneck option discovery over continuous state-action spaces.
result MO2 outperforms recent option learning methods on complex long-horizon continuous control tasks.
A new 1-transfer space construction for hyperbolic groups.
problem Proving the Farrell-Jones conjecture for hyperbolic groups.
method Explicit construction of a 1-transfer space.
result A new method to prove the Farrell-Jones conjecture for hyperbolic groups.
Combines NES and PPO to enhance exploration in various environments.
problem Improving exploration in reinforcement learning environments.
method Parameter transfer and parameter space noise methods for combining NES and PPO.
result PPO benefits from both NES methods in discrete and continuous control tasks.
Survey examines challenges and solutions in sim-to-real transfer for robotics.
problem Challenges in transferring robotic systems from simulation to real-world environments.
method Leveraging techniques like domain randomization, real-to-sim transfer, state and action abstractions, and sim-real co-training.
result Promising results in closing the reality gap across various robotic domains.
ExTra uses past experience to guide exploration in reinforcement learning.
problem Navigating new tasks with limited prior knowledge.
method Transfer-guided exploration using bisimulation distances and Softmax sampling.
result ExTra outperforms traditional exploration methods in gridworld environments.
Designs a framework to transfer causal models between similar environments.
problem Transferability of causal models between different but similar environments.
method Object-oriented representations and continuous optimization for structure learning.
result Demonstrates advantages in gridworld settings using reinforcement learning.
A new method for task transfer in reinforcement learning using expert preferences.
problem Inconveniently obtaining expert demonstrations and cost functions for task transfer.
method Develops a novel framework that uses expert preferences to select relevant demonstrations and learns the target cost function and trajectory distribution.
result Demonstrated effectiveness through simulations on various benchmarks.
This paper tackles online strategic decision making with asymmetry and knowledge transportability.
problem Strategic decision making with information asymmetry and knowledge transportability challenges.
method Developed a sample-efficient algorithm for online learning under these conditions.
result Proved sample complexity of O(1/ε2) for learning an ε-optimal policy. Dreaming model enables robots to learn policies from images.
problem Learning to control robots based on images is challenging.
method Learning a future representation/scene regressor conditioned on robot actions.
result Dreaming model enables policy learning that transfers to real-world.
PTF accelerates RL by reusing source policies without measuring task similarity.
problem Leveraging prior knowledge for faster RL.
method Adaptive Policy Transfer Framework (PTF) for RL.
result Significantly accelerates RL learning process and surpasses state-of-the-art methods.
Offline RL tackles resource-constrained online deployment with improved policy transfer.
problem Training policies with limited online features using a rich offline dataset.
method Introduce a policy transfer algorithm that first trains a teacher agent with full offline features and then transfers knowledge to a student agent with limited online features.
result Consistent improvement in performance over baseline methods on resource-constrained datasets.
Growing action spaces accelerates learning in complex tasks.
problem Efficient learning in tasks with large combinatorial action spaces.
method Curriculum of progressively growing action spaces, off-policy reinforcement learning, data transfer.
result Efficacy demonstrated in control tasks and large-scale StarCraft tasks.
Adversarial adaptation improves reinforcement learning performance.
problem Slow training and exploration in reinforcement learning.
method Adversarial domain adaptation for feature initialization.
result Significant improvement in reinforcement learning performance.
Proposes a curriculum learning algorithm to maximize cumulative return in reinforcement learning.
problem Maximizing cumulative return in reinforcement learning tasks.
method Task sequencing algorithm maximizing cumulative return, using curriculum learning to minimize suboptimal actions.
result Significantly better performance on cumulative return maximization compared to metaheuristic algorithms.
New algorithm improves knowledge transfer in dynamic decision-making.
problem Utilizing data from existing ventures to improve decision-making in new ventures.
method Proposes Transferred Fitted Q-Iteration algorithm for estimating optimal action-state function Q∗. result Significantly improved final learning error of Q∗ function. New algorithm reduces sample complexity for new tasks by leveraging prior knowledge.
problem Designing reinforcement learning agents that reduce sample complexity for new tasks.
method Designing an algorithm that quickly identifies an accurate solution by seeking informative state-action pairs from related tasks, using a generative model.
result PAC bounds on sample complexity demonstrate the benefits of using prior knowledge.
We propose a framework that learns a representation transferable across different domains and tasks in a label efficient manner. Our approach battles domain shift with a domain adversarial loss, and generalizes the embedding to novel task using a metric learning-based approach. Our model is simultaneously optimized on …
Paper tackles action delays in reinforcement learning, proposing a delay-aware framework.
problem Action delays degrade reinforcement learning performance in real-world systems.
method Formal definition of delay-aware MDP, transformation into standard MDP with augmented states, delay-aware model-based reinforcement learning framework.
result Proposed framework is more efficient in training and transferable between systems with various delay durations.
Paper tackles robust knowledge transfer in parallel RL tasks.
problem Transfer knowledge from low-tier to high-tier tasks in parallel RL without shared dynamics or reward functions.
method Identifies Optimal Value Dominance condition and proposes online learning algorithms for both tasks.
result Achieves constant regret on partial states and near-optimal regret when tasks are dissimilar.
We propose a general notion of algebraic gauge theory obtained via extracting the main properties of classical gauge theory. Building on a recent work on transferring curved A∞-structures we show that, under certain technical conditions, algebraic gauge theories can be transferred along chain contractions. Sp…
Paper tackles skill transfer in RL for morphologically different agents.
problem Transfer skills between morphologically different reinforcement learning agents.
method Proposes a paired variational encoder-decoder model (PVED) for subspace learning.
result Demonstrates improved skill transfer efficiency compared to state-of-the-art methods.
New cohomology theory shows compact Lie group actions are Morita invariant.
problem Establishing Morita invariance for cohomology of compact Lie group actions.
method Using bibundles to transfer coefficient systems between Morita equivalent groupoids.
result Twisted Bredon-Illman cohomology is Morita invariant for compact Lie group actions.
Integrates skills and world models for efficient task solving and transfer.
problem Quickly solve new tasks in complex environments using reusable knowledge.
method Leverages partial amortization for fast adaptation and online skill planning.
result Improved sample efficiency in single tasks and transfer between tasks.
New framework for modular reinforcement learning reduces sample complexity.
problem Achieving independent credit assignment in reinforcement learning.
method Defining modular credit assignment as minimizing algorithmic mutual information, introducing modularity criterion for causal analysis.
result Single-step temporal difference action-value methods meet the modularity criterion, improving sample efficiency.
For a pinched Hadamard manifold X and a discrete group of isometries Γ of X, the critical exponent δΓ is the exponential growth rate of the orbit of a point in X under the action of Γ. We show that the critical exponent for any family N of normal subgroups of Γ0 has the same coarse behaviour…
Novel framework proves fast RL convergence in continuous spaces.
problem Analyzing stability in continuous state-action RL.
method Introduces a novel framework to analyze stability properties of RL.
result Highlights two key stability properties and demonstrates their satisfaction in RL.
Paper tackles reinforcement learning for StarCraft II, achieving high win rates.
problem Grand challenge of reinforcement learning due to huge state and action space, long-time horizon.
method Hierarchical reinforcement learning approach with two levels of abstraction.
result Achieved over 93% winning rate against level-7 AI, demonstrating strong generalization.
DARL framework tackles partial domain adaptation by selecting source instances for positive transfer.
problem Tackles the challenge of selecting source instances for positive transfer in partial domain adaptation.
method Proposes a Domain Adversarial Reinforcement Learning (DARL) framework that uses deep Q-learning and domain adversarial learning to select source instances and learn domain-invariant features.
result Demonstrates superior performance over existing methods for partial domain adaptation on several benchmark datasets.
Model improves information transfer from visual streams.
problem Challenges in unsupervised learning from continuous visual data.
method Inspired by physics, maximizes mutual information through temporal process.
result Focus of attention enhances information transfer from input stream.
Extends SORTE to multivariate risk functions.
problem Analyzing systemic risk in financial institutions or insurance-reinsurance markets.
method Develops a new framework for multivariate utility functions and applies duality theory.
result Proves existence, uniqueness, and Nash Equilibrium property of Multivariate Systemic Optimal Risk Transfer Equilibrium.
Investigates offline RL in factorisable action spaces, overcoming overestimation bias.
problem Overestimation bias in value estimates for unseen state-action pairs.
method Value-decomposition approach in DecQN, adapted for factorised discrete action spaces.
result Demonstrates the effectiveness of factorised approach in offline RL.
lamBERT learns language and actions using multimodal BERT.
problem Learning language and actions in complex environments.
method Extending BERT to multimodal representation and integrating with reinforcement learning.
result lamBERT model achieved higher rewards in multitask and transfer settings.
This work combines autoencoder transfer learning with MSCP for accurate aerodynamic predictions.
problem Data scarcity in aerodynamic modeling limits the use of high-fidelity simulations.
method Autoencoder-based transfer learning with MSCP for uncertainty-aware data fusion.
result The model achieves high accuracy with minimal high-fidelity training data and robust uncertainty bands.
Paper shows geometric properties preserved by compactifications in relation to coarse structures and group actions.
problem Geometric properties preserved by compactifications in relation to coarse structures and group actions.
method Analyzes compactifications of spaces with coarse structures and group actions, proving preservation of geometric properties.
result Geometric properties are preserved by compactifications when coarse structures and group actions are involved.
New approach discovers active learning strategies for various domains.
problem Finding efficient active learning strategies across different data types.
method Formalized annotation as a Markov decision process, designed universal state and action spaces, introduced a reward function, and used reinforcement learning to find optimal strategies.
result Learned strategies consistently outperform existing methods on multiple unrelated domains.
LEAP nets model power grid disruptions for rapid response.
problem Modeling and predicting power grid disruptions.
method Transfer learning neural network embedding approach.
result LEAP nets can rapidly assess human operators' actions in emergencies.
Perceptor Gradients learns symbolic representations from raw data.
problem Learning transferable symbolic representations from raw data.
method Decomposes policy into perceptor network and task encoding program.
result Efficiently learns symbolic representations for control tasks.
The paper proposes a method to transfer skills between tasks using disentangled latent policies.
problem Transfer learning in reinforcement learning struggles with diverse tasks without explicit supervision.
method Learning a small set of policies in a disentangled latent space that can be recombined to solve many tasks.
result Disentangled latent policies enable quick performance on many diverse tasks.
Study Lie 2-group actions on Riemannian groupoids, proving existence and developing geometric Killing vector fields.
problem Understanding isometric actions of Lie 2-groups on Riemannian groupoids.
method Exhibit properties, prove existence, construct bi-invariant metrics, provide infinitesimal description.
result Existence of 2-equivariant Slice Theorem and Equivariant Tubular Neighborhood Theorem.
RODE learns roles to simplify multi-agent tasks.
problem Efficiently discovering roles for complex multi-agent tasks.
method Clustering actions based on effects, bi-level learning hierarchy, integrating action effects into role policies.
result RODE outperforms state-of-the-art MARL algorithms on StarCraft II benchmarks.
Paper introduces a new value function for state transitions and optimal policy learning.
problem Learning optimal policies from state transitions and actions.
method Develops a forward dynamics model to maximize a novel value function Q(s,s′). result Demonstrates benefits in value function transfer, redundant action spaces, and off-policy learning.
We consider a sequential learning problem with Gaussian payoffs and side information: after selecting an action i, the learner receives information about the payoff of every action j in the form of Gaussian observations whose mean is the same as the mean payoff, but the variance depends on the pair (i,j) (and may…