A new method for faster learning in reinforcement learning.
problem Learning from multiple tasks with different goals.
method Universal Successor Representations (USR) and USR Approximator (USRA).
result Agents initialized with USRA trained on USR can achieve goals faster than random initialization.
VUSFA improves transfer learning for target-driven navigation in AI2THOR.
problem Improving transfer reinforcement learning for complex visual navigation tasks.
method Introducing SFDP and Variational Information Bottlenecks to A3C agent.
result VUSFA achieves state-of-the-art performance and generalizability.
A new approach for exploration in RL using the successor representation.
problem Developing theoretically justified algorithms for exploration in RL.
method The successor representation (SR) and substochastic successor representation (SSR) to incentivize exploration and count observations.
result An algorithm that performs as well as sample-efficient approaches and achieves state-of-the-art performance in Atari games.
USFs capture dynamics for faster RL task transfer.
problem Applying knowledge from one task to another.
method Proposed Universal Successor Features (USFs) for RL.
result USFs accelerate training and transfer knowledge.
Successor Options discovers reusable skills using landmark states.
problem Discovering reusable skills in reinforcement learning.
method Leverages Successor Representations to build a state space model and learns intra-option policies using a novel pseudo-reward.
result Demonstrates the approach's efficacy on grid-worlds and high-dimensional robotic control environments.
A model learns successor representations in uncertain environments.
problem Learning effective strategies in partially observable, noisy environments.
method Neurally plausible model using distributional successor features.
result Distributional successor features support reinforcement learning in noisy environments.
Successor Features improve transfer in RL by decoupling feature and reward.
problem Improving feature representation for task transfer in reinforcement learning.
method Decouples feature representation from reward function, allowing domain transfer.
result Advantages and limitations of Successor Features for transfer identified.
USFAs combine UVFAs, SFs, and GPI for scalable, instant RL generalisation.
problem Generalizing to unseen tasks in reinforcement learning.
method Combining universal value function approximators, successor features, and generalized policy improvement.
result Demonstrates practical benefits and transfer abilities in a complex 3D environment.
Paper introduces a new distributional successor measure for reinforcement learning.
problem Learning the distributional consequences of behavior in reinforcement learning.
method Formulates distributional successor measure as a distribution over distributions, proposes algorithm to learn it from data.
result Demonstrates zero-shot risk-sensitive policy evaluation.
Deep RL approach improves MIS for complex environments.
problem Improving off-policy evaluation for complex environments.
method Uses successor representation from deep RL to decouple reward and dynamics.
result Empirically stable and applicable to high-dimensional domains.
This research formalizes inductive generalization and proposes a new learning paradigm called Inductive Learning.
problem Generalization from easy to hard tasks, especially out-of-domain generalization.
method Formalizes inductive generalization, introduces Inductive Learning, and outlines steps to adapt techniques for learning model successors.
result A new learning paradigm (Inductive Learning) that emphasizes induction and universal properties of learning and computation.
Model Features improve transfer in reinforcement learning by clustering states.
problem Improving knowledge transfer between tasks with shared transition dynamics.
method Introduces Model Features, a feature representation that clusters behaviourally equivalent states.
result Learning Successor Features is equivalent to learning a Model-Reduction.
Learning robust value functions given raw observations and rewards is now possible with model-free and model-based deep reinforcement learning algorithms. There is a third alternative, called Successor Representations (SR), which decomposes the value function into two components -- a reward predictor and a successor ma…
Algorithm improves transfer learning by inferring successor maps.
problem Machine learning challenges in multi-task scenarios.
method Combining factorized representations and nonparametric memory-based approaches.
result Improves transfer capabilities and outperforms other algorithms.
State2vec improves RL by learning state embeddings that generalize across policies.
problem Inefficient generalization across policies in RL.
method Extends node2vec to learn state embeddings accounting for discounted future state transitions.
result Captures the geometry of the state space, leading to sample-efficient value function approximation.
Develops a new method for optimizing policies in hierarchical models.
problem Optimizing complex policies in hierarchical models.
method Applies second-order methods in the space of state-action paths.
result The natural path gradient method can be computed exactly and reflects state-space hierarchy.
New algorithm learns expert reward structures from batch data.
problem Learning expert reward structures from batch data without dynamics models.
method Deep Successor Feature Networks (DSFN) and transition-regularized imitation network.
result Superior performance on control benchmarks and sepsis management.
SR improves learning speed in dynamic environments.
problem Infeasibility of hand-modeling complex, dynamic systems.
method SR for accelerating learning in GVF-based systems.
result SR improves sample efficiency and learning speed.
This paper develops source traces for faster TD learning.
problem Improving temporal difference learning speed and generalization.
method Introduces source traces as a backward view of successor representations, enabling TD errors to be propagated to potential causal states.
result Demonstrates faster generalization and improved performance of source traces compared to previous methods.
Proto-value networks improve deep reinforcement learning representations using auxiliary tasks.
problem Improving deep reinforcement learning representations with auxiliary tasks.
method Derived a new family of auxiliary tasks based on the successor measure, combined with off-policy learning rule.
result Proto-value networks produce rich features comparable to established algorithms using only linear approximation and a small number of interactions.
VISR learns controllable features for fast task inference.
problem Generalizing behaviors beyond explicitly learned set for subsequent tasks.
method Combines Successor Features and Variational Inference.
result Achieves human-level performance on 14 Atari games.
SF-DQN improves RL transfer by learning successor features.
problem Transfer RL with shared dynamics but different reward functions.
method Decomposes Q-function into SF and reward mapping; uses GPI for policy improvement.
result SF-DQN with GPI converges faster and generalizes better than traditional RL methods.
Identifies optimal base features for zero-shot adaptation in reinforcement learning.
problem Unclear what constitutes a good set of base features for a wide range of downstream tasks.
method Identifies optimal base features based on downstream performance, without assuming downstream tasks are linear.
result Optimal base features are the same across three task families, differing from Laplacian eigenfunctions.
CAST predicts distribution-valued time series by stabilizing and transporting simplex-supported successors.
problem Forecasting distribution-valued time series with structural failure modes.
method CAST (Causal Anchored Simplex Transport) uses successors retrieved from causal context, stabilized with a persistence anchor, and locally transported on ordered supports.
result CAST outperforms baselines on eleven public and simulated benchmarks, achieving best average rank on both one-step KL and autoregressive rollout JSD.
In classical curve theory, the geometry of a curve in three dimensions is essentially characterized by their invariants, curvature and torsion. When they are given, the problem of finding a corresponding curve is known as 'solving natural equations'. Explicit solutions are known only for a handful of curve classes, inc…
New framework for efficient query-based imitation learning.
problem Aligning agent policy with human expert behavior without prior knowledge.
method Adversarial reward query with successor representation.
result Significantly outperforms uncertainty-based methods in query efficiency.
Proposes method to discover diverse near-optimal policies in reinforcement learning.
problem Finding different solutions to the same problem in reinforcement learning.
method Formalizes problem as CMDP, uses Successor Features, proposes new diversity rewards.
result Proposed method discovers diverse near-optimal policies that are robust and distinct.
In classical curve theory, the geometry of a curve in three dimensions is essentially characterized by their invariants, curvature and torsion. When they are given, the problem of finding a corresponding curve is known as 'solving natural equations'. Explicit solutions are known only for a handful of curve classes, inc…
This work improves understanding of reinforcement learning state representations.
problem Lack of precise characterization of how and when state representations generalize.
method Developed a bound on the generalization error based on effective dimension.
result Bound quantifies the tension between generalization and approximation.
Develops a bialgebra theory for post-Lie algebras using geometric interpretations and bilinear forms.
problem Characterizing and understanding post-Lie algebras and their associated structures.
method Utilizes Manin triples and generalized Hessian Lie groups to define and characterize post-Lie algebras with nondegenerate symmetric invariant bilinear forms.
result Establishes a bialgebra theory for post-Lie algebras via the Manin triple approach, including new algebraic structures like pp-post-Lie algebras.
Our work proves CSF can recover ground-truth features in RL, improving understanding of feature learning.
problem Understanding the role of representation and mutual information in reinforcement learning.
method Investigates Contrastive Successor Features (CSF) method for identifiable representation learning in reinforcement learning.
result Proves CSF can recover ground-truth features up to a linear transformation.
Drinfel'd used associators to construct families of universal representations of braid groups. We consider semi-associators (i.e., we drop the pentagonal axiom and impose a normalization in degree one). We show that the process may be reversed, to obtain semi-associators from universal representations of 3-braids. We v…
Study extends Vogel's universality to torus knots in adjoint representation.
problem Applying Vogel's universality to knot invariants in adjoint representation theory.
method Extending Vogel's parameters to include torus knots T[m,n] and focusing on T[4,n] with odd n. result Unified description of adjoint invariants for torus knots T[4,n] with odd n. We present a universal knot polynomials for 2- and 3-strand torus knots in adjoint representation, by universalization of appropriate Rosso-Jones formula. According to universality, these polynomials coincide with adjoined colored HOMFLY and Kauffman polynomials at SL and SO/Sp lines on Vogel's plane, and give their ex…
Based on the analogies between knot theory and number theory, we study a deformation theory for SL_2-representations of knot groups, following after Mazur's deformation theory of Galois representations. Firstly, by employing the pseudo-SL_2-representations, we prove the existence of the universal deformation of a given…
New method improves data efficiency in reinforcement learning by composing skills.
problem Improving data efficiency in reinforcement learning by composing previously mastered skills.
method Extending policy improvement to maximum entropy framework, introducing successor features, and explicitly learning divergence between base policies.
result Proposes a novel approach that outperforms or matches existing methods in various tasks.
Study homeomorphism groups of ordinals, proving strong distortion and normal generators.
problem Understanding algebraic and geometric properties of homeomorphism groups of ordinals.
method Analyzing successor ordinals with connections to permutation groups and manifolds.
result Proves strong distortion and normal generators for homeomorphism groups of ordinals.
SU improves exploration in reinforcement learning, surpassing human performance on Atari games.
problem Challenges in scaling PSRL for reinforcement learning with neural networks.
method Design and implementation of Successor Uncertainties (SU) algorithm.
result SU outperforms human performance on Atari games and surpasses RVF competitor Bootstrapped DQN.
Revives Vogel's diagrammatic technique for universal Lie algebra computations.
problem The universality of Lie algebra quantities remains open, despite many being described.
method Diagrammatic algebra based on Vogel's Λ-algebra.
result Diagrammatic technique enables truly universal computations in Lie theory.
New universal automorphic functions capture monstrous moonshine.
problem Developing a universal framework for automorphic functions.
method Reformulating old results, constructing new coordinates, and defining central extensions.
result New invariant 1-forms and representations for universal Teichmüller space.
The paper shows how neural networks can approximate PDEs with polynomial scaling in dimension.
problem Understanding the complexity of approximating PDE solutions with neural networks.
method Developed a proof technique to simulate gradient descent using neural networks.
result Neural network parameters scale polynomially with input dimension for approximating PDE solutions.
Universal connection constructed using diffeology theory.
problem Natural connection on bundles of paths on manifolds.
method Diffeological construction of Singer's universal connection.
result Functorial equivalence between holonomy categories and diffeological bundle-connection pairs.
We study the twisted knot module for the universal deformation of an SL2-representation of a knot group, and introduce an associated L-function, which may be seen as an analogue of the algebraic p-adic L-function associated to the Selmer module for the universal deformation of a Galois representation. We…
By now it is well established that the quantum dimensions of descendants of the adjoint representation can be described in a universal form, independent of a particular family of simple Lie algebras. The Rosso-Jones formula then implies a universal description of the adjoint knot polynomials for torus knots, which in p…
This work decouples language from math problems to enable cross-language learning.
problem Current machine learning representations are language dependent.
method Inspired by linguistics, the work learns language agnostic representations.
result Models trained on one language achieve similar accuracies in other languages.
Sharp lower bound on GHHs' representation power of CPWL functions.
problem Proving the minimum number of nestings for GHHs to represent arbitrary CPWL functions.
method Using a key lemma about finite sums of periodic functions, proving necessity of n nestings.
result Proving necessity of n nestings for GHHs to achieve universal representation power.
Paper characterizes and constructs universal approximators for neural networks.
problem Limited understanding of universal approximation in neural networks.
method Characterization, representation, construction method, existence result for any universal approximator.
result Improved capabilities of feed-forward architecture to approximate continuous functions.
Universal neural networks learn across diverse vision tasks.
problem Machine vision systems lack a universal representation.
method Investigated neural networks' capacity across various vision domains.
result Single neural network performs as well as specialized networks on multiple domains.