VUSFA improves transfer learning for target-driven navigation in AI2THOR.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
USFs capture dynamics for faster RL task transfer.
The ability of a reinforcement learning (RL) agent to learn about many reward functions at the same time has many potential benefits, such as the decomposition of complex tasks into simpler ones, the exchange of information between tasks, and the reuse of skills. We focus on one aspect in particular, namely the ability…
One question central to Reinforcement Learning is how to learn a feature representation that supports algorithm scaling and re-use of learned information from different tasks. Successor Features approach this problem by learning a feature representation that satisfies a temporal constraint. We present an implementation…
Identifies optimal base features for zero-shot adaptation in reinforcement learning.
The objective of transfer reinforcement learning is to generalize from a set of previous tasks to unseen new tasks. In this work, we focus on the transfer scenario where the dynamics among tasks are the same, but their goals differ. Although general value function (Sutton et al., 2011) has been shown to be useful for k…
Animals need to devise strategies to maximize returns while interacting with their environment based on incoming noisy sensory observations. Task-relevant states, such as the agent's location within an environment or the presence of a predator, are often not directly observable but must be inferred using available sens…
This research formalizes inductive generalization and proposes a new learning paradigm called Inductive Learning.
It has been established that diverse behaviors spanning the controllable subspace of an Markov decision process can be trained by rewarding a policy for being distinguishable from other policies \citep{gregor2016variational, eysenbach2018diversity, warde2018unsupervised}. However, one limitation of this formulation is …
SF-DQN improves RL transfer by learning successor features.
A key question in Reinforcement Learning is which representation an agent can learn to efficiently reuse knowledge between different tasks. Recently the Successor Representation was shown to have empirical benefits for transferring knowledge between tasks with shared transition dynamics. This paper presents Model Featu…
Proposes method to discover diverse near-optimal policies in reinforcement learning.
State2vec improves RL by learning state embeddings that generalize across policies.
In this paper we introduce a simple approach for exploration in reinforcement learning (RL) that allows us to develop theoretically justified algorithms in the tabular case but that is also extendable to settings where function approximation is required. Our approach is based on the successor representation (SR), which…
We introduce a novel apprenticeship learning algorithm to learn an expert's underlying reward structure in off-policy model-free \emph{batch} settings. Unlike existing methods that require a dynamics model or additional data acquisition for on-policy evaluation, our algorithm requires only the batch data of observed ex…
Learning robust value functions given raw observations and rewards is now possible with model-free and model-based deep reinforcement learning algorithms. There is a third alternative, called Successor Representations (SR), which decomposes the value function into two components -- a reward predictor and a successor ma…
The options framework in reinforcement learning models the notion of a skill or a temporally extended sequence of actions. The discovery of a reusable set of skills has typically entailed building options, that navigate to bottleneck states. This work adopts a complementary approach, where we attempt to discover option…
Paper introduces a new distributional successor measure for reinforcement learning.
Deep RL approach improves MIS for complex environments.
CAST predicts distribution-valued time series by stabilizing and transporting simplex-supported successors.
In classical curve theory, the geometry of a curve in three dimensions is essentially characterized by their invariants, curvature and torsion. When they are given, the problem of finding a corresponding curve is known as 'solving natural equations'. Explicit solutions are known only for a handful of curve classes, inc…
New framework for efficient query-based imitation learning.
In classical curve theory, the geometry of a curve in three dimensions is essentially characterized by their invariants, curvature and torsion. When they are given, the problem of finding a corresponding curve is known as 'solving natural equations'. Explicit solutions are known only for a handful of curve classes, inc…
Our work proves CSF can recover ground-truth features in RL, improving understanding of feature learning.
Proto-value networks improve deep reinforcement learning representations using auxiliary tasks.
Develops a new method for optimizing policies in hierarchical models.
Posterior sampling for reinforcement learning (PSRL) is an effective method for balancing exploration and exploitation in reinforcement learning. Randomised value functions (RVF) can be viewed as a promising approach to scaling PSRL. However, we show that most contemporary algorithms combining RVF with neural network f…
Paper relaxes symmetry conditions for universal feature selection in noisy data.
Study homeomorphism groups of ordinals, proving strong distortion and normal generators.
Composing previously mastered skills to solve novel tasks promises dramatic improvements in the data efficiency of reinforcement learning. Here, we analyze two recent works composing behaviors represented in the form of action-value functions and show that they perform poorly in some situations. As part of this analysi…
We study the approximation properties of random ReLU features through their reproducing kernel Hilbert space (RKHS). We first prove a universality theorem for the RKHS induced by random features whose feature maps are of the form of nodes in neural networks. The universality result implies that the random ReLU features…
The generative learning phase of Autoencoder (AE) and its successor Denosing Autoencoder (DAE) enhances the flexibility of data stream method in exploiting unlabelled samples. Nonetheless, the feasibility of DAE for data stream analytic deserves in-depth study because it characterizes a fixed network capacity which can…
The study proves Gaussian universality of deep random features learning.
The paper identifies universal features for high-dimensional data inference.
UniFeat is an open-source Java tool for feature selection.
Quantum kernels can be efficiently embedded into classical feature spaces.
Quantum machine learning models can approximate any continuous function.
Path signatures adapted for Lie groups improve action recognition in computer vision.
Tab-TRM uses recursive model for insurance pricing on tabular data.
Unified theorem for deep and shallow joint-equivariant machines.
Random feature models approximate functions in Banach spaces efficiently.
Phishing as one of the most well-known cybercrime activities is a deception of online users to steal their personal or confidential information by impersonating a legitimate website. Several machine learning-based strategies have been proposed to detect phishing websites. These techniques are dependent on the features …
TabPFN-2.5 boosts tabular AI performance, especially for large datasets.
Effective feature representation is key to the predictive performance of any algorithm. This paper introduces a meta-procedure, called Non-Euclidean Upgrading (NEU), which learns feature maps that are expressive enough to embed the universal approximation property (UAP) into most model classes while only outputting fea…
We present a probabilistic model of events in continuous time in which each event triggers a Poisson process of successor events. The ensemble of observed events is thereby modeled as a superposition of Poisson processes. Efficient inference is feasible under this model with an EM algorithm. Moreover, the EM algorithm …
New insights into model robustness for random features and NTK models.
Transformers enable in-context learning with guarantees for a wide range of tasks.
Here we propose using the successor representation (SR) to accelerate learning in a constructive knowledge system based on general value functions (GVFs). In real-world settings like robotics for unstructured and dynamic environments, it is infeasible to model all meaningful aspects of a system and its environment by h…