Tabular Q-Learning with learned state abstractions solves continuous control tasks.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
NHC learns scalable algorithmic solutions from diverse tasks.
Algorithm finds latent structure in value functions for improved reinforcement learning.
Abstraction is a fundamental part when learning behavioral models of systems. Usually the process of abstraction is manually defined by domain experts. This paper presents a method to perform automatic abstraction for network protocols. In particular a weakly supervised clustering algorithm is used to build an abstract…
Abstraction of Markov Decision Processes is a useful tool for solving complex problems, as it can ignore unimportant aspects of an environment, simplifying the process of learning an optimal policy. In this paper, we propose a new algorithm for finding abstract MDPs in environments with continuous state spaces. It is b…
In this paper, we develop a framework to obtain graph abstractions for decision-making by an agent where the abstractions emerge as a function of the agent's limited computational resources. We discuss the connection of the proposed approach with information-theoretic signal compression, and formulate a novel optimizat…
We introduce an algorithm for model-based hierarchical reinforcement learning to acquire self-contained transition and reward models suitable for probabilistic planning at multiple levels of abstraction. We call this framework Planning with Abstract Learned Models (PALM). By representing subtasks symbolically using a n…
We present PubMed 200k RCT, a new dataset based on PubMed for sequential sentence classification. The dataset consists of approximately 200,000 abstracts of randomized controlled trials, totaling 2.3 million sentences. Each sentence of each abstract is labeled with their role in the abstract using one of the following …
Abstraction plays a key role in concept learning and knowledge discovery; this paper is concerned with computational abstraction. In particular, we study the nature of abstraction through a group-theoretic approach, formalizing it as symmetry-driven---as opposed to data-driven---hierarchical clustering. Thus, the resul…
A key feature of intelligent behavior is the ability to learn abstract strategies that transfer to unfamiliar problems. Therefore, we present a novel architecture, based on memory-augmented networks, that is inspired by the von Neumann and Harvard architectures of modern computers. This architecture enables the learnin…
Abstract Neural Networks (ANNs) improve DNN verification efficiency.
This work defines a complexity measure for BAMDP planning and introduces state abstraction for more efficient approximate planning.
We present an algorithm, HOMER, for exploration and reinforcement learning in rich observation environments that are summarizable by an unknown latent state space. The algorithm interleaves representation learning to identify a new notion of kinematic state abstraction with strategic exploration to reach new states usi…
Scalable verifier for recurrent neural networks using polyhedral abstractions.
HO2 learns options from data efficiently, improving robot manipulation tasks.
Proposes method to learn state abstractions that generalize across environments.
New method simplifies data analysis.
We develop a method to learn abstract causal graphs from interventional data.
Paper explores how analysts balance rule-based and situational aspects of data analytics.
This paper presents a way of solving Markov Decision Processes that combines state abstraction and temporal abstraction. Specifically, we combine state aggregation with the options framework and demonstrate that they work well together and indeed it is only after one combines the two that the full benefit of each is re…
Study develops a new algorithm for assessing clinical trial abstracts.
In this paper, we propose a deep multimodal fusion network to fuse multiple modalities (face, iris, and fingerprint) for person identification. The proposed deep multimodal fusion algorithm consists of multiple streams of modality-specific Convolutional Neural Networks (CNNs), which are jointly optimized at multiple fe…
In the past few years, neural abstractive text summarization with sequence-to-sequence (seq2seq) models have gained a lot of popularity. Many interesting techniques have been proposed to improve seq2seq models, making them capable of handling different challenges, such as saliency, fluency and human readability, and ge…
SPEDER extracts state-action abstraction from dynamics for reinforcement learning.
Abstracts index for ML4H workshop at NeurIPS 2019.
The paper aims to mathematically define and learn abstractions from data.
TASID learns policies in high-dimensional settings with abstract simulator knowledge.
Unified framework for causal models at different levels of abstraction.
Abstract MDPs enable strategic exploration and fast reward transfer in complex environments.
Abstract: A new approach to technical indicators without lag.
Mid-training improves RL by identifying compact action abstractions.
Improved reinforcement learning with deep learning.
Paper analyzes history-based RL methods for MDPs, introduces a theoretical framework and practical algorithm.
New approach to abstract neural network representations using renormalization group.
Study investigates how simple speech sounds can form abstract categories.
Although exploration in reinforcement learning is well understood from a theoretical point of view, provably correct methods remain impractical. In this paper we study the interplay between exploration and approximation, what we call approximate exploration. Our main goal is to further our theoretical understanding of …
A distinctive property of human and animal intelligence is the ability to form abstractions by neglecting irrelevant information which allows to separate structure from noise. From an information theoretic point of view abstractions are desirable because they allow for very efficient information processing. In artifici…
Mapper-GIN simplifies 3D point cloud classification with lightweight structure.
Temporal abstraction refers to the ability of an agent to use behaviours of controllers which act for a limited, variable amount of time. The options framework describes such behaviours as consisting of a subset of states in which they can initiate, an internal policy and a stochastic termination condition. However, mu…
The paper proposes a principle for dynamically adjusting the granularity of reinforcement learning abstractions.
Graph Neural Networks align with dynamic programming, improving algorithmic reasoning.
We prove a definable version of the Whitney embedding theorem for abstract-definable manifolds with , namely: every abstract-definable manifold is abstract-definable embedded into , for some positive integer . As a consequence, we show that every abstract-de…
This paper simplifies OPE in large state spaces using state abstractions.
Learning transferable knowledge across similar but different settings is a fundamental component of generalized intelligence. In this paper, we approach the transfer learning challenge from a causal theory perspective. Our agent is endowed with two basic yet general theories for transfer learning: (i) a task shares a c…
CIB compresses variables causally, preserving key causal interactions.
Automatic summarization of natural language is a current topic in computer science research and industry, studied for decades because of its usefulness across multiple domains. For example, summarization is necessary to create reviews such as this one. Research and applications have achieved some success in extractive …
Constellation learns group-level visual relationships for abstract reasoning.
In this work, we consider the problem of autonomously discovering behavioral abstractions, or options, for reinforcement learning agents. We propose an algorithm that focuses on the termination condition, as opposed to -- as is common -- the policy. The termination condition is usually trained to optimize a control obj…