New reinforcement learning framework for adapting to new actions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Generically learns movement control policies from exploration data.
Novel IRL method identifies suboptimal medical decisions in ICU data.
Novel framework proves fast RL convergence in continuous spaces.
Portfolio traders strive to identify dynamic portfolio allocation schemes so that their total budgets are efficiently allocated through the investment horizon. This study proposes a novel portfolio trading strategy in which an intelligent agent is trained to identify an optimal trading action by using deep Q-learning. …
New RL method learns from state transitions without actions.
UTE improves reinforcement learning by measuring action uncertainty, enhancing policy learning efficiency.
Parameterised actions in reinforcement learning are composed of discrete actions with continuous action-parameters. This provides a framework for solving complex domains that require combining high-level actions with flexible control. The recent P-DQN algorithm extends deep Q-networks to learn over such action spaces. …
POTEC tackles off-policy learning in large action spaces, improving effectiveness.
Detecting temporal extents of human actions in videos is a challenging computer vision problem that requires detailed manual supervision including frame-level labels. This expensive annotation process limits deploying action detectors to a limited number of categories. We propose a novel method, called WSGN, that learn…
Recommender system, as an essential part of modern e-commerce, consists of two fundamental modules, namely Click-Through Rate (CTR) and Conversion Rate (CVR) prediction. While CVR has a direct impact on the purchasing volume, its prediction is well-known challenging due to the Sample Selection Bias (SSB) and Data Spars…
A novel framework interprets driving patterns using Action phases clustering.
In this paper, we describe a novel approach to imitation learning that infers latent policies directly from state observations. We introduce a method that characterizes the causal effects of latent actions on observations while simultaneously predicting their likelihood. We then outline an action alignment procedure th…
Imitation by observation is an approach for learning from expert demonstrations that lack action information, such as videos. Recent approaches to this problem can be placed into two broad categories: training dynamics models that aim to predict the actions taken between states, and learning rewards or features for com…
A scheme robust to action erasures improves MAB performance.
New model predicts drug effects across various cell types using causal imputation.
In this work we describe a novel deep reinforcement learning architecture that allows multiple actions to be selected at every time-step in an efficient manner. Multi-action policies allow complex behaviours to be learnt that would otherwise be hard to achieve when using single action selection techniques. We use both …
A new framework for structured bandits using influence diagrams and variational Thompson sampling.
The choice of the control frequency of a system has a relevant impact on the ability of reinforcement learning algorithms to learn a highly performing policy. In this paper, we introduce the notion of action persistence that consists in the repetition of an action for a fixed number of decision steps, having the effect…
Programmatic Motion Concepts learn human actions from paired videos.
New RL approach uses future state and action visitation measures for better exploration.
For a particular class of backgrounds, equations of motion for string sigma models targeted in mutually dual Poisson-Lie groups are equivalent. This phenomenon is called the Poisson-Lie T-duality. On the level of the corresponding string effective actions, the situation becomes more complicated due to the presence of t…
In recent years, deep reinforcement learning has been shown to be adept at solving sequential decision processes with high-dimensional state spaces such as in the Atari games. Many reinforcement learning problems, however, involve high-dimensional discrete action spaces as well as high-dimensional state spaces. This pa…
DFL framework improves action and outcome fairness in policy learning.
Existing imitation learning approaches often require that the complete demonstration data, including sequences of actions and states, are available. In this paper, we consider a more realistic and difficult scenario where a reinforcement learning agent only has access to the state sequences of an expert, while the expe…
Imitation learning targets deriving a mapping from states to actions, a.k.a. policy, from expert demonstrations. Existing methods for imitation learning typically require any actions in the demonstrations to be fully available, which is hard to ensure in real applications. Though algorithms for learning with unobservab…
Novel ML approach solves complex warehouse routing problem.
We introduce a novel type of stabilization map on the configuration spaces of a graph, which increases the number of particles occupying an edge. There is an induced action on homology by the polynomial ring generated by the set of edges, and we show that this homology module is finitely generated. An analogue of class…
New Lagrangian and special Lagrangian examples found in complex space.
Paper tackles RL with continuous actions and unmeasured confounders.
In this study, a novel feature coding method that exploits invariance for transformations represented by a finite group of orthogonal matrices is proposed. We prove that the group-invariant feature vector contains sufficient discriminative information when learning a linear classifier using convex loss minimization. Ba…
This paper investigates the adversarial Bandits with Knapsack (BwK) online learning problem, where a player repeatedly chooses to perform an action, pays the corresponding cost, and receives a reward associated with the action. The player is constrained by the maximum budget that can be spent to perform actions, an…
This paper optimizes slate decision systems for large action spaces.
A new RL algorithm POWR learns world models to estimate action-values.
We rephrase the problem of 3D reconstruction from images in terms of intersections of projections of orbits of custom built Lie groups actions. We then use an algorithmic method based on moving frames "a la Fels-Olver" to obtain a fundamental set of invariants of these groups actions. The invariants are used to define …
New approach for open ad hoc teamwork using graph-based policy learning.
A new method combines online and offline learning to tackle contextual bandits with missing action support.
We introduce Dynamic Planning Networks (DPN), a novel architecture for deep reinforcement learning, that combines model-based and model-free aspects for online planning. Our architecture learns to dynamically construct plans using a learned state-transition model by selecting and traversing between simulated states and…
Develops a new RL algorithm for medical treatment regimes.
Understanding procedural text requires tracking entities, actions and effects as the narrative unfolds. We focus on the challenging real-world problem of action-graph extraction from material science papers, where language is highly specialized and data annotation is expensive and scarce. We propose a novel approach, T…
New estimator improves off-policy evaluation for large action spaces.
Facial expression analysis based on machine learning requires large number of well-annotated data to reflect different changes in facial motion. Publicly available datasets truly help to accelerate research in this area by providing a benchmark resource, but all of these datasets, to the best of our knowledge, are limi…
A new method for disentangling action sequences improves model stability.
Composing previously mastered skills to solve novel tasks promises dramatic improvements in the data efficiency of reinforcement learning. Here, we analyze two recent works composing behaviors represented in the form of action-value functions and show that they perform poorly in some situations. As part of this analysi…
TOFU-POV tackles partially observed linear bandits, achieving sublinear regret with low-dimensional action vectors.
Action guidance helps agents learn true objectives in games with sparse rewards.
Decouples critic chunk length from policy to improve policy reactivity and performance.
Proposes a novel algorithm for multi-objective reinforcement learning.