New approach handles stochastic and partially-observable environments using discrete autoencoders and Monte Carlo tree search.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Proposes ABIDES-Gym for financial markets simulation.
We explore non-acyclic GFlowNets in discrete settings.
We demonstrate the use of conditional autoregressive generative models (van den Oord et al., 2016a) over a discrete latent space (van den Oord et al., 2017b) for forward planning with MCTS. In order to test this method, we introduce a new environment featuring varying difficulty levels, along with moving goals and obst…
Deep neural network learns discrete state abstractions for efficient planning.
In this paper we study a new reinforcement learning setting where the environment is non-rewarding, contains several possibly related objects of various controllability, and where an apt agent Bob acts independently, with non-observable intentions. We argue that this setting defines a realistic scenario and we present …
Introduces Conditional Action Trees to simplify RL action spaces.
We study the motion of discrete interfaces driven by ferromagnetic interactions in a two-dimensional periodic environment by coupling the minimizing movements approach by Almgren, Taylor and Wang and a discrete-to-continuous analysis. The case of a homogeneous environment has been recently treated by Braides, Gelli and…
Despite remarkable successes, Deep Reinforcement Learning (DRL) is not robust to hyperparameterization, implementation details, or small environment changes (Henderson et al. 2017, Zhang et al. 2018). Overcoming such sensitivity is key to making DRL applicable to real world problems. In this paper, we identify sensitiv…
Adversarial methods for imitation learning have been shown to perform well on various control tasks. However, they require a large number of environment interactions for convergence. In this paper, we propose an end-to-end differentiable adversarial imitation learning algorithm in a Dyna-like framework for switching be…
The distributional perspective on reinforcement learning (RL) has given rise to a series of successful Q-learning algorithms, resulting in state-of-the-art performance in arcade game environments. However, it has not yet been analyzed how these findings from a discrete setting translate to complex practical application…
D2D-SPL uses discrete states and a classifier to train RL faster.
WayDCM predicts trajectories considering long-term goals, improving accuracy.
Empowerment quantifies the influence an agent has on its environment. This is formally achieved by the maximum of the expected KL-divergence between the distribution of the successor state conditioned on a specific action and a distribution where the actions are marginalised out. This is a natural candidate for an intr…
XLVINs improve deep reinforcement learning by combining self-supervised learning and neural algorithmic reasoning.
We study the motion of discrete interfaces driven by ferromagnetic interactions in a two-dimensional low-contrast periodic environment, by coupling the minimizing movements approach by Almgren, Taylor and Wang and a discrete-to-continuum analysis. As in a recent paper by Braides and Scilla dealing with high-contrast pe…
AdaCat improves density estimation and planning in autoregressive models.
Data visualization and interaction with large data sets is known to be essential and critical in many businesses today, and the same applies to research and teaching, in this case, when exploring large and complex mathematical objects. GAP is a computer algebra system for computational discrete algebra with an emphasis…
We address the problem of Bayesian reinforcement learning using efficient model-based online planning. We propose an optimism-free Bayes-adaptive algorithm to induce deeper and sparser exploration with a theoretical bound on its performance relative to the Bayes optimal policy, with a lower computational complexity. Th…
CEA augments reinforcement learning by generating counterfactual experiences.
Bayesian approach learns causal concepts from diverse social surveys.
We investigate the performance of dynamic portfolios constructed using more than 21,000 technical trading rules on 12 categorical and country-specific markets over the 2004-2015 study period, on rolling forward structures of different lengths. We also introduce a discrete false discovery rate (DFRD+/-) method for contr…
UDRL fails to converge in stochastic environments with episodic resets.
Surrogate models speed up RL training in dynamic systems.
The study explores how agents learn and adapt preferences in dynamic environments.
The paper tackles finding optimal treatment sequences in continuous state spaces.
Paper improves Bayesian regret bounds for Thompson Sampling in reinforcement learning.
Hybrid SAC improves RL for video games with discrete, continuous actions.
The health state assessment and remaining useful life (RUL) estimation play very important roles in prognostics and health management (PHM), owing to their abilities to reduce the maintenance and improve the safety of machines or equipment. However, they generally suffer from this problem of lacking prior knowledge to …
New algorithm improves online learning with reduced discretization.
Paper robustifies reinforcement learning with risk-averse methods.
We introduce two approaches for combining neural evolution strategy (NES) and proximal policy optimization (PPO): parameter transfer and parameter space noise. Parameter transfer is a PPO agent with parameters transferred from a NES agent. Parameter space noise is to directly add noise to the PPO agent`s parameters. We…
Algorithm optimizes system design and control for better rewards.
The behavioral dynamics of multi-agent systems have a rich and orderly structure, which can be leveraged to understand these systems, and to improve how artificial agents learn to operate in them. Here we introduce Relational Forward Models (RFM) for multi-agent learning, networks that can learn to make accurate predic…
Reinforcement learning is considered to be a strong AI paradigm which can be used to teach machines through interaction with the environment and learning from their mistakes, but it has not yet been successfully used for automotive applications. There has recently been a revival of interest in the topic, however, drive…
In a wide variety of applications, humans interact with a complex environment by means of asynchronous stochastic discrete events in continuous time. Can we design online interventions that will help humans achieve certain goals in such asynchronous setting? In this paper, we address the above problem from the perspect…
Learning goal-directed behavior in environments with sparse feedback is a major challenge for reinforcement learning algorithms. The primary difficulty arises due to insufficient exploration, resulting in an agent being unable to learn robust value functions. Intrinsically motivated agents can explore new behavior for …
Widely-used deep reinforcement learning algorithms have been shown to fail in the batch setting--learning from a fixed data set without interaction with the environment. Following this result, there have been several papers showing reasonable performances under a variety of environments and batch settings. In this pape…
Efficient exploration is an unsolved problem in Reinforcement Learning which is usually addressed by reactively rewarding the agent for fortuitously encountering novel situations. This paper introduces an efficient active exploration algorithm, Model-Based Active eXploration (MAX), which uses an ensemble of forward mod…
GNMC reduces XCSF population size while preserving function approximation and policy accuracy.
Many decision-making problems naturally exhibit pronounced structures inherited from the characteristics of the underlying environment. In a Markov decision process model, for example, two distinct states can have inherently related semantics or encode resembling physical state configurations. This often implies locall…
Stochastic Q-learning tackles large action spaces with reduced computation.
The paper proposes a deep learning technique for structured and composable representations.
New approach predicts under latent shifts using high-dimensional images.
Algorithm creates synthetic experiences to enhance Deep Reinforcement Learning.
We provide bounds on control learning error in stochastic systems.
We study the problem of identifying the policy space of a learning agent, having access to a set of demonstrations generated by its optimal policy. We introduce an approach based on statistical testing to identify the set of policy parameters the agent can control, within a larger parametric policy space. After present…
We propose and study a general framework for regularized Markov decision processes (MDPs) where the goal is to find an optimal policy that maximizes the expected discounted total reward plus a policy regularization term. The extant entropy-regularized MDPs can be cast into our framework. Moreover, under our framework, …