Optimized RL algorithms perform well on offline datasets, outperforming fully trained agents.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Reverse Experience Replay improves Deep Q-learning for sparse rewards.
Algorithm creates synthetic experiences to enhance Deep Reinforcement Learning.
This project combines recent advances in experience replay techniques, namely, Combined Experience Replay (CER), Prioritized Experience Replay (PER), and Hindsight Experience Replay (HER). We show the results of combinations of these techniques with DDPG and DQN methods. CER always adds the most recent experience to th…
A simple DQN-based multi-agent RL system for binary actions.
Adaptive synchronization improves deep reinforcement learning performance.
Study shows DQN's performance degrades with temporal dependence in data.
Experience replay is widely used in deep reinforcement learning algorithms and allows agents to remember and learn from experiences from the past. In an effort to learn more efficiently, researchers proposed prioritized experience replay (PER) which samples important transitions more frequently. In this paper, we propo…
A new method prioritizes and recycles experiences for better reinforcement learning.
Extends XCS with Experience Replay for improved sample efficiency in single-step tasks.
Despite the great empirical success of deep reinforcement learning, its theoretical foundation is less well understood. In this work, we make the first attempt to theoretically understand the deep Q-network (DQN) algorithm (Mnih et al., 2015) from both algorithmic and statistical perspectives. In specific, we focus on …
Enhanced DQN model boosts trading performance with advanced techniques.
We propose Deep Q-Networks (DQN) with model-based exploration, an algorithm combining both model-free and model-based approaches that explores better and learns environments with sparse rewards more efficiently. DQN is a general-purpose, model-free algorithm and has been proven to perform well in a variety of tasks inc…
A critical and challenging problem in reinforcement learning is how to learn the state-action value function from the experience replay buffer and simultaneously keep sample efficiency and faster convergence to a high quality solution. In prior works, transitions are uniformly sampled at random from the replay buffer o…
Modern deep reinforcement learning methods have departed from the incremental learning required for eligibility traces, rendering the implementation of the -return difficult in this context. In particular, off-policy methods that utilize experience replay remain problematic because their random sampling of minibatch…
This paper investigates learning sparse representations and action-value functions simultaneously in deep reinforcement learning.
A lightweight FPGA-based reinforcement learning approach for edge devices.
We propose Episodic Backward Update (EBU) - a novel deep reinforcement learning algorithm with a direct value propagation. In contrast to the conventional use of the experience replay with uniform random sampling, our agent samples a whole episode and successively propagates the value of a state to its previous states.…
Optimal trade execution is an important problem faced by essentially all traders. Much research into optimal execution uses stringent model assumptions and applies continuous time stochastic control to solve them. Here, we instead take a model free approach and develop a variation of Deep Q-Learning to estimate the opt…
DQN outperforms static policies in a dynamic fee environment for automated market makers.
Langevin DQN achieves deep exploration using Gaussian noise.
NROWAN-DQN improves stability and exploration in noisy networks.
A novel framework optimizes experience replay for reinforcement learning.
MB-DQN uses different backup lengths for improved reinforcement learning.
The deep Q-network (DQN) and return-based reinforcement learning are two promising algorithms proposed in recent years. DQN brings advances to complex sequential decision problems, while return-based algorithms have advantages in making use of sample trajectories. In this paper, we propose a general framework to combin…
New modifiers improve noisy RNN replay in hippocampal networks.
Learning an effective representation for high-dimensional data is a challenging problem in reinforcement learning (RL). Deep reinforcement learning (DRL) such as Deep Q networks (DQN) achieves remarkable success in computer games by learning deeply encoded representation from convolution networks. In this paper, we pro…
New insights into experience replay in RL algorithms.
DQNs can approximate optimal Q-functions with high accuracy on compact sets.
Efficient exploration in complex environments remains a major challenge for reinforcement learning. We propose bootstrapped DQN, a simple algorithm that explores in a computationally and statistically efficient manner through use of randomized value functions. Unlike dithering strategies such as epsilon-greedy explorat…
Deep reinforcement learning algorithms have shown an impressive ability to learn complex control policies in high-dimensional tasks. However, despite the ever-increasing performance on popular benchmarks, policies learned by deep reinforcement learning algorithms can struggle to generalize when evaluated in remarkably …
We present a new replay-based method of continual classification learning that we term "conditional replay" which generates samples and labels together by sampling from a distribution conditioned on the class. We compare conditional replay to another replay-based continual learning paradigm (which we term "marginal rep…
RS-DQN protects RL agents from adversarial attacks.
This paper improves DQN agents' robustness to adversarial perturbations.
Curious Replay improves model-based reinforcement learning agents' adaptability.
A new RL method improves performance on Atari games without complex techniques.
SF-DQN improves RL transfer by learning successor features.
We present a novel algorithm to train a deep Q-learning agent using natural-gradient techniques. We compare the original deep Q-network (DQN) algorithm to its natural-gradient counterpart, which we refer to as NGDQN, on a collection of classic control domains. Without employing target networks, NGDQN significantly outp…
ReaPER improves learning efficiency by prioritizing reliable experiences.
Unified study of stateful replay for streaming learning, reducing forgetting by 2-3x.
In this paper, we propose a replay attack spoofing detection system for automatic speaker verification using multitask learning of noise classes. We define the noise that is caused by the replay attack as replay noise. We explore the effectiveness of training a deep neural network simultaneously for replay attack spoof…
Combines Hebbian and DQN for better POMDP problem solving.
Parameterised actions in reinforcement learning are composed of discrete actions with continuous action-parameters. This provides a framework for solving complex domains that require combining high-level actions with flexible control. The recent P-DQN algorithm extends deep Q-networks to learn over such action spaces. …
Generative replay improves continual learning by using generated data as negative examples.
Efficient actor-critic learning with shared experience replay improves data efficiency.
Recent research has shown that although Reinforcement Learning (RL) can benefit from expert demonstration, it usually takes considerable efforts to obtain enough demonstration. The efforts prevent training decent RL agents with expert demonstration in practice. In this work, we propose Active Reinforcement Learning wit…
The paper improves reinforcement learning stability and efficiency with a new theoretical framework.
DQN outperforms traditional stock market strategies by 30%.