MB-DQN uses different backup lengths for improved reinforcement learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Multi-step methods such as Retrace() and -step -learning have become a crucial component of modern deep reinforcement learning agents. These methods are often evaluated as a part of bigger architectures and their evaluations rarely include enough samples to draw statistically significant conclusions about thei…
A new algorithm reduces training time for distributed machine learning by dynamically assigning backup workers.
Decouples critic chunk length from policy to improve policy reactivity and performance.
The study proposes using TD error for selecting σ in Q(σ, λ).
The project intends to model the stability of power system with a deep learning algorithm to the problem, aiming to delay the removal of the fault. The so-called "fail-delay cut-off" refers to the occurrence of N-1 backup protection action on the backbone network of the system, resulting in longer time for the removal …
A reinforcement learning framework combining value function and tree search planner for strategic and tactical decisions.
Modeling supply chain disruptions from climate hazards with adaptive firms.
New RL approach infers optimal policies via variational inference.
Modern reinforcement learning algorithms reach super-human performance on many board and video games, but they are sample inefficient, i.e. they typically require significantly more playing experience than humans to reach an equal performance level. To improve sample efficiency, an agent may build a model of the enviro…
Off-policy reinforcement learning aims to leverage experience collected from prior policies for sample-efficient learning. However, in practice, commonly used off-policy approximate dynamic programming methods based on Q-learning and actor-critic methods are highly sensitive to the data distribution, and can make only …
Reinforcement learning has attracted great attention recently, especially policy gradient algorithms, which have been demonstrated on challenging decision making and control tasks. In this paper, we propose an active multi-step TD algorithm with adaptive stepsizes to learn actor and critic. Specifically, our model cons…
Recently, a new multi-step temporal learning algorithm, called , unifies -step Tree-Backup (when ) and -step Sarsa (when ) by introducing a sampling parameter . However, similar to other multi-step temporal-difference learning algorithms, needs much memory consumption and computation tim…
MLaaS (ML-as-a-Service) offerings by cloud computing platforms are becoming increasingly popular. Hosting pre-trained machine learning models in the cloud enables elastic scalability as the demand grows. But providing low latency and reducing the latency variance is a key requirement. Variance is harder to control in a…
Combining deep model-free reinforcement learning with on-line planning is a promising approach to building on the successes of deep RL. On-line planning with look-ahead trees has proven successful in environments where transition models are known a priori. However, in complex environments where transition models need t…
Paper studies offline RL with linear approx, focusing on inherent Bellman error.
Value-based reinforcement learning (RL) methods like Q-learning have shown success in a variety of domains. One challenge in applying Q-learning to continuous-action RL problems, however, is the continuous action maximization (max-Q) required for optimal Bellman backup. In this work, we develop CAQL, a (class of) algor…
The paper analyzes off-policy TD-learning using generalized Bellman operators and provides finite-sample bounds.
Reducing the latency variance in machine learning inference is a key requirement in many applications. Variance is harder to control in a cloud deployment in the presence of stragglers. In spite of this challenge, inference is increasingly being done in the cloud, due to the advent of affordable machine learning as a s…
Off-policy reinforcement learning with eligibility traces is challenging because of the discrepancy between target policy and behavior policy. One common approach is to measure the difference between two policies in a probabilistic way, such as importance sampling and tree-backup. However, existing off-policy learning …
Reinforcement learning methods carry a well known bias-variance trade-off in n-step algorithms for optimal control. Unfortunately, this has rarely been addressed in current research. This trade-off principle holds independent of the choice of the algorithm, such as n-step SARSA, n-step Expected SARSA or n-step Tree bac…
SUNRISE improves off-policy RL algorithms by integrating ensemble methods.
QAM uses adjoint matching to optimize continuous-action RL policies efficiently.
Reinforcement learning is a promising approach to synthesizing policies for challenging robotics tasks. A key problem is how to ensure safety of the learned policy---e.g., that a walking robot does not fall over or that an autonomous car does not run into an obstacle. We focus on the setting where the dynamics are know…
Simplifies BCQ to match and outperform state-of-the-art in offline RL benchmarks.
FOCAL tackles offline meta-reinforcement learning with efficient task inference and behavior regularization.
RiskNet predicts penalties in unreliable communication networks using GNNs.
Paper proposes a second-order method for faster SVI convergence.
Improved risk-sensitive RL with exponential Bellman equation and better regret bounds.
Many currently deployed Reinforcement Learning agents work in an environment shared with humans, be them co-workers, users or clients. It is desirable that these agents adjust to people's preferences, learn faster thanks to their help, and act safely around them. We argue that most current approaches that learn from hu…
Safe autonomous decisions made with machine learning predictions using Conformal Decision Theory.
Knoop enhances variable selection with over-parameterization and knockoffs.
Adversarial training is a principled approach for training robust neural networks. Despite of tremendous successes in practice, its theoretical properties still remain largely unexplored. In this paper, we provide new theoretical insights of gradient descent based adversarial training by studying its computational prop…
Q-chunking improves RL for long tasks by chunking actions.
Gradient descent biases towards stable rank networks for nearly-orthogonal data.
TT-DAC-PS: A deterministic actor-critic approach for optimal trade execution
Hydropower reduces system electricity price and volatility, especially at extreme levels.
Work addresses long-term accuracy issues in IoT air quality sensors.
Machine learning (ML) techniques are increasingly applied to decision-making and control problems in Cyber-Physical Systems among which many are safety-critical, e.g., chemical plants, robotics, autonomous vehicles. Despite the significant benefits brought by ML techniques, they also raise additional safety issues beca…
Paper introduces a method to evaluate abstaining classifiers by considering missing predictions as counterfactuals.
New algorithm learns sparse linear MDPs with polynomial interactions, improving sample complexity.
New method calculates Shapley values for uncertain functions.
New set-valued star-shaped risk measures introduced for better risk assessment.
The paper introduces Absolute Shapley Value to handle negative contributions in machine learning model training.
Study on 2-valued dynamics on complex plane, showing some dynamics can't be group actions.
Formula for Z_2-valued index of symmetric operators on manifolds.
New method converts p-values to e-values for more efficient CP and aggregation.
Introduces joint Shapley values to measure feature importance in models.