Proposes a method to enhance exploration in RL using temporal difference uncertainties.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Efficient exploration is necessary to achieve good sample efficiency for reinforcement learning in general. From small, tabular settings such as gridworlds to large, continuous and sparse reward settings such as robotic object manipulation tasks, exploration through adding an uncertainty bonus to the reward function ha…
This paper studies directed exploration for reinforcement learning agents by tracking uncertainty about the value of each available action. We identify two sources of uncertainty that are relevant for exploration. The first originates from limited data (parametric uncertainty), while the second originates from the dist…
New method quantifies uncertainty in reinforcement learning models.
A simple uncertainty measure improves deep bandit performance.
Efficient exploration improves large language model performance with fewer queries.
REN addresses uncertainty in user feedbacks for better recommendation systems.
Proposes EVE for efficient exploration in reinforcement learning.
New algorithm improves graph-based active learning by identifying unexplored regions.
BO algorithms improve binary and preferential optimization by distinguishing between types of uncertainty.
Model based predictions of future trajectories of a dynamical system often suffer from inaccuracies, forcing model based control algorithms to re-plan often, thus being computationally expensive, suboptimal and not reliable. In this work, we propose a model agnostic method for estimating the uncertainty of a model?s pr…
We consider an agent's uncertainty about its environment and the problem of generalizing this uncertainty across observations. Specifically, we focus on the problem of exploration in non-tabular reinforcement learning. Drawing inspiration from the intrinsic motivation literature, we use density models to measure uncert…
New approach uses autoregressive models to explore and quantify uncertainty in decision-making.
Posterior sampling for reinforcement learning (PSRL) is an effective method for balancing exploration and exploitation in reinforcement learning. Randomised value functions (RVF) can be viewed as a promising approach to scaling PSRL. However, we show that most contemporary algorithms combining RVF with neural network f…
Accurately estimating uncertainties in neural network predictions is of great importance in building trusted DNNs-based models, and there is an increasing interest in providing accurate uncertainty estimation on many tasks, such as security cameras and autonomous driving vehicles. In this paper, we focus on the two mai…
Thompson sampling is an efficient algorithm for sequential decision making, which exploits the posterior uncertainty to address the exploration-exploitation dilemma. There has been significant recent interest in integrating Bayesian neural networks into Thompson sampling. Most of these methods rely on global variable u…
This paper tackles uncertainty in deep learning for construction of prediction intervals.
Proposes H-UCRL for efficient model-based RL with sublinear regret.
We study the problem of safe learning and exploration in sequential control problems. The goal is to safely collect data samples from operating in an environment, in order to learn to achieve a challenging control goal (e.g., an agile maneuver close to a boundary). A central challenge in this setting is how to quantify…
Graph Posterior Network improves uncertainty estimation for node classification in interdependent graphs.
This paper tackles combinatorial optimization under uncertainty with limited feedback.
Algorithms that tackle deep exploration -- an important challenge in reinforcement learning -- have relied on epistemic uncertainty representation through ensembles or other hypermodels, exploration bonuses, or visitation count distributions. An open question is whether deep exploration can be achieved by an incrementa…
We consider the exploration/exploitation problem in reinforcement learning. For exploitation, it is well known that the Bellman equation connects the value at any time-step to the expected value at subsequent time-steps. In this paper we consider a similar \textit{uncertainty} Bellman equation (UBE), which connects the…
The paper explores how prior functions and bootstrapping improve ensemble uncertainty estimation.
A new multi-objective RL framework improves intrinsic exploration performance.
Reinforcement learning agents are faced with two types of uncertainty. Epistemic uncertainty stems from limited data and is useful for exploration, whereas aleatoric uncertainty arises from stochastic environments and must be accounted for in risk-sensitive applications. We highlight the challenges involved in simultan…
Study designs steering rewards for MFGs with unknown dynamics and model uncertainty.
Efficient exploration remains a challenging problem in reinforcement learning, especially for those tasks where rewards from environments are sparse. A commonly used approach for exploring such environments is to introduce some "intrinsic" reward. In this work, we focus on model uncertainty estimation as an intrinsic r…
Estimates of predictive uncertainty are important for accurate model-based planning and reinforcement learning. However, predictive uncertainties---especially ones derived from modern deep learning systems---can be inaccurate and impose a bottleneck on performance. This paper explores which uncertainties are needed for…
It is well known that quantifying uncertainty in the action-value estimates is crucial for efficient exploration in reinforcement learning. Ensemble sampling offers a relatively computationally tractable way of doing this using randomized value functions. However, it still requires a huge amount of computational resour…
DEUP directly predicts epistemic uncertainty, improving model optimization and exploration.
In statistical dialogue management, the dialogue manager learns a policy that maps a belief state to an action for the system to perform. Efficient exploration is key to successful policy optimisation. Current deep reinforcement learning methods are very promising but rely on epsilon-greedy exploration, thus subjecting…
Deep Bayesian Bandits improve personalized ads by balancing exploration and exploitation.
This paper explores Bayesian Neural Network posteriors, uncovering symmetries and their impact.
The study explores machine learning for predicting customer propensity-to-pay uncertainty.
We study dynamic allocation problems for discrete time multi-armed bandits under uncertainty, based on the the theory of nonlinear expectations. We show that, under strong independence of the bandits and with some relaxation in the definition of optimality, a Gittins allocation index gives optimal choices. This involve…
New method improves GP uncertainty quantification for misspecified priors.
We investigate statistical uncertainty quantification for reinforcement learning (RL) and its implications in exploration policy. Despite ever-growing literature on RL applications, fundamental questions about inference and error quantification, such as large-sample behaviors, appear to remain quite open. In this paper…
In distributional reinforcement learning (RL), the estimated distribution of value function models both the parametric and intrinsic uncertainties. We propose a novel and efficient exploration method for deep RL that has two components. The first is a decaying schedule to suppress the intrinsic uncertainty. The second …
The investor is interested in the expected return and he is also concerned about the risk and the uncertainty assumed by the investment. One of the most popular concepts used to measure the risk and the uncertainty is the variance and/or the standard-deviation. In this paper we explore the following issues: Is the stan…
D2D converts CLDs into SDMs to explore leverage points under uncertainty.
Paper optimizes industrial refrigeration using adaptive exploration.
Proposes UICR to improve novelty in recommendation systems without sacrificing relevance.
Study values and optimizes forestry leases under risk and uncertainty.
Large models collapse epistemic uncertainty, challenging traditional wisdom.
Trading off exploration and exploitation in an unknown environment is key to maximising expected return during learning. A Bayes-optimal policy, which does so optimally, conditions its actions not only on the environment state but on the agent's uncertainty about the environment. Computing a Bayes-optimal policy is how…
This paper applies quantum theory to cost accounting, focusing on WIP valuation.
Hypermodels improve exploration efficiency and accuracy.