The paper explores MAB strategies for very short horizons, introducing new methods and showing improved performance.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study explores strategies for randomized allocation in delayed rewards bandits.
New strategies avoid forced exploration for unimodal bandits.
We consider a scenario where an agent has multiple available strategies to explore an unknown environment. For each new interaction with the environment, the agent must select which exploration strategy to use. We provide a new strategy-agnostic method that treat the situation as a Multi-Armed Bandits problem where the…
A new strategy for identifying the best arm in Gaussian bandits with improved exploration.
New strategies improve multi-agent decision-making on irregular networks.
New method learns adaptive exploration strategies for dynamic tasks.
In this paper we consider the problem of how a reinforcement learning agent that is tasked with solving a sequence of reinforcement learning problems (a sequence of Markov decision processes) can use knowledge acquired early in its lifetime to improve its ability to solve new problems. We argue that previous experience…
Efficient exploration is one of the key challenges for reinforcement learning (RL) algorithms. Most traditional sample efficiency bounds require strategic exploration. Recently many deep RL algorithms with simple heuristic exploration strategies that have few formal guarantees, achieve surprising success in many domain…
The paper tackles pure exploration in multi-armed bandits with low rank structure using oblivious sampling.
Paper proposes efficient sample collection strategy for RL.
Proposes EE-Net for neural exploration in contextual bandits.
The article proposes optimal learning strategies for machine learning-based reliability analysis.
Meta-Reinforcement learning approaches aim to develop learning procedures that can adapt quickly to a distribution of tasks with the help of a few examples. Developing efficient exploration strategies capable of finding the most useful samples becomes critical in such settings. Existing approaches towards finding effic…
This paper explores how RL enhances HFT strategies in volatile markets.
Efficient exploration remains a major challenge for reinforcement learning. One reason is that the variability of the returns often depends on the current state and action, and is therefore heteroscedastic. Classical exploration strategies such as upper confidence bound algorithms and Thompson sampling fail to appropri…
We present an active learning architecture that allows a robot to actively learn which data collection strategy is most efficient for acquiring motor skills to achieve multiple outcomes, and generalise over its experience to achieve new outcomes. The robot explores its environment both via interactive learning and goal…
Approach collects missing outcomes to improve fairness in classification.
We describe MELEE, a meta-learning algorithm for learning a good exploration policy in the interactive contextual bandit setting. Here, an algorithm must take actions based on contexts, and learn based only on a reward signal from the action taken, thereby generating an exploration/exploitation trade-off. MELEE address…
Regularization-induced exploration improves contextual bandit performance.
A new active learning method for Gaussian process models.
This paper proposes an exploration method for deep reinforcement learning based on parameter space noise. Recent studies have experimentally shown that parameter space noise results in better exploration than the commonly used action space noise. Previous methods devised a way to update the diagonal covariance matrix o…
Paper introduces a new method for risk-sensitive investment management using RL.
Novel meta-RL strategy improves efficiency in learning novel tasks.
New model-free algorithms learn representations for low-rank MDPs efficiently.
We show how an ensemble of -functions can be leveraged for more effective exploration in deep reinforcement learning. We build on well established algorithms from the bandit setting, and adapt them to the -learning setting. We propose an exploration strategy based on upper-confidence bounds (UCB). Our experimen…
Improved RL algorithm stabilizes unknown linear systems with polynomial regret.
A new exploration strategy for contextual bandits reduces regret and is computationally efficient.
Study shows time matters in automated trading, improving simple strategies over complex ones.
Proposes qPO, a new acquisition strategy for batched Bayesian optimization that maximizes the probability of including the optimum.
We investigate statistical uncertainty quantification for reinforcement learning (RL) and its implications in exploration policy. Despite ever-growing literature on RL applications, fundamental questions about inference and error quantification, such as large-sample behaviors, appear to remain quite open. In this paper…
This paper explores portfolio management strategies to maximize alpha and minimize beta.
Efficient exploration in complex environments remains a major challenge for reinforcement learning. We propose bootstrapped DQN, a simple algorithm that explores in a computationally and statistically efficient manner through use of randomized value functions. Unlike dithering strategies such as epsilon-greedy explorat…
This paper acts as a collection of various trading strategies and useful pieces of market information that might help to implement such strategies. This list is meant to be comprehensive (though by no means exhaustive) and hence we only provide pointers and give further sources to explore each strategy further. To set …
In this paper, we propose an information-theoretic exploration strategy for stochastic, discrete multi-armed bandits that achieves optimal regret. Our strategy is based on the value of information criterion. This criterion measures the trade-off between policy information and obtainable rewards. High amounts of policy …
New approach incentivizes strategic agents to explore, making exploration almost free.
Meta-agent learns effective exploration from offline data.
Unified framework for randomized exploration in cooperative MARL.
The inference of correlated signal fields with unknown correlation structures is of high scientific and technological relevance, but poses significant conceptual and numerical challenges. To address these, we develop the correlated signal inference (CSI) algorithm within information field theory (IFT) and discuss its n…
Paper proposes a RL approach for ALM with superior performance.
Enhanced options trading strategies using advanced portfolio optimization.
In this article we explore an alternative approach to address deep exploration and we introduce the ISL algorithm, which is efficient at performing deep exploration. Similarly to maximum entropy RL, we derive the algorithm by augmenting the traditional RL objective with a novel regularization term. A distinctive featur…
SUPE combines unlabeled data with RL to efficiently explore tasks.
Deep learning has enabled traditional reinforcement learning methods to deal with high-dimensional problems. However, one of the disadvantages of deep reinforcement learning methods is the limited exploration capacity of learning agents. In this paper, we introduce an approach that integrates human strategies to increa…
We present a new algorithm that significantly improves the efficiency of exploration for deep Q-learning agents in dialogue systems. Our agents explore via Thompson sampling, drawing Monte Carlo samples from a Bayes-by-Backprop neural network. Our algorithm learns much faster than common exploration strategies such as …
Two-stage recommender systems struggle with exploration, leading to linear regret.
We consider parallel asynchronous Markov Chain Monte Carlo (MCMC) sampling for problems where we can leverage (stochastic) gradients to define continuous dynamics which explore the target distribution. We outline a solution strategy for this setting based on stochastic gradient Hamiltonian Monte Carlo sampling (SGHMC) …
Large textual corpora are often represented by the document-term frequency matrix whose elements are the frequency of terms; however, this matrix has two problems: sparsity and high dimensionality. Four dimension reduction strategies are used to address these problems. Of the four strategies, unsupervised feature trans…