Paper tackles overestimation bias in continuous control, improving performance by 25%.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Q-Learning overestimation bias influenced by learning rate, discount factor, and reward signal.
MFVI can overestimate predictive variance compared to the exact posterior
The breakthrough of deep Q-Learning on different types of environments revolutionized the algorithmic design of Reinforcement Learning to introduce more stable and robust algorithms, to that end many extensions to deep Q-Learning algorithm have been proposed to reduce the variance of the target values and the overestim…
Method learns neural network to overestimate reference function with guarantees.
A widely applicable Bayesian information criterion (Watanabe, 2013) is applicable for both regular and singular models in the model selection problem. This criterion tends to overestimate the log marginal likelihood. We identify an overestimating term of a widely applicable Bayesian information criterion. Adjustment of…
The paper examines various RL algorithms to address overestimation and noise issues.
Bayesian models overestimate clusters, but practical summaries can correct this.
In value-based reinforcement learning methods such as deep Q-learning, function approximation errors are known to lead to overestimated value estimates and suboptimal policies. We show that this problem persists in an actor-critic setting and propose novel mechanisms to minimize its effects on both the actor and the cr…
We correct a mistake in the published version of our paper. Our new conclusion is that the "implied leverage effect" for single stocks is underestimated by option markets for short maturities and overestimated for long maturities, while it is always overestimated for OEX options, except for the shortest maturities wher…
The paper examines skill estimation and variance under model misspecification in IRT.
Study shows statistical biases can mislead transformer models, impairing their generalization.
New model improves volatility forecasting by reducing overestimation and underestimation.
A key problem in research on adversarial examples is that vulnerability to adversarial examples is usually measured by running attack algorithms. Because the attack algorithms are not optimal, the attack algorithms are prone to overestimating the size of perturbation needed to fool the target model. In other words, the…
GUM tackles MARL by avoiding overestimation through state-marginal restriction.
Q-Distribution Guided Q-Learning corrects overestimation of uncertain OOD actions in offline RL.
Catastrophic forgetting is a critical challenge in training deep neural networks. Although continual learning has been investigated as a countermeasure to the problem, it often suffers from the requirements of additional network components and the limited scalability to a large number of tasks. We propose a novel appro…
Compensation methods correct overestimation of adversarial robustness in neural networks.
EMIX minimizes surprise in multi-agent reinforcement learning.
Investigates offline RL in factorisable action spaces, overcoming overestimation bias.
SPQR improves Q-ensemble diversity in reinforcement learning.
The study aims to prevent unfair content presentation in recommender systems.
Automates bias control in reinforcement learning algorithms.
Many real world tasks require multiple agents to work together. Multi-agent reinforcement learning (RL) methods have been proposed in recent years to solve these tasks, but current methods often fail to efficiently learn policies. We thus investigate the presence of a common weakness in single-agent RL, namely value fu…
Markov random fields (MRFs) are difficult to evaluate as generative models because computing the test log-probabilities requires the intractable partition function. Annealed importance sampling (AIS) is widely used to estimate MRF partition functions, and often yields quite accurate results. However, AIS is prone to ov…
In classical Q-learning, the objective is to maximize the sum of discounted rewards through iteratively using the Bellman equation as an update, in an attempt to estimate the action value function of the optimal policy. Conventionally, the loss function is defined as the temporal difference between the action value and…
Paper analyzes finite-time convergence of double Q-learning.
Proposes a conservative LR estimator for infrequent data near a frequency threshold.
A neural network method improves CVA computations for complex financial portfolios.
New technique prevents Q-learning collapse by maximizing diversity among ensembles.
A novel Q-learning variant reduces underestimation bias in deep actor-critic methods for reinforcement learning.
In this paper, we use a new approach to prove that the largest eigenvalue of the sample covariance matrix of a normally distributed vector is bigger than the true largest eigenvalue with probability 1 when the dimension is infinite. We prove a similar result for the smallest eigenvalue.
In the presence of a layer of metaprobabilities (from uncertainty concerning the parameters), the asymptotic tail exponent corresponds to the lowest possible tail exponent regardless of its probability. The problem explains "Black Swan" effects, i.e., why measurements tend to chronically underestimate tail contribution…
Bayesian methods often misinterpret data and asymptotic concepts.
Modern neural networks are highly non-robust against adversarial manipulation. A significant amount of work has been invested in techniques to compute lower bounds on robustness through formal guarantees and to build provably robust models. However, it is still difficult to get guarantees for larger networks or robustn…
We study the phenomenon of bias amplification in classifiers, wherein a machine learning model learns to predict classes with a greater disparity than the underlying ground truth. We demonstrate that bias amplification can arise via an inductive bias in gradient descent methods that results in the overestimation of the…
Paper addresses underestimation bias in double Q-learning, proposing a method to improve learning performance.
Suppose an investor aims at Delta hedging a European contingent claim in a jump-diffusion model, but incorrectly specifies the stock price's volatility and jump sensitivity, so that any hedging strategy is calculated under a misspecified model. When does the erroneously computed strategy super-replicate the t…
Cold posteriors improve Bayesian neural networks by reducing overestimation of aleatoric uncertainty.
A new Q-learning variant reduces underestimation bias in deep reinforcement learning.
Selective state-adaptive regularization improves offline RL performance.
What do binary (or probabilistic) forecasting abilities have to do with overall performance? We map the difference between (univariate) binary predictions, bets and "beliefs" (expressed as a specific "event" will happen/will not happen) and real-world continuous payoffs (numerical benefits or harm from an event) and sh…
We propose a robust risk measurement approach that minimizes the expectation of overestimation plus underestimation costs. We consider uncertainty by taking the supremum over a collection of probability measures, relating our approach to dual sets in the representation of coherent risk measures. We provide results that…
This article presents results from the first statistically significant study of traffic forecasts in transportation infrastructure projects. The sample used is the largest of its kind, covering 210 projects in 14 nations worth US$59 billion. The study shows with very high statistical significance that forecasters gener…
We focus on variational inference in dynamical systems where the discrete time transition function (or evolution rule) is modelled by a Gaussian process. The dominant approach so far has been to use a factorised posterior distribution, decoupling the transition function from the system states. This is not exact in gene…
Improved RL policies from offline data with relaxed BC constraints.
Information theoretic criteria (ITC) have been widely adopted in engineering and statistics for selecting, among an ordered set of candidate models, the one that better fits the observed sample data. The selected model minimizes a penalized likelihood metric, where the penalty is determined by the criterion adopted. Wh…
Small initialization improves tensor recovery from noisy data.