Paper addresses underestimation bias in double Q-learning, proposing a method to improve learning performance.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Machine learning algorithms can misrepresent training data, study finds.
A new Q-learning variant reduces underestimation bias in deep reinforcement learning.
A novel Q-learning variant reduces underestimation bias in deep actor-critic methods for reinforcement learning.
Research shows bias in machine learning can be due to algorithmic flaws, not just data.
New theory shows how learning algorithms can create a bias towards negative outcomes.
New method corrects risk estimation bias, improving backtesting results.
New algorithm corrects risk estimation bias for heavy-tailed data.
LatentNN corrects neural network attenuation bias in astronomical data.
Importance-weighting is a popular and well-researched technique for dealing with sample selection bias and covariate shift. It has desirable characteristics such as unbiasedness, consistency and low computational complexity. However, weighting can have a detrimental effect on an estimator as well. In this work, we empi…
New method corrects bias in estimating entropic risk for better decision-making.
Artificial Intelligence (AI) is an important driving force for the development and transformation of the financial industry. However, with the fast-evolving AI technology and application, unintentional bias, insufficient model validation, immature contingency plan and other underestimated threats may expose the company…
Machine learning models predict brain age with systematic bias, corrected in this study.
Social bias in machine learning has drawn significant attention, with work ranging from demonstrations of bias in a multitude of applications, curating definitions of fairness for different contexts, to developing algorithms to mitigate bias. In natural language processing, gender bias has been shown to exist in contex…
From scientific experiments to online A/B testing, the previously observed data often affects how future experiments are performed, which in turn affects which data will be collected. Such adaptivity introduces complex correlations between the data and the collection procedure. In this paper, we prove that when the dat…
There is an extensive historical dataset on real GDP per capita prepared by Angus Maddison. This dataset covers the period since 1870 with continuous annual estimates in developed countries. All time series for individual economies have a clear structural break between 1940 and 1950. The behavior before 1940 and after …
Assessing the fairness of a decision making system with respect to a protected class, such as gender or race, is challenging when class membership labels are unavailable. Probabilistic models for predicting the protected class based on observable proxies, such as surname and geolocation for race, are sometimes used to …
The paper explores how the depth of neural networks affects their ability to represent data accurately.
Performance of investment managers are evaluated in comparison with benchmarks, such as financial indices. Due to the operational constraint that most professional databases do not track the change of constitution of benchmark portfolios, standard tests of performance suffer from the "look-ahead benchmark bias," when t…
Second-order methods fail to fully quantify epistemic uncertainty, leading to biased predictions.
In general, underestimation of risk is something which should be avoided as far as possible. Especially in financial asset management, equity risk is typically characterized by the measure of portfolio variance, or indirectly by quantities which are derived from it. Since there is a linear dependency of the variance an…
The estimation of risk measures recently gained a lot of attention, partly because of the backtesting issues of expected shortfall related to elicitability. In this work we shed a new and fundamental light on optimal estimation procedures of risk measures in terms of bias. We show that once the parameters of a model ne…
Copula is a powerful tool to model multivariate data. We propose the modelling of intraday financial returns of multiple assets through copula. The problem originates due to the asynchronous nature of intraday financial data. We propose a consistent estimator of the correlation coefficient in case of Elliptical copula …
A new method reduces bias in adaptive Lasso estimates.
Study compares imputation methods' effects on IML confidence intervals.
New EiV models correct bias in operator learning with noisy data.
Smaller actor-critic models lead to performance degradation and overfitting, highlighting the critic's role in value underestimation.
The paper analyzes how factorized Gaussian approximations underestimate uncertainty in variational inference.
Study evaluates the impact of academic support center's face-to-face assistance on student performance.
The paper examines skill estimation and variance under model misspecification in IRT.
The study uses a multi-armed bandit model to analyze and mitigate hiring discrimination.
Unsupervised recalibration (URC) is a general way to improve the accuracy of an already trained probabilistic classification or regression model upon encountering new data while deployed in the field. URC does not require any ground truth associated with the new field data. URC merely observes the model's predictions a…
This paper compares VaR estimation methods under tail misspecification, finding importance sampling underestimates VaR.
New model improves volatility forecasting by reducing overestimation and underestimation.
Two new methods improve block-sparse signal recovery from noisy data.
SaML guides ML models to avoid survey biases.
When using the K-nearest neighbors method, one often ignores uncertainty in the choice of K. To account for such uncertainty, Holmes and Adams (2002) proposed a Bayesian framework for K-nearest neighbors (KNN). Their Bayesian KNN (BKNN) approach uses a pseudo-likelihood function, and standard Markov chain Monte Carlo (…
We study the consistency of sample mean-variance portfolios of arbitrarily high dimension that are based on Bayesian or shrinkage estimation of the input parameters as well as weighted sampling. In an asymptotic setting where the number of assets remains comparable in magnitude to the sample size, we provide a characte…
MFVI can overestimate predictive variance compared to the exact posterior
Typically, operational risk losses are reported above some threshold. This paper studies the impact of ignoring data truncation on the 0.999 quantile of the annual loss distribution for operational risk for a broad range of distribution parameters and truncation levels. Loss frequency and severity are modelled by the P…
Bayesian model identifies health disparities in disease progression.
This article presents results from the first statistically significant study of cost escalation in transportation infrastructure projects. Based on a sample of 258 transportation infrastructure projects worth US$90 billion and representing different project types, geographical regions, and historical periods, it is fou…
Improved forecasting of financial risk using Diffusion-Copula framework.
A new method combines VI and IS to improve Bayesian inference accuracy.
Improved noise estimation in latent neural SDEs enhances model accuracy.
We investigate the possible drawbacks of employing the standard Pearson estimator to measure correlation coefficients between financial stocks in the presence of non-stationary behavior, and we provide empirical evidence against the well-established common knowledge that using longer price time series provides better, …
In recent years research on credit risk modelling has mainly focused on default probabilities. Recovery rates are usually modelled independently, quite often they are even assumed constant. Then, however, the structural connection between recovery rates and default probabilities is lost and the tails of the loss distri…
Predictive models ground many state-of-the-art developments in statistical brain image analysis: decoding, MVPA, searchlight, or extraction of biomarkers. The principled approach to establish their validity and usefulness is cross-validation, testing prediction on unseen data. Here, I would like to raise awareness on e…