Study proposes an alternative method to measure societal biases using smoothed co-occurrence relations.
problem Measuring societal biases using word embeddings can introduce irrelevant concepts.
method Proposes an alternative approach using smoothed first-order co-occurrence relations.
result First-order approach shows higher correlations with actual gender bias statistics.
The study examines methods to correct measurement error in nutritional epidemiology studies.
problem Measurement error in nutritional studies leads to biased and underconfident estimates.
method The article reviews various bias-correction models for exposure variables in nutritional epidemiology.
result Bias-correction methods are essential for accurate inference in nutritional studies.
The Pinned AUC metric hides unintended bias when class distributions vary.
problem Unintended bias in classification models.
method Examines the Pinned AUC metric and its limitations.
result Pinned AUC can obscure different types of unintended bias.
Corrects bias in feature importance measures of tree-based methods.
problem Bias in feature importance measures of tree-based methods.
method Corrects bias by incorporating out-of-sample split-improvement.
result Better summaries and screening tools of feature importance.
Study shows how to reduce variational inference bias by concentrating likelihood ratio distribution.
problem Bias and variance issues in variational inference.
method Upper bound variational gap using dispersion measure of likelihood ratio, suggesting methods to reduce bias.
result Reducing bias in variational inference can be achieved by making likelihood ratio distribution more concentrated.
Study uncovers bias in image classification models using attribution maps.
problem Data bias in image classification models.
method Created an artificial dataset with known bias, trained CNNs, and used attribution maps to inspect decisions.
result Different attribution map techniques highlight bias better than others, and metrics support bias identification.
Study measures gender bias in machine translation using multiple reference points.
problem Measuring and identifying gender bias in machine translation.
method Used an optimal non-biased translator, reference points from occupational statistics and survey.
result Found bias against both genders, but more against women, and found occupations have a greater effect than adjectives.
Bias is essential for machine learning success, quantifiable and conserved.
problem The necessity and quantification of bias in machine learning success.
method Quantifying bias relative to possible datasets and demonstrating its role in increasing success probability.
result Bias is a conserved quantity and essential for favorably biasing towards a fixed target.
AI bias arises from human-defined goals, not algorithmic flaws.
problem AI bias due to human-defined goals in LLMs.
method Purpose-conditioned cognition and revealing downstream use of LLM outputs.
result AI bias can be reduced by purpose-aware prompting but not fully by regularization.
Machine learning forecasts show bias at long horizons, contrary to standard tests.
problem Forecast efficiency tests misinterpret machine learning performance.
method Theoretical and empirical analysis of regularization and measurement noise.
result Machine learning forecasts exhibit overreaction at longer horizons, not bias.
This paper analyzes the bias of inexact MCMC methods in high dimensions.
problem Understanding the bias of inexact MCMC methods in high-dimensional spaces.
method Establishing bounds on Wasserstein distances between inexact MCMC methods and target distributions.
result The asymptotic bias of ULA and uHMC depends on key quantities related to the target distribution or the stationary probability measure of the scheme.
New model reduces bias in cosmic shear measurements.
problem Bias in cosmic shear measurements due to non-well-defined ellipticity.
method Hybrid physical and deep learning Hierarchical Bayesian Model.
result Unbiased estimate of shear on realistic galaxies.
Study improves confidence measures in medical imaging pipelines by addressing bias.
problem Bias in metric-based imaging pipelines compromises the efficiency of prediction intervals.
method Formalized symmetric and asymmetric CP formulations, analyzed bias effects, and validated empirically.
result Symmetric intervals are inflated by bias, while asymmetric intervals remain unaffected.
LatentNN corrects neural network attenuation bias in astronomical data.
problem Neural networks underestimate extreme values due to measurement errors.
method Jointly optimizes network parameters and latent input values.
result LatentNN reduces attenuation bias across various signal-to-noise ratios.
Research shows bias in machine learning can be due to algorithmic flaws, not just data.
problem Underestimation bias in machine learning algorithms.
method Initial research to understand factors contributing to bias in classification algorithms.
result Regularization methods to address overfitting can also accentuate bias.
The paper explores the trade-off between bias and variance in high-dimensional models.
problem Understanding the unavoidable trade-off between bias and variance in high-dimensional statistical models.
method Proposes a general strategy to obtain lower bounds on the variance of estimators with a specified bias, and applies it to various statistical models.
result Shows the extent to which the bias-variance trade-off is unavoidable and quantifies the performance loss for methods that do not balance it.
Over the last decade there has been increasing concern about the biases embodied in traditional evaluation methods for Natural Language Processing/Learning, particularly methods borrowed from Information Retrieval. Without knowledge of the Bias and Prevalence of the contingency being tested, or equivalently the expecta…
New measure corrects news bias in NLP stock return forecasting.
problem Improving stock return and volatility forecasting accuracy.
method Hype-Adjusted Probability Measure, sentiment score equation.
result Significantly improved forecast accuracy for U.S. semiconductor tickers.
Improved GSPGS estimators reduce bias in noisy function measurements.
problem Reduced bias in noisy function measurements.
method Generalized Simultaneous Perturbation-based Gradient Search (GSPGS) with various estimators.
result Estimators requiring more function measurements have lower bias.
New fairness measures account for prediction uncertainties to detect bias.
problem Fairness of ML models is not well-defined and measures are limited.
method Introduce new fairness measures based on aleatoric and epistemic uncertainties.
result Uncertainty-based measures reveal bias not captured by existing measures.
Geometric framework analyzes bias in variational inference for posterior functionals.
problem Analyzing the bias of posterior functionals under variational approximations.
method Developed a geometric framework to evaluate the bias of posterior functionals using the variational tangent space.
result The leading-order bias of a posterior functional is determined by its component orthogonal to the variational tangent space.
New models reduce regional inequality by adjusting exchange range and asset distribution bias.
problem Reduction of regional inequality in economic systems.
method Proposed new asset exchange models with spatial exchange range and local support bias to adjust asset distribution and circulation rates.
result Achieved asset distribution from over-concentration to exponential and eventually normal, reducing Gini coefficient.
New method debiases feature importance in Random Forests.
problem MDI feature importance measure incorrectly assigns high importance to noisy features.
method Derive a new analytical expression for MDI and propose MDI-oob debiased feature importance measure.
result MDI-oob achieves state-of-the-art performance in feature selection from Random Forests.
We study sampling as optimization in the space of measures. We focus on gradient flow-based optimization with the Langevin dynamics as a case study. We investigate the source of the bias of the unadjusted Langevin algorithm (ULA) in discrete time, and consider how to remove or reduce the bias. We point out the difficul…
Many modern Artificial Intelligence (AI) systems make use of data embeddings, particularly in the domain of Natural Language Processing (NLP). These embeddings are learnt from data that has been gathered "from the wild" and have been found to contain unwanted biases. In this paper we make three contributions towards me…
To improve the efficiency of Monte Carlo estimation, practitioners are turning to biased Markov chain Monte Carlo procedures that trade off asymptotic exactness for computational speed. The reasoning is sound: a reduction in variance due to more rapid sampling can outweigh the bias introduced. However, the inexactness …
The paper introduces a new bias measure, infra-marginality, to quantify unfairness in group fairness.
problem The trade-off between group fairness and individual-level bias in decision-making.
method Proposes a new notion of η-infra-marginality, proves its independence from accuracy, and provides practical methods to measure and avoid it. result High accuracy does not lead to high infra-marginality, but maximizing group fairness often increases infra-marginality.
Measures of implied volatility roughness corrected for bias.
problem Bias in measuring implied volatility roughness.
method Examined implied volatility of short-term options and VIX index, corrected for bias.
result Corrected measures indicate appropriate proxies for underlying volatility.
New measure quantifies task difficulty for machine learning models.
problem Quantifying the inherent difficulty of machine learning tasks.
method Inductive bias complexity measure.
result Tasks requiring generalization over many dimensions are more difficult.
Paper introduces metrics to detect unintended bias in text classifiers.
problem Unintended bias in machine learning classifiers.
method Threshold-agnostic metrics considering various score distribution variations across groups.
result New metrics reveal subtle unintended bias in public models.
Study suggests using information flow measures to target interventions in neural networks.
problem Identifying neural network edges that can be pruned to reduce bias.
method Used M-information flow framework to measure and compare information flows about true labels and protected attributes, and evaluated pruning effects on bias reduction. result Pruning edges with larger information flows about protected attributes reduces bias at the output.
We discuss the origin of multiscaling in financial time-series and investigate how to best quantify it. Our methodology consists in separating the different sources of measured multifractality by analysing the multi/uni-scaling behaviour of synthetic time-series with known properties. We use the results from the synthe…
New method corrects bias in feature importance measures of GBM.
problem Bias in feature importance measures of GBM.
method Cross-validated unbiased base learners.
result Significant improvement in feature importance measures with minimal computational cost.
We use tools from geometric statistics to analyze the usual estimation procedure of a template shape. This applies to shapes from landmarks, curves, surfaces, images etc. We demonstrate the asymptotic bias of the template shape estimation using the stratified geometry of the shape space. We give a Taylor expansion of t…
Bias and flexibility trade off in learning algorithms.
problem Understanding the trade-off between bias and expressivity in learning algorithms.
method Measuring expressivity using entropy on algorithm outcome distributions, and deriving bounds on bias and expressivity.
result There is a necessary trade-off between bias and expressivity in learning algorithms.
A new confidence measure improves self-training in biased data.
problem Improving self-training in biased data.
method Proposes a new confidence measure, T-similarity, based on ensemble diversity of linear classifiers.
result Empirically shows the benefit of T-similarity for pseudo-labeling policies on various datasets.
New method corrects risk estimation bias, improving backtesting results.
problem Underestimation of risk by existing methods, especially in small samples.
method Proposes a new algorithm for bias correction using generalized Pareto distributions.
result The new algorithm leads to improved efficiency in estimating risk with heavy tails or heteroscedasticity.
New algorithm corrects risk estimation bias for heavy-tailed data.
problem Underestimation of risk in banking and insurance due to bias in estimation procedures.
method Proposes a new algorithm for bias correction and applies it to generalized Pareto distributions.
result The algorithm leads to more accurate risk estimation, especially in heavy-tailed data.
Entropy asymmetry affects regularization in ERM, leading to biased solutions.
problem Analyzing the impact of relative entropy asymmetry in ERM regularization.
method Examined Type-I and Type-II ERM-RER, comparing their solutions and properties.
result Type-II ERM-RER regularization introduces a strong bias against training data.
New method to evaluate visual explanations from neural networks.
problem Lack of consensus on measuring effectiveness of visual explanations.
method Proposed a new procedure for evaluating explanations using a range of sources.
result Demonstrated the benefit of combining different sources and the impact of bias parameters.
Improved sampling from complex distributions with reduced bias.
problem Reducing bias in high-dimensional sampling algorithms.
method Hierarchical entropy analysis to weaken assumptions and expand scope.
result Bias reduction in low-dimensional marginals scales with lower dimension, not full dimension.
New algorithms optimize spectral risk measures, improving interpolation between average and worst-case performance.
problem Optimizing spectral risk measures for learning systems.
method Developed stochastic algorithms to optimize spectral risk measures by characterizing their subdifferential and addressing challenges like biasedness of subgradient estimates and non-smoothness.
result Our approach outperforms out-of-the-box stochastic subgradient and dual averaging methods in optimizing spectral risk measures.
Look-Ahead-Bench evaluates financial LLMs for lookahead bias, revealing significant differences in model performance.
problem Measuring and mitigating lookahead bias in financial LLMs.
method Standardized benchmark evaluating model behavior in practical financial scenarios, analyzing performance decay across market regimes.
result Standard LLMs exhibit significant lookahead bias, while Pitinf models show improved generalization and reasoning abilities.
Measurements made by satellite remote sensing, Moderate Resolution Imaging Spectroradiometer (MODIS), and globally distributed Aerosol Robotic Network (AERONET) are compared. Comparison of the two datasets measurements for aerosol optical depth values show that there are biases between the two data products. In this pa…
Scalable algorithm for computing Wasserstein-2 barycenters without bias.
problem Computing Wasserstein-2 barycenters efficiently and accurately.
method Input convex neural networks and cycle-consistency regularization.
result Our approach avoids introducing bias and does not require minimax optimization.
New algorithms estimate Hessians using random directions for faster stochastic optimization.
problem Efficiently estimating Hessians for stochastic optimization.
method Generalized Hessian estimators using random directions and noisy function measurements.
result Asymptotically unbiased estimators with lower bias for more measurements.
New methods reduce bias in estimating calibration error.
problem Reducing bias in estimating calibration error.
method Synthesizing model outputs and using equal-mass bins.
result Two reliable calibration-error estimators found: debiased estimator and ECE_sweep.
Paper presents a new algorithm to approximate Wasserstein-2 barycenters without bias.
problem Approximating Wasserstein-2 barycenters of continuous measures.
method Generative model approach using arbitrary neural networks.
result The method does not introduce bias and is applicable to large-scale tasks.