Study identifies and measures biases in legal case data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper tackles bandit problems with biased offline data by using causal methods.
Gradient-based methods can be biased by distributional asymmetries in bivariate categorical data.
Paper explores how knowledge distillation transfers inductive biases between models.
Machine learning models learn what we teach them to learn. Machine learning is at the heart of recommender systems. If a machine learning model is trained on biased data, the resulting recommender system may reflect the biases in its recommendations. Biases arise at different stages in a recommender system, from existi…
New method improves model robustness to biased data.
The paper tackles sampling biases by ensuring minority groups are adequately represented in training data.
Many machine learning algorithms are trained and evaluated by splitting data from a single source into training and test sets. While such focus on in-distribution learning scenarios has led to interesting advancement, it has not been able to tell if models are relying on dataset biases as shortcuts for successful predi…
This paper tackles confounding biases in data augmentation.
Study shows statistical biases can mislead transformer models, impairing their generalization.
Multiple fairness constraints have been proposed in the literature, motivated by a range of concerns about how demographic groups might be treated unfairly by machine learning classifiers. In this work we consider a different motivation; learning from biased training data. We posit several ways in which training data m…
The study examines how social biases are reinforced in machine learning models used for credit scoring.
In this paper, we propose a new framework for mitigating biases in machine learning systems. The problem of the existing mitigation approaches is that they are model-oriented in the sense that they focus on tuning the training algorithms to produce fair results, while overlooking the fact that the training data can its…
In many applications, different populations are compared using data that are sampled in a biased manner. Under sampling biases, standard methods that estimate the difference between the population means yield unreliable inferences. Here we develop an inference method that is resilient to sampling biases and is able to …
The effectiveness of machine learning algorithms depends on the quality and amount of data and the operationalization and interpretation by the human analyst. In humanitarian response, data is often lacking or overburdening, thus ambiguous, and the time-scarce, volatile, insecure environments of humanitarian activities…
The paper analyzes time-dependent streaming data with biased gradient estimates and proposes improved stochastic optimization methods.
How do we learn from biased data? Historical datasets often reflect historical prejudices; sensitive or protected attributes may affect the observed treatments and outcomes. Classification algorithms tasked with predicting outcomes accurately from these datasets tend to replicate these biases. We advocate a causal mode…
New method corrects skewed confidence for PbN classification.
This study simulates biases in classifiers to assess fairness.
We evaluate the folk wisdom that algorithmic decision rules trained on data produced by biased human decision-makers necessarily reflect this bias. We consider a setting where training labels are only generated if a biased decision-maker takes a particular action, and so "biased" training data arise due to discriminato…
In this paper, we show that popular Generative Adversarial Networks (GANs) exacerbate biases along the axes of gender and skin tone when given a skewed distribution of face-shots. While practitioners celebrate synthetic data generation using GANs as an economical way to augment data for training data-hungry machine lea…
Study contextual online pricing with biased offline data, achieving optimal regret bounds.
Noise in imputed values corrects biases in machine learning models.
New method learns collective variables using autoencoders for molecular simulations.
Biased sampling and missing data complicates statistical problems ranging from causal inference to reinforcement learning. We often correct for biased sampling of summary statistics with matching methods and importance weighting. In this paper, we study nearest neighbor matching (NNM), which makes estimates of populati…
Paper proposes synthetic data generator to study and mitigate bias in machine learning.
Algorithm improves binary classification of biased grouped data.
Balance corrects biased survey data for more accurate insights.
Neural networks with random hidden nodes have gained increasing interest from researchers and practical applications. This is due to their unique features such as very fast training and universal approximation property. In these networks the weights and biases of hidden nodes determining the nonlinear feature mapping a…
Recommender systems are used in variety of domains affecting people's lives. This has raised concerns about possible biases and discrimination that such systems might exacerbate. There are two primary kinds of biases inherent in recommender systems: observation bias and bias stemming from imbalanced data. Observation b…
Study shows how online personalization can lead to unfair models due to biased user responses.
Kernel methods are popular in clustering due to their generality and discriminating power. However, we show that many kernel clustering criteria have density biases theoretically explaining some practically significant artifacts empirically observed in the past. For example, we provide conditions and formally prove the…
Improved speech enhancement with larger neural networks using novel embeddings and biases.
New method uses biased MD to create accurate MLIPs.
As machine learning black boxes are increasingly being deployed in domains such as healthcare and criminal justice, there is growing emphasis on building tools and techniques for explaining these black boxes in an interpretable manner. Such explanations are being leveraged by domain experts to diagnose systematic error…
Model proposes neural network for continuous time dynamics with inductive biases.
Regularized training of an autoencoder typically results in hidden unit biases that take on large negative values. We show that negative biases are a natural result of using a hidden layer whose responsibility is to both represent the input data and act as a selection mechanism that ensures sparsity of the representati…
Study evaluates if LLMs have company-specific biases in financial sentiment analysis.
Paper proposes using unlabeled data for fair decision-making.
We study the problem of ranking from crowdsourced pairwise comparisons. Answers to pairwise tasks are known to be affected by the position of items on the screen, however, previous models for aggregation of pairwise comparisons do not focus on modeling such kind of biases. We introduce a new aggregation model factorBT …
GCNs favor high-degree nodes, leading to biased performance; a new method mitigates this.
The use of synthetic data generated by Generative Adversarial Networks (GANs) has become quite a popular method to do data augmentation for many applications. While practitioners celebrate this as an economical way to get more synthetic data that can be used to train downstream classifiers, it is not clear that they re…
This work uncovers how model and data biases interact to cause unfairness in fraud detection.
Study data biases to predict algorithmic discrimination, developing a Data Bias Profile.
Chinchilla Approach 2 biases neural scaling law estimates, leading to unnecessary compute costs.
Improves fairness in machine learning by adding underrepresented group data.
Many modern Artificial Intelligence (AI) systems make use of data embeddings, particularly in the domain of Natural Language Processing (NLP). These embeddings are learnt from data that has been gathered "from the wild" and have been found to contain unwanted biases. In this paper we make three contributions towards me…
LLMs show biases in investment analysis, leading to unreliable recommendations.