New method learns collective variables using autoencoders for molecular simulations.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Improves likelihood-free inference by using a new sampling approach to avoid biased data collection.
Many machine learning algorithms are trained and evaluated by splitting data from a single source into training and test sets. While such focus on in-distribution learning scenarios has led to interesting advancement, it has not been able to tell if models are relying on dataset biases as shortcuts for successful predi…
Post-ADC inference corrects bias in statistical inference after active data collection.
Study shows noisy data collection in ImageNet leads to biased model performance.
Machine learning models learn what we teach them to learn. Machine learning is at the heart of recommender systems. If a machine learning model is trained on biased data, the resulting recommender system may reflect the biases in its recommendations. Biases arise at different stages in a recommender system, from existi…
The paper tackles sampling biases by ensuring minority groups are adequately represented in training data.
Study optimizes data collection from biased, costly sources to minimize risk.
Macromolecular and biomolecular folding landscapes typically contain high free energy barriers that impede efficient sampling of configurational space by standard molecular dynamics simulation. Biased sampling can artificially drive the simulation along pre-specified collective variables (CVs), but success depends crit…
From scientific experiments to online A/B testing, the previously observed data often affects how future experiments are performed, which in turn affects which data will be collected. Such adaptivity introduces complex correlations between the data and the collection procedure. In this paper, we prove that when the dat…
Estimators computed from adaptively collected data do not behave like their non-adaptive brethren. Rather, the sequential dependence of the collection policy can lead to severe distributional biases that persist even in the infinite data limit. We develop a general method -- -decorrelation -- for transformi…
Balance corrects biased survey data for more accurate insights.
Citizen science projects are successful at gathering rich datasets for various applications. However, the data collected by citizen scientists are often biased --- in particular, aligned more with the citizens' preferences than with scientific objectives. We propose the Shift Compensation Network (SCN), an end-to-end l…
CausalSim corrects bias in trace-driven simulations for more accurate results.
Proposes CRA framework for certifying fair predictive models.
Study shows how online personalization can lead to unfair models due to biased user responses.
Estimates peeking effects in p-values to correct bias.
In binary classification, there are situations where negative (N) data are too diverse to be fully labeled and we often resort to positive-unlabeled (PU) learning in these scenarios. However, collecting a non-representative N set that contains only a small portion of all possible N data can often be much easier in prac…
Scientific and business practices are increasingly resulting in large collections of randomized experiments. Analyzed together, these collections can tell us things that individual experiments in the collection cannot. We study how to learn causal relationships between variables from the kinds of collections faced by m…
Systematic discriminatory biases present in our society influence the way data is collected and stored, the way variables are defined, and the way scientific findings are put into practice as policy. Automated decision procedures and learning algorithms applied to such data may serve to perpetuate existing injustice or…
New validation method prevents privacy breaches and biases in federated learning.
Accumulation of standardized data collections is opening up novel opportunities for holistic characterization of genome function. The limited scalability of current preprocessing techniques has, however, formed a bottleneck for full utilization of contemporary microarray collections. While short oligonucleotide arrays …
Paper addresses bias in search intent affecting click behavior.
New method detects and mitigates historical bias in data.
A trade-off between accuracy and fairness is almost taken as a given in the existing literature on fairness in machine learning. Yet, it is not preordained that accuracy should decrease with increased fairness. Novel to this work, we examine fair classification through the lens of mismatched hypothesis testing: trying …
Industrial recommender systems deal with extremely large action spaces -- many millions of items to recommend. Moreover, they need to serve billions of users, who are unique at any point in time, making a complex user state space. Luckily, huge quantities of logged implicit feedback (e.g., user clicks, dwell time) are …
This study examines how learning algorithms affect collective action in machine learning.
UBM transfers bias mitigation from upstream to downstream tasks efficiently.
Electronic health records (EHR) are rich heterogeneous collection of patient health information, whose broad adoption provides great opportunities for systematic health data mining. However, heterogeneous EHR data types and biased ascertainment impose computational challenges. Here, we present mixEHR, an unsupervised g…
In many countries information on expectations collected through consumer confidence surveys are used in macroeconomic policy formulation. Unfortunately, before doing so, the consistency of responses is often not taken into account, leading to biases creeping in and affecting the reliability of the indices hence created…
Study identifies and measures biases in legal case data.
The paper tackles bandit problems with biased offline data by using causal methods.
New online method for statistical inference with matrix context in decision-making.
Synthetic data mimics real-world demographics for fairness testing.
The paper proposes a method to align AI models using conformal risk control.
Gradient-based methods can be biased by distributional asymmetries in bivariate categorical data.
Data collection in economically constrained countries often necessitates using approximate and biased measurements due to the low-cost of the sensors used. This leads to potentially invalid predictions and poor policies or decision making. This is especially an issue if methods from resource-rich regions are applied wi…
The results of data mining endeavors are majorly driven by data quality. Throughout these deployments, serious show-stopper problems are still unresolved, such as: data collection ambiguities, data imbalance, hidden biases in data, the lack of domain information, and data incompleteness. This paper is based on the prem…
Experiment shows cognitive biases impact human-AI collaboration, highlighting the need for diverse evaluator samples.
Paper explores how knowledge distillation transfers inductive biases between models.
New method improves model robustness to biased data.
Proposes a new Hawkes process bandit model for disaster search and rescue.
Data that is gathered adaptively --- via bandit algorithms, for example --- exhibits bias. This is true both when gathering simple numeric valued data --- the empirical means kept track of by stochastic bandit algorithms are biased downwards --- and when gathering more complicated data --- running hypothesis tests on c…
The study examines machine learning classification algorithms and their generalizability using Framingham Heart Study data.
This paper tackles confounding biases in data augmentation.
New framework embeds physics in coarse-grained models without big data.
Study shows statistical biases can mislead transformer models, impairing their generalization.
Multiple fairness constraints have been proposed in the literature, motivated by a range of concerns about how demographic groups might be treated unfairly by machine learning classifiers. In this work we consider a different motivation; learning from biased training data. We posit several ways in which training data m…