Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

24487195 · Jun 202019922001200920182026
48 results for ethnicity sensitive

Automated author disambiguation using crowdsourced data and semi-supervised learning.

problem Grouping scientific publications by the same author, accounting for homonyms and synonyms.
method Exploits crowdsourced annotations for training an accurate classifier and clustering publications semi-supervisedly.
result Improves recall and tailors disambiguation to non-Western author names.

TaCo prevents non-linear classifiers from detecting sensitive attributes.

problem Ensuring fairness in NLP models by preventing sensitive attribute detection.
method Targeted Concept Erasure (TaCo) removes sensitive information from final latent representations, even against non-linear classifiers.
result TaCo outperforms state-of-the-art methods in reducing sensitive attribute prediction accuracy while preserving overall task performance.

Models predict race and ethnicity from names, improving accuracy over census data.

problem Inferring race and ethnicity from names, especially when first names are available.
method Modeling the relationship between characters in a name and race/ethnicity using Long Short-Term Memory.
result Long Short-Term Memory model achieves out-of-sample accuracy of 0.85.

The study improves colorectal cancer survivability prediction by considering ethnicity.

problem Improving colorectal cancer survivability prediction using machine learning.
method Machine learning techniques applied to SEER cancer incidence database, comparing different ethnicities.
result Models perform better on single-ethnicity populations and provide different feature importance rankings.

The study uses transfer learning to compare surgical outcomes across racial/ethnic subgroups.

problem Difficulty in comparing surgical outcomes due to racial/ethnic and geographic differences.
method Causal inference framework and transfer learning to incorporate data from multiple populations.
result Racial and ethnic differences in surgical outcomes are found, with non-Hispanic Black patients experiencing wide variability.

The paper proposes a method to improve fairness in classification without using sensitive features directly.

problem Balancing accuracy and fairness in automated decision-making systems.
method Combining Multitask Learning with fairness constraints to train group-specific classifiers.
result The method achieves substantial improvements in both accuracy and fairness on real datasets.

Research aims to ensure fair classification across explicit and implicit sensitive features.

problem Ensuring fairness in machine learning models when sensitive features are not explicitly provided.
method Defined explicit and implicit cohorts, used clustering of embeddings, modified loss function.
result Improved classification parity across explicit and implicit sensitive features.

This research quantifies cross-sectoral inequalities using latent class analysis.

problem Addressing multiple and intersecting forms of inequality in various sectors.
method Innovative latent class analysis approach to quantify discrepancies.
result Significant discrepancies found among minority ethnic groups and between them and non-minority groups.

Paper finds gender classification accuracy varies by skin type, not ethnicity.

problem Unequal performance of face classification services across skin types and genders.
method Stability experiments, image manipulation, and post-hoc explanation techniques.
result Lip, eye, and cheek structure differences, not skin type, cause gender classification discrepancies.

Convolutional embedded networks improve clustering and ethnicity prediction from genetic variants.

problem Identifying population groups and predicting geographic ethnicity from genetic variants.
method Proposed convolutional embedded clustering and autoencoder classifier for genetic variant data.
result Our approach outperforms state-of-the-art methods in accuracy and scalability.

Novel approach for robust domain generalization in health studies.

problem Challenges in making statistical inferences about underrepresented minority groups.
method Structured tensor completion for multi-dimensional domain generalization in linear regression models.
result Established rigorous theoretical guarantees and demonstrated minimax optimality.

The study examines how social biases are reinforced in machine learning models used for credit scoring.

problem Reinforcement of societal biases in machine learning algorithms for credit scoring.
method Analysis of machine learning models predicting gender or ethnicity based on loan applications data.
result Machine learning models can reflect and reinforce social biases present in the data.

New fair regression method improves fairness in chronic kidney disease classification.

problem Mitigating societal bias in health care for multiple groups.
method Penalized fair regression framework for multiple groups, with penalties for true positive rate disparity.
result Achieves fairness-accuracy frontier beyond existing methods in simulations and real-world data.

The paper addresses fairness in machine learning by adjusting input distributions.

problem Reducing disparate impact in machine learning models over different groups.
method The approach involves learning a counterfactual distribution to adjust input variables for disadvantaged groups.
result The method can reduce disparate impact without training a new model.

A reliable human skin detection method that is adaptable to different human skin colours and illu- mination conditions is essential for better human skin segmentation. Even though different human skin colour detection solutions have been successfully applied, they are prone to false skin detection and are not able to c…

2014-10-14abs ↗pdf ↗

Extends multivariate regression for tensor-variate data, identifying brain regions and facial characteristics.

problem Challenges in fitting regression models with multivariate responses and covariates.
method Low-rank tensor formats on regression coefficients and tensor-variate normal distribution for errors.
result Maximum likelihood estimators for tensor-on-tensor regression via block-relaxation algorithms.

Paper predicts demographics at finer geographic resolutions using geotagged tweets.

problem Limited traditional survey methods for demographics estimates at finer geographic resolutions.
method Adapting prior work to predict gender and race/ethnicity counts at the blockgroup-level.
result Achieves high correlations (0.671 for gender, 0.692 for race) compared to prior work.

Proposes a general framework for fairness-aware learning using f-divergences.

problem Ensuring fairness in classifier predictions without compromising accuracy.
method Introduces a general framework using f-divergences and provides a unified analysis of the upper bound of the estimation error.
result Guarantees low dependencies on unseen samples for any f-divergence.

A fair policy for hiring candidates from different groups is proposed in a linear contextual bandit problem.

problem Selecting candidates from different sensitive groups in a fair manner.
method A greedy policy that constructs a ridge regression estimate and computes relative rank using empirical cumulative distribution function.
result The greedy policy achieves fair pseudo-regret of order dT\sqrt{dT} after TT rounds, satisfying demographic parity.

Introduces MPR to measure and optimize representation across intersectional groups in retrieval.

problem Harmful stereotypes, cultural erasure, and social disparities in image search and retrieval.
method Develops MPR metric, practical estimation methods, theoretical guarantees, and optimization algorithms.
result Optimizing MPR yields more proportional representation across multiple intersectional groups, often with minimal retrieval accuracy compromise.

5D AI model detects bad loans without biased features, improving consumer protection.

problem Detecting bad loans without biased features and improving consumer protection.
method Machine learning, BiMOPT features, European Banking Authority principles, AI principles, historical and validation datasets.
result 5D correctly detected 1,461 bad loans out of 1,613 (Sensitivity = 0.91, Prevalence = 0.0253, Positive Predictive Value = 0.19).

Prevents sensitive data generation in diffusion models using labeled and unlabeled data.

problem Generating sensitive data in diffusion models using unlabeled data.
method Positive-Unlabeled Diffusion Models, approximating ELBO with labeled and unlabeled data.
result Prevents the generation of sensitive data without compromising image quality.

Model detects cyberbullying by analyzing participant-vocabulary consistency.

problem Identifying cyberbullying on social media platforms.
method Formulated an objective function based on participant-vocabulary consistency to detect cyberbullying.
result The model can detect new bullying vocabulary, victims, and bullies.

This work provides efficient algorithms for approximating ℓ_p sensitivities and related statistics.

problem Estimating the importance of datapoints in high-dimensional datasets.
method Efficient algorithms for computing α-approximation of ℓ_1 sensitivities and total sensitivity using importance sampling and sensitivity computations.
result Real-world datasets have significantly lower intrinsic effective dimensionality than theoretical predictions.

Unified framework for CVA sensitivities, hedging, and risk assessment.

problem Computing and managing Credit Value Adjustment (CVA) sensitivities and risks.
method Probabilistic machine learning and refined regression on simulated data, validated by Monte Carlo methods.
result Identification of optimal sensitivities for practical tasks like hedging and risk assessment.

Proposes a method to increase diversity without sacrificing meritocracy.

problem Systemic bias in datasets affecting diversity and meritocracy.
method Optimally flipping outcome labels and training classification models simultaneously.
result The price of diversity is low and sometimes negative, enhancing diversity without significantly affecting meritocracy.