System identifies health risks using semantic and machine learning.
problem Identifying risk factors associated with health conditions in subpopulations.
method Developed a combined semantic and machine learning system using a health risk ontology and knowledge graph.
result Dynamic discovery of risk factors and their subpopulations.
Improves flu prediction by blending environment and population info.
problem Challenges in using data from one environment in another due to feature variability and population subgroup differences.
method Population-aware hierarchical Bayesian domain adaptation framework with multiple invariant components.
result Model improves flu prediction in new environments with unlabelled data.
New methods resolve conflicting treatment effect estimates in health tech assessments.
problem Conflicting conclusions from different sponsors analyzing the same data.
method Arbitrated indirect treatment comparisons (ArMAIC) targeting a common target population.
result Estimates treatment effects in a common target population, resolving the MAIC paradox.
Diabetes is a major public health problem in the United States, affecting roughly 30 million people. Diabetes complications, along with the mental health comorbidities that often co-occur with them, are major drivers of high healthcare costs, poor outcomes, and reduced treatment adherence in diabetes. Here, we evaluate…
SaML guides ML models to avoid survey biases.
problem ML models trained on survey data often ignore survey design metadata.
method Nine-step guideline for incorporating survey design metadata in ML lifecycle.
result SaML provides valid population inference from survey data.
Although there are millions of transgender people in the world, a lack of information exists about their health issues. This issue has consequences for the medical field, which only has a nascent understanding of how to identify and meet this population's health-related needs. Social media sites like Twitter provide ne…
Method fuses low and high-resolution data for better health estimates.
problem Improving high-resolution health estimates from mixed data sources.
method Fusion of unbiased low-resolution and potentially biased high-resolution data, learning a distribution consistent with sampling bias.
result Significant reduction in bias in high-resolution estimates.
EKG-based models show better stability across patient populations than EHR-based models.
problem Model generalization issues in EHR and EKG-based predictive models.
method Two tests to measure model generalization, comparing EHR and EKG data.
result EKG-based models are more stable across different patient populations.
LHIEM model predicts health, income, and employment over years.
problem Lack of path dependency in health policy simulations.
method Discrete-time microsimulation with Markov chain modules.
result Validates health care financing proposal through detailed modeling.
A lack of information exists about the health issues of lesbian, gay, bisexual, transgender, and queer (LGBTQ) people who are often excluded from national demographic assessments, health studies, and clinical trials. As a result, medical experts and researchers lack a holistic understanding of the health disparities fa…
Bayesian networks learn sub-population differences from data.
problem Inference from a single network structure can be misleading when data populations are heterogeneous.
method A mixture of Bayesian networks where component probabilities depend on individual characteristics.
result Identifies both network structures and demographic predictors of sub-population membership.
Generative adversarial networks transform streetscape images to improve health and wellbeing.
problem Improving health and wellbeing outcomes through better streetscape design.
method Generative adversarial networks were used to translate Google Street View images, preserving structure while changing the style from bad health to good health areas.
result Translated images show that good health areas have more green space and compact urban design, while good social capital areas have more footpaths and less fencing.
Novel imputation method for EHRs with structured and sporadic missingness.
problem Missing data in integrated EHR datasets for clinical applications.
method Macomss, a novel imputation framework for structurally and heterogeneously missing data.
result Macomss outperforms existing methods in imputation and downstream prediction accuracy.
Proposes a new approach to improve disease prediction by considering who generates the data.
problem Improving disease prediction across different datasets considering who generates the data.
method Formulates domain adaptation as a multi-source hierarchical Bayesian framework.
result Improves prediction accuracy in target datasets with largely unlabelled data.
Novel approach for robust domain generalization in health studies.
problem Challenges in making statistical inferences about underrepresented minority groups.
method Structured tensor completion for multi-dimensional domain generalization in linear regression models.
result Established rigorous theoretical guarantees and demonstrated minimax optimality.
This paper presents an example of how demographical characteristics of patients influence their susceptibility to certain medical conditions. In this paper, we investigate the association of health conditions to age of patients in a heterogeneous population. We show that besides the symptoms a patients is having, the a…
Proposes a federated transfer learning method to improve precision medicine models for underrepresented populations.
problem Underrepresentation of minorities in precision medicine research leads to underperforming risk prediction models.
method Two-way federated transfer learning strategy integrating diverse populations and healthcare institutions.
result Improves risk prediction models for underrepresented populations, reducing performance gaps.
Bayesian meta-learning improves health prediction models across similar diseases.
problem Inter- and intra-task variability in healthcare predictions due to disease heterogeneity and patient differences.
method Bayesian meta-learning approach that models task similarity to mitigate negative transfer and improve generalizability.
result Significant generalizability improvements in stroke prediction tasks using electronic health record data.
AICov integrates population covariates for better COVID-19 forecasting.
problem Forecasting COVID-19 with broader social context.
method Integrative deep learning framework with LSTM and multiple data sources.
result Improved prediction of COVID-19 cases and deaths with population risk factors.
Algorithm optimizes lockdown policies balancing health and economy.
problem Balancing health and economic impacts of lockdowns during pandemics.
method Reinforcement learning to automatically compute lockdown policies.
result Algorithm learns optimal lockdown policies from disease and population data.
Estimates causal effect of managed care plans on NYC Medicaid spending.
problem Generalizing causal estimates to a target population not well-represented by randomized studies.
method Conditional cross-design synthesis estimators combining randomized and observational data.
result Estimates causal effect of managed care plans on health care spending.
Study shows racial bias in health data, which can be reduced with simple techniques.
problem Racial bias in health indicators measured by the Medical Expenditure Panel Survey (MEPS).
method Used publicly available and nationally representative MEPS data to show bias in predictive models for care management.
result Racial bias can be significantly reduced using simple mitigation techniques.
Machine learning models outperform traditional actuarial methods in predicting health insurance costs.
problem Improving accuracy in health insurance pricing to identify concession opportunities.
method Developed and evaluated two machine learning models at the patient and employer-group levels.
result Machine learning models outperformed traditional actuarial models by 20% in predicting costs.
Much recent research aims to identify evidence for Drug-Drug Interactions (DDI) and Adverse Drug reactions (ADR) from the biomedical scientific literature. In addition to this "Bibliome", the universe of social media provides a very promising source of large-scale data that can help identify DDI and ADR in ways that ha…
Language models improve clinical prediction models using EHR data.
problem Limited patient data for training clinical prediction models.
method Using patient representation schemes from natural language processing.
result 3.5% mean improvement in AUROC on five prediction tasks.
Deep autoencoders improve health tweet clustering.
problem Accurate clustering of health-related tweets.
method Deep convolutional autoencoders for learning compact tweet representations.
result Clustering performance significantly outperforms conventional methods.
Risk prediction is central to both clinical medicine and public health. While many machine learning models have been developed to predict mortality, they are rarely applied in the clinical literature, where classification tasks typically rely on logistic regression. One reason for this is that existing machine learning…
Deep neural network predicts health costs better than traditional models.
problem Accurate prediction of healthcare costs for optimal cost management.
method Developed a deep neural network to predict future health care costs from health insurance claims records.
result Deep neural network outperformed ridge regression and Morbi-RSA models in cost prediction.
Bayesian model identifies health disparities in disease progression.
problem Health disparities bias disease progression models.
method Interpretable Bayesian model accounting for three disparities.
result Model identifies and corrects for health disparities.
Bayesian predictive inference analyzes a dataset to make predictions about new observations. When a model does not match the data, predictive accuracy suffers. We develop population empirical Bayes (POP-EB), a hierarchical framework that explicitly models the empirical population distribution as part of Bayesian analys…
Study identifies diverse health states of opioid users to improve policy.
problem Diverse health states of opioid users lead to ineffective policy interventions.
method Probabilistic topic modeling of medical histories.
result Learned phenotypes predict future opioid use and prescription variability.
Study historical cholera epidemics and simulate long-term mortality impacts.
problem Long-term impacts of mortality shocks on longevity.
method Historical analysis of cholera epidemics and mathematical modeling of stochastic Individual-Based models.
result Simulated long-term mortality impacts following a mortality shock.
Supervised learning improves disease outbreak detection accuracy.
problem Early detection of infectious disease outbreaks to protect public health.
method Developed a supervised learning approach based on hidden Markov models for disease outbreak detection.
result Reduces false positive rate by up to 50% while maintaining sensitivity.
The wide implementation of electronic health record (EHR) systems facilitates the collection of large-scale health data from real clinical settings. Despite the significant increase in adoption of EHR systems, this data remains largely unexplored, but presents a rich data source for knowledge discovery from patient hea…
CRE discovers interpretable subgroups with heterogeneous treatment effects.
problem Identifying subgroups with notable treatment effect heterogeneity.
method Causal Rule Ensemble (CRE) using an ensemble-of-trees approach.
result CRE offers interpretable decision rules and high stability in subgroup discovery.
Stable health predictions need deconfounding test set features.
problem Stability of predictions in health machine learning is compromised by selection biases.
method Deconfounding the test set features improves prediction stability across different environments.
result Improved stability achieved by deconfounding test set features.
In this work we investigate intra-day patterns of activity on a population of 7,261 users of mobile health wearable devices and apps. We show that: (1) using intra-day step and sleep data recorded from passive trackers significantly improves classification performance on self-reported chronic conditions related to ment…
AI improves precision health through adaptive interventions.
problem Improving healthcare through personalized and dynamic treatments.
method Reinforcement learning (RL) for adaptive interventions in digital health.
result RL shows promise in dynamic healthcare problems.
Models for predicting the risk of cardiovascular events based on individual patient characteristics are important tools for managing patient care. Most current and commonly used risk prediction models have been built from carefully selected epidemiological cohorts. However, the homogeneity and limited size of such coho…
Predicts hospital patient discharge within 24 hours to optimize resource allocation.
problem Improving hospital resource management and patient care by prioritizing discharge.
method Used eight years of EHR data to train models predicting 24-hour discharge.
result Best models achieved AUROC of 0.85 and AUPRC of 0.53, well calibrated.
Early diagnosis is important for type 2 diabetes (T2D) to improve patient prognosis, prevent complications and reduce long-term treatment costs. We present a novel risk profiling approach based exclusively on health expenditure data that is available to Belgian mutual health insurers. We used expenditure data related t…
Study shows high-rise buildings in Dhaka affect mental health, especially lower-income residents.
problem Mental health risks due to high-rise buildings in Dhaka.
method Computer vision pipeline to analyze sky visibility, greenery, and colors in streets.
result Lower-income residents suffer more from lack of sky visibility and greenery in their environment.
Hidden Markov Models classify cough events with high accuracy.
problem Identifying coughing events in noisy environments for health monitoring.
method Used Hidden Markov Models (HMMs) for cough classification.
result Multivariate HMMs achieved 92% AUR in classifying cough events.
Researchers developed a generic model to account for structural variability in SHM.
problem Variability in natural frequency due to operational and environmental conditions limits SHM technologies.
method An overlapping mixture of Gaussian processes (OMGP) was used to generate a generic representation of normal condition.
result The OMGP model provided a generic representation (form) to characterise the normal condition of structures.
Imaging fluorescent disease biomarkers in tissues and skin is a non-invasive method to screen for health conditions. We report an automated process that combines intraoral fluorescent porphyrin biomarker imaging, clinical examinations and machine learning for correlation of systemic health conditions with periodontal d…
LEARNER improves low-rank matrix estimation using source population data.
problem Improving low-rank matrix estimation in target populations with diverse data sources.
method LEARNER uses similarity in latent spaces between source and target populations to enhance estimation.
result LEARNER often outperforms benchmark methods, especially with higher signal-to-noise ratios in the source population.
Multiple cause-of-death data provides a valuable source of information that can be used to enhance health standards by predicting health related trajectories in societies with large populations. These data are often available in large quantities across U.S. states and require Big Data techniques to uncover complex hidd…
Generative model predicts remaining life of damaged structures.
problem Prognosis of damage and remaining useful life of structures.
method Generative adversarial networks (GANs) in a population-based SHM framework.
result Algorithm provides confident predictions about structures' remaining useful life.