Multitask learning improves phenotyping in EHR data, but its benefits vary by phenotype complexity.
problem Improving phenotyping accuracy in EHR data using multitask learning.
method Investigated multitask learning for phenotyping rare and common phenotypes in EHR data using neural nets and logistic regression.
result Multitask learning with neural nets consistently outperforms single-task neural nets for rare phenotypes but underperforms for common phenotypes.
Deep learning outperforms classical methods in patient phenotyping.
problem Classifying patients into medical conditions using clinical notes.
method Comparison of CNNs, n-gram models, and cTAKES-based approaches on 10 phenotyping tasks.
result CNNs achieve an average F1-score of 76, significantly outperforming other methods.
OMTL uses ontology to learn from imbalanced EHR data.
problem Imbalanced and small usable data for phenotypes in EHRs.
method Ontology-driven multi-task learning framework.
result Improved learning performance on phenotypes.
We develop a model to cluster time-series data with interval censoring, improving disease phenotyping.
problem Noise and interval censoring hinder clustering in disease phenotyping.
method Deep generative, continuous-time model that clusters time-series data while correcting for censorship.
result Our model corrects for interval censoring and recovers known clinical subtypes.
Model identifies key problems in HIV patients' records.
problem Complex and time-consuming task of identifying patient problems from electronic health records.
method Unsupervised phenotyping approach that jointly learns phenotypes from structured and unstructured data.
result Learned phenotypes and their relatedness are clinically valid and surpass existing methods.
SWoTTeD discovers hidden temporal patterns in EHR data.
problem Complex temporal patterns in EHR data.
method Sliding Window for Temporal Tensor Decomposition (SWoTTeD) with constraints and regularizations.
result SWoTTeD achieves at least as accurate reconstruction as state-of-the-art models and extracts meaningful temporal phenotypes.
Exponential growth in Electronic Healthcare Records (EHR) has resulted in new opportunities and urgent needs for discovery of meaningful data-driven representations and patterns of diseases in Computational Phenotyping research. Deep Learning models have shown superior performance for robust prediction in computational…
SS3M learns disease phenotypes from few labels.
problem Lack of supervised data for disease phenotyping.
method Semi-Supervised Mixed Membership Model (SS3M).
result SS3M learns interpretable disease phenotypes.
Given genetic variations and various phenotypical traits, such as Magnetic Resonance Imaging (MRI) features, we consider two important and related tasks in biomedical research: i)to select genetic and phenotypical markers for disease diagnosis and ii) to identify associations between genetic and phenotypical data. Thes…
A new method uses literature constraints to improve phenotyping from EHR data.
problem Improving phenotyping from electronic health records (EHR) data.
method Constrained tensor factorization with literature constraints.
result Improved phenotyping results for hypertensive patients.
VBphenoR uses variational Bayes for EHR-based patient phenotyping.
problem Phenotyping patients from EHR data for targeted treatments.
method Variational Bayes Gaussian Mixture Model (GMM) and logistic regression.
result Closed-form inference for efficient patient phenotype determination.
The linking genotype to phenotype is the fundamental aim of modern genetics. We focus on study of links between gene expression data and phenotype data through integrative analysis. We propose three approaches. 1) The inherent complexity of phenotypes makes high-throughput phenotype profiling a very difficult and labor…
The widely used genetic pleiotropic analysis of multiple phenotypes are often designed for examining the relationship between common variants and a few phenotypes. They are not suited for both high dimensional phenotypes and high dimensional genotype (next-generation sequencing) data. To overcome these limitations, we …
Paper develops federated tensor factorization for phenotyping without sharing patient data.
problem Deriving phenotypes across multiple hospitals without patient-level data sharing.
method Secure data harmonization and federated computation using ADMM.
result Method generates phenotypes similar to centralized training while respecting privacy.
Bayesian model enhances phenotype discovery in asthma EHRs.
problem Lack of interpretability in unsupervised learning phenotyping of EHR data.
method Operationalized a Bayesian latent class framework with clinical knowledge priors.
result Identified an asthma sub-phenotype with elevated eosinophil levels and allergy markers.
Method predicts brain regions for neuroimaging phenotypes.
problem Predicting phenotypes from brain networks without static community structure.
method Supervised community detection using block-structured regularization and ADMM optimization.
result The method identifies task-specific brain regions that improve phenotype prediction.
DPFact preserves privacy while collaboratively factorizing EHR tensors.
problem Privacy-preserving tensor factorization for EHRs.
method Differential privacy and collaborative learning.
result DPFact achieves higher accuracy and efficiency under privacy constraints.
Active learning with Gaussian processes improves crop phenotype data collection.
problem Scalability issue in high throughput phenotyping for crop improvement.
method Active learning algorithm with Gaussian Process model.
result Superior performance compared to current practices on sorghum data.
Unsupervised clustering reveals novel TBI phenotypes.
problem Inadequate categorization of traumatic brain injury (TBI) based on symptoms.
method Applied unsupervised learning with GLRM feature selection.
result Identified four novel TBI phenotypes with distinct feature profiles.
New model predicts traits from gene expression, accounting for heterogeneity and gene networks.
problem Predicting phenotypes from gene expression data, considering heterogeneity and gene networks.
method Developed a novel model that considers heterogeneity and gene regulatory networks.
result Model performs well on prediction and provides clusters and gene regulatory networks.
Deep learning quantifies butterfly phenotypes, validating evolutionary theory.
problem Capturing comprehensive phenotypic information of butterflies.
method Deep convolutional triplet network for phenotypic distance calculation.
result Euclidean phenotypic distances support classical mimicry theory.
auton-survival simplifies survival analysis for healthcare data.
problem Handling censored time-to-event data in healthcare.
method Open-source package for survival regression, adjustment, counterfactual estimation, phenotyping, and treatment effects.
result Demonstrates auton-survival's ability to support complex health and epidemiological questions.
This paper reviews methods for discovering patient subgroups from EHR data.
problem Discovering subgroups of patients and co-occurring medical conditions from EHR data.
method Low-rank data approximation methods like matrix and tensor decompositions.
result These methods provide transparent and interpretable insights into patient phenotypes.
New method handles correlated genes for better genomic prediction.
problem Technical issues with highly correlated genes in prediction models.
method Grouping algorithm that treats correlated genes as a group and uses their common patterns.
result Significantly outperforms standard models in prediction and feature selection.
New method learns parameter groups and structures in multi-response models.
problem Discovering unknown grouping structures in multi-response models.
method Proposes two convex regularization formulations and optimization approaches.
result Validated on simulations and real datasets, providing more accurate parameter estimation.
Develops methods for GWAS of high dimensional phenotypes using summary statistics.
problem Lack of methods to model pleiotropy in multi-phenotype GWAS.
method Bayesian inference model using summary statistics, fast computation, and biologically informed priors.
result Demonstrates utility in metabolite GWAS with interpretable pathway-level inference.
ODBAE detects complex phenotypes in biological data.
problem Challenges in identifying complex phenotypes from high-dimensional biological data.
method ODBAE (Outlier Detection using Balanced Autoencoders) identifies influential and high leverage points in latent relationships among multiple physiological parameters.
result ODBAE reveals novel metabolism-related genes and uncovers coordinated abnormalities across metabolic indicators.
Systematic review of electronic health record phenotyping approaches.
problem Detecting patient cohorts using electronic health records.
method Comprehensive literature review of preprocessing and modeling approaches.
result Natural language processing shows promise for electronic phenotyping.
TASTE combines static and temporal data for phenotyping EHRs.
problem Phenotyping EHRs with both static and temporal data.
method Jointly models static and temporal tensors using PARAFAC2 and non-negative matrix factorization, alternatingly solving sub-problems.
result TASTE outperforms existing methods in speed and clinical meaningfulness of phenotypes.
Study identifies three sub-phenotypes of AKI with different severity.
problem Tackles the heterogeneity of AKI to improve targeted interventions.
method Used a memory network-based deep learning approach on EHR data.
result Identified three distinct sub-phenotypes of AKI with varying severity.
Study uses LCA to identify ARDS sub-phenotypes improving predictive models.
problem Complex and heterogeneous nature of ARDS makes early recognition difficult.
method Applied latent class analysis to identify sub-groups, then built predictive models.
result Significantly improved prediction performance for two sub-phenotypes of ARDS.
Paper introduces tCNNS model for predicting drug cell line interactions.
problem Predicting phenotypic drug responses on cancer cell lines.
method tCNNS model using SMILES format for drugs and cancer cell lines.
result Achieves 0.84 for R2 and 0.92 for Rp. Scientists interact with deep learning models to avoid misleading results.
problem Deep neural networks can misinterpret data and achieve high performance by exploiting confounding factors.
method Introduce explanatory interactive learning (XIL) where scientists revise models based on explanations.
result XIL helps prevent misleading results and encourages model trust.
WEST uses EHRs and expert cases to improve rare disease phenotyping.
problem Limited labeled data for rare diseases.
method Weakly supervised transformer model trained on probabilistic silver-standard labels.
result WEST outperforms existing methods in phenotype classification and subphenotyping.
Study develops electronic phenotypes of ICU patient acuity.
problem Limited time for patient acuity assessments and imprecise clinical trajectory prediction.
method Developed electronic phenotypes using automated variable retrieval in electronic health records.
result Identified three phenotypes: persistently stable, persistently unstable, and transitioning from unstable to stable.
Binary PheNorm extends phenotype labeling for EHRs using binary silver labels.
problem Lack of gold-standard phenotype labels in EHR studies.
method Proposes Binary PheNorm, an extension that uses binary silver labels directly in phenotype scoring.
result Binary PheNorm achieved strong discrimination using binary labels alone and improved performance when combined with count labels.
The increased affordability of whole genome sequencing has motivated its use for phenotypic studies. We address the problem of learning interpretable models for discrete phenotypes from whole genomes. We propose a general approach that relies on the Set Covering Machine and a k-mer representation of the genomes. We sho…
Method uses network biology to construct gene expression models for cancer.
problem Building models for cancer phenotypes using gene expression data.
method Unsupervised construction of computational graphs based on protein-protein networks.
result The method outperforms other models in cancer phenotype analysis.
New method phenotypes sleep apnea patients using time series analysis.
problem Traditional diagnosis of sleep apnea is insufficient for capturing its multi-faceted outcomes.
method Fuzzy clustering in time and frequency domains, and persistent homology for topological analysis.
result Phenotyping patients improves understanding of sleep apnea.
Machine learning predicts plant phenotypes from soil microbiome data.
problem Predicting plant phenotypes from soil microbiome data.
method Two models (random forest and Bayesian neural network) were used to predict plant phenotypes from soil properties and microbial population density.
result Human decisions and normalization strategies significantly impact model performance.
Transfer learning improves clinical time series prediction with limited data.
problem Training deep RNNs for clinical tasks requires large labeled data and tuning.
method Transfer learning from pre-trained RNNs on multiple tasks to new tasks.
result Features from pre-trained RNNs improve model performance and robustness.
Study developed phenotypes for ICU patients' brain dysfunction states.
problem Underdiagnosis of acute brain dysfunction in ICU patients.
method Created algorithms to quantify and cluster brain dysfunction states.
result Developed three phenotypes of ICU patients' brain dysfunction states.
Unsupervised learning uncovers hidden patterns in health data.
problem Limited scalability and accuracy of supervised learning in identifying complex clinical patterns.
method Derives from Lasko et al. method, implemented in Apache Spark and Python, generalized for MIMIC-III lab data.
result Unsupervised learning finds patterns in EHRs that reveal subtypes of diseases.
GSU tests association between complex genotypes and phenotypes.
problem Testing association between complex genotypes and phenotypes.
method GSU is a similarity-based test using Laplacian kernel for complex objects.
result GSU identified three genes associated with Alzheimer's Disease.
As an increasing number of genome-wide association studies reveal the limitations of attempting to explain phenotypic heritability by single genetic loci, there is growing interest for associating complex phenotypes with sets of genetic loci. While several methods for multi-locus mapping have been proposed, it is often…
Tree-based regularization improves latent variable inference from related datasets.
problem Inferring latent variables from multiple related datasets in causal systems.
method Tree-Based Regularization (TBR) for sparse changes across environments.
result TBR identifies true latent variables up to simple transformations under sparse changes.
Deep learning predicts ICD codes with high accuracy for patient phenotyping.
problem Variability in ICD code assignment by coders.
method Deep learning model trained on demographics, lab results, and medications.
result Model predictions outperform coder assigned ICD codes in accuracy.
New algorithm improves plant breeding by clustering soybean genotypes more accurately and efficiently.
problem Low accuracy and high computational complexity in clustering plant genotypes.
method Spectral Clustering with Pivotal Sampling for phenotypic data.
result Our algorithm achieves substantially more accuracy than existing methods.