Develops a new statistical framework for analyzing genetic pleiotropy in high-dimensional phenotypes.
problem Limited analysis of genetic pleiotropy for high-dimensional phenotypes and genotypes.
method Sparse structural equation models (SEMs) extended to sparse functional SEMs, incorporating both common and rare variants, and using functional data analysis and ADMM techniques.
result Higher power to detect true causal genetic pleiotropic structures compared to existing methods.
Develops methods for GWAS of high dimensional phenotypes using summary statistics.
problem Lack of methods to model pleiotropy in multi-phenotype GWAS.
method Bayesian inference model using summary statistics, fast computation, and biologically informed priors.
result Demonstrates utility in metabolite GWAS with interpretable pathway-level inference.
ODBAE detects complex phenotypes in biological data.
problem Challenges in identifying complex phenotypes from high-dimensional biological data.
method ODBAE (Outlier Detection using Balanced Autoencoders) identifies influential and high leverage points in latent relationships among multiple physiological parameters.
result ODBAE reveals novel metabolism-related genes and uncovers coordinated abnormalities across metabolic indicators.
The linking genotype to phenotype is the fundamental aim of modern genetics. We focus on study of links between gene expression data and phenotype data through integrative analysis. We propose three approaches. 1) The inherent complexity of phenotypes makes high-throughput phenotype profiling a very difficult and labor…
UMAP visualizes patient phenotypes from EHR data for emergency triage.
problem Interpreting high-dimensional EHR data for rapid patient triage.
method UMAP for non-linear dimensionality reduction, Gaussian mixture models for clustering.
result UMAP reveals clinically relevant patient phenotypes from EHR data.
This paper reviews methods for discovering patient subgroups from EHR data.
problem Discovering subgroups of patients and co-occurring medical conditions from EHR data.
method Low-rank data approximation methods like matrix and tensor decompositions.
result These methods provide transparent and interpretable insights into patient phenotypes.
Active learning with Gaussian processes improves crop phenotype data collection.
problem Scalability issue in high throughput phenotyping for crop improvement.
method Active learning algorithm with Gaussian Process model.
result Superior performance compared to current practices on sorghum data.
A new method uses literature constraints to improve phenotyping from EHR data.
problem Improving phenotyping from electronic health records (EHR) data.
method Constrained tensor factorization with literature constraints.
result Improved phenotyping results for hypertensive patients.
Bayesian model enhances phenotype discovery in asthma EHRs.
problem Lack of interpretability in unsupervised learning phenotyping of EHR data.
method Operationalized a Bayesian latent class framework with clinical knowledge priors.
result Identified an asthma sub-phenotype with elevated eosinophil levels and allergy markers.
Framework analyzes leaf vein architecture using deep learning and statistical methods.
problem Discards structural information in leaf venation studies.
method Integrates deep learning and statistical techniques to represent and analyze leaf vascular architecture.
result Identifies significant gene-environment interactions in leaf vascular architecture.
Binary PheNorm extends phenotype labeling for EHRs using binary silver labels.
problem Lack of gold-standard phenotype labels in EHR studies.
method Proposes Binary PheNorm, an extension that uses binary silver labels directly in phenotype scoring.
result Binary PheNorm achieved strong discrimination using binary labels alone and improved performance when combined with count labels.
Linear Mixed Models (LMMs) are important tools in statistical genetics. When used for feature selection, they allow to find a sparse set of genetic traits that best predict a continuous phenotype of interest, while simultaneously correcting for various confounding factors such as age, ethnicity and population structure…
GSU tests association between complex genotypes and phenotypes.
problem Testing association between complex genotypes and phenotypes.
method GSU is a similarity-based test using Laplacian kernel for complex objects.
result GSU identified three genes associated with Alzheimer's Disease.
Developed a model for multi-SNP, multi-trait association mapping.
problem Complex traits are high-dimensional, hindering simple association mapping.
method Nonparametric Bayesian reduced rank regression model.
result Improves statistical power to identify genetic associations.
New KNN test improves association analysis of high-dimensional sequencing data.
problem Challenges in using neural networks for high-dimensional sequencing data analysis.
method Kernel-based neural network (KNN) test for complex association analysis.
result KNN test outperforms SKAT in detecting non-linear and interaction effects.
Study uses LCA to identify ARDS sub-phenotypes improving predictive models.
problem Complex and heterogeneous nature of ARDS makes early recognition difficult.
method Applied latent class analysis to identify sub-groups, then built predictive models.
result Significantly improved prediction performance for two sub-phenotypes of ARDS.
Deep learning predicts ICD codes with high accuracy for patient phenotyping.
problem Variability in ICD code assignment by coders.
method Deep learning model trained on demographics, lab results, and medications.
result Model predictions outperform coder assigned ICD codes in accuracy.
WEST uses EHRs and expert cases to improve rare disease phenotyping.
problem Limited labeled data for rare diseases.
method Weakly supervised transformer model trained on probabilistic silver-standard labels.
result WEST outperforms existing methods in phenotype classification and subphenotyping.
SS3M learns disease phenotypes from few labels.
problem Lack of supervised data for disease phenotyping.
method Semi-Supervised Mixed Membership Model (SS3M).
result SS3M learns interpretable disease phenotypes.
Multitask learning improves phenotyping in EHR data, but its benefits vary by phenotype complexity.
problem Improving phenotyping accuracy in EHR data using multitask learning.
method Investigated multitask learning for phenotyping rare and common phenotypes in EHR data using neural nets and logistic regression.
result Multitask learning with neural nets consistently outperforms single-task neural nets for rare phenotypes but underperforms for common phenotypes.
VBphenoR uses variational Bayes for EHR-based patient phenotyping.
problem Phenotyping patients from EHR data for targeted treatments.
method Variational Bayes Gaussian Mixture Model (GMM) and logistic regression.
result Closed-form inference for efficient patient phenotype determination.
Proposes a model to decompose feature-level variation in high-dimensional data.
problem Interpreting complex high-dimensional data for understanding feature-level variability.
method Covariate Gaussian Process Latent Variable Model (c-GPLVM) for structured kernel decomposition.
result Extracts low-dimensional structures from high-dimensional data sets while explaining feature-level variability.
Deep learning outperforms classical methods in patient phenotyping.
problem Classifying patients into medical conditions using clinical notes.
method Comparison of CNNs, n-gram models, and cTAKES-based approaches on 10 phenotyping tasks.
result CNNs achieve an average F1-score of 76, significantly outperforming other methods.
Paper develops federated tensor factorization for phenotyping without sharing patient data.
problem Deriving phenotypes across multiple hospitals without patient-level data sharing.
method Secure data harmonization and federated computation using ADMM.
result Method generates phenotypes similar to centralized training while respecting privacy.
Model identifies key problems in HIV patients' records.
problem Complex and time-consuming task of identifying patient problems from electronic health records.
method Unsupervised phenotyping approach that jointly learns phenotypes from structured and unstructured data.
result Learned phenotypes and their relatedness are clinically valid and surpass existing methods.
New algorithm improves plant breeding by clustering soybean genotypes more accurately and efficiently.
problem Low accuracy and high computational complexity in clustering plant genotypes.
method Spectral Clustering with Pivotal Sampling for phenotypic data.
result Our algorithm achieves substantially more accuracy than existing methods.
Unsupervised clustering reveals novel TBI phenotypes.
problem Inadequate categorization of traumatic brain injury (TBI) based on symptoms.
method Applied unsupervised learning with GLRM feature selection.
result Identified four novel TBI phenotypes with distinct feature profiles.
Modern biotechnologies often result in high-dimensional data sets with much more variables than observations (n ≪ p). These data sets pose new challenges to statistical analysis: Variable selection becomes one of the most important tasks in this setting. We assess the recently proposed flexible framework for variab…
Deep learning quantifies butterfly phenotypes, validating evolutionary theory.
problem Capturing comprehensive phenotypic information of butterflies.
method Deep convolutional triplet network for phenotypic distance calculation.
result Euclidean phenotypic distances support classical mimicry theory.
New method selects key features from millions of biological data points.
problem Scalability issue in feature selection for ultra-high dimensional biological data.
method Scaled up HSIC Lasso to handle millions of features.
result Achieves high accuracy with only 20 out of one million features.
The paper tackles high-dimensional mixed linear regression with unknown parameters and proposes methods for estimation, confidence intervals, and hypothesis testing.
problem High-dimensional mixed linear regression with unknown parameters and covariance structure.
method Iterative high-dimensional EM algorithm for estimating regression vectors, debiased estimators for individual coordinates, and large-scale multiple testing procedure.
result Asymptotic normality of debiased estimators and FDR control for hypothesis testing.
AI improves healthcare diagnostics and predictions.
problem Data heterogeneity and model limitations in AI for health.
method Review of AI applications in health informatics.
result AI enhances disease diagnosis and prediction.
Systematic review of electronic health record phenotyping approaches.
problem Detecting patient cohorts using electronic health records.
method Comprehensive literature review of preprocessing and modeling approaches.
result Natural language processing shows promise for electronic phenotyping.
Paper introduces a method to distill interpretable phenotypes from deep learning models for healthcare.
problem Lack of interpretability in deep learning models for clinical decision-making.
method Interpretable Mimic Learning using Gradient Boosting Trees.
result Obtains similar or better performance than deep learning models while providing interpretable phenotypes.
TASTE combines static and temporal data for phenotyping EHRs.
problem Phenotyping EHRs with both static and temporal data.
method Jointly models static and temporal tensors using PARAFAC2 and non-negative matrix factorization, alternatingly solving sub-problems.
result TASTE outperforms existing methods in speed and clinical meaningfulness of phenotypes.
Study identifies three sub-phenotypes of AKI with different severity.
problem Tackles the heterogeneity of AKI to improve targeted interventions.
method Used a memory network-based deep learning approach on EHR data.
result Identified three distinct sub-phenotypes of AKI with varying severity.
Unified model learns joint and individual features from brain imaging data.
problem Integrating structural and functional connectivity data for behavioral phenotypes.
method Cross-Modal Joint-Individual Variational Network (CM-JIVNet) with multi-head attention fusion.
result CM-JIVNet outperforms in cross-modal reconstruction and behavioral trait prediction.
SWoTTeD discovers hidden temporal patterns in EHR data.
problem Complex temporal patterns in EHR data.
method Sliding Window for Temporal Tensor Decomposition (SWoTTeD) with constraints and regularizations.
result SWoTTeD achieves at least as accurate reconstruction as state-of-the-art models and extracts meaningful temporal phenotypes.
Paper develops a method for causal representation learning from irregular tensors.
problem Complex patterns in high-dimensional, irregular tensor data.
method Novel causal formulation and CaRTeD framework integrating temporal causal representation learning with irregular tensor decomposition.
result Framework provides theoretical guarantees and outperforms state-of-the-art techniques.
Paper introduces tCNNS model for predicting drug cell line interactions.
problem Predicting phenotypic drug responses on cancer cell lines.
method tCNNS model using SMILES format for drugs and cancer cell lines.
result Achieves 0.84 for R2 and 0.92 for Rp. OMTL uses ontology to learn from imbalanced EHR data.
problem Imbalanced and small usable data for phenotypes in EHRs.
method Ontology-driven multi-task learning framework.
result Improved learning performance on phenotypes.
Scientists interact with deep learning models to avoid misleading results.
problem Deep neural networks can misinterpret data and achieve high performance by exploiting confounding factors.
method Introduce explanatory interactive learning (XIL) where scientists revise models based on explanations.
result XIL helps prevent misleading results and encourages model trust.
We develop a model to cluster time-series data with interval censoring, improving disease phenotyping.
problem Noise and interval censoring hinder clustering in disease phenotyping.
method Deep generative, continuous-time model that clusters time-series data while correcting for censorship.
result Our model corrects for interval censoring and recovers known clinical subtypes.
Study develops electronic phenotypes of ICU patient acuity.
problem Limited time for patient acuity assessments and imprecise clinical trajectory prediction.
method Developed electronic phenotypes using automated variable retrieval in electronic health records.
result Identified three phenotypes: persistently stable, persistently unstable, and transitioning from unstable to stable.
Fast and cheaper next generation sequencing technologies will generate unprecedentedly massive and highly-dimensional genomic and epigenomic variation data. In the near future, a routine part of medical record will include the sequenced genomes. A fundamental question is how to efficiently extract genomic and epigenomi…
ENN method uses expectile regression for genetic data analysis of complex diseases.
problem Discover additional genetic variants contributing to complex diseases.
method Developed an expectile neural network (ENN) method integrating expectile regression and neural networks.
result ENN method outperforms existing expectile regression in discovering genetic variants predisposing to sub-populations.
The increased affordability of whole genome sequencing has motivated its use for phenotypic studies. We address the problem of learning interpretable models for discrete phenotypes from whole genomes. We propose a general approach that relies on the Set Covering Machine and a k-mer representation of the genomes. We sho…
New method handles correlated genes for better genomic prediction.
problem Technical issues with highly correlated genes in prediction models.
method Grouping algorithm that treats correlated genes as a group and uses their common patterns.
result Significantly outperforms standard models in prediction and feature selection.