Cluster analysis aims at separating patients into phenotypically heterogenous groups and defining therapeutically homogeneous patient subclasses. It is an important approach in data-driven disease classification and subtyping. Acute coronary syndrome (ACS) is a syndrome due to sudden decrease of coronary artery blood f…
Deep learning clusters patient time-series data for better prognosis.
problem Clustering time-series data for patient phenotyping and prognosis.
method Deep predictive clustering with novel loss functions for future outcome distribution.
result Model achieves superior clustering performance and identifies meaningful patient subgroups.
This study addresses the issue of predicting the glaucomatous visual field loss from patient disease datasets. Our goal is to accurately predict the progress of the disease in individual patients. As very few measurements are available for each patient, it is difficult to produce good predictors for individuals. A rece…
New method interprets deep embeddings for diabetes patient clustering.
problem Interpreting deep embeddings for disease progression.
method Patient clustering approach using deep embeddings.
result Clinically meaningful insights into diabetes progression patterns.
Driven by the multi-level structure of human intracranial electroencephalogram (iEEG) recordings of epileptic seizures, we introduce a new variant of a hierarchical Dirichlet Process---the multi-level clustering hierarchical Dirichlet Process (MLC-HDP)---that simultaneously clusters datasets on multiple levels. Our sei…
Bayesian Supervised Causal Clustering identifies patient subgroups for personalized decision-making.
problem Finding patient subgroups with similar characteristics for personalized decision-making.
method Bayesian Supervised Causal Clustering (BSCC) that identifies homogenous subgroups based on treatment effects.
result BSCC identifies subgroups with similar covariate profiles and treatment effects.
Patient journeys are compared to find clusters of similar disease trajectories.
problem Discovering shared health outcomes among patient journeys.
method Comparing longitudinal health data to identify clusters of similar patient trajectories.
result Clusters of patient journeys with similar health outcomes can be identified.
UMAP visualizes patient phenotypes from EHR data for emergency triage.
problem Interpreting high-dimensional EHR data for rapid patient triage.
method UMAP for non-linear dimensionality reduction, Gaussian mixture models for clustering.
result UMAP reveals clinically relevant patient phenotypes from EHR data.
The ability to accurately forecast and control inpatient census, and thereby workloads, is a critical and longstanding problem in hospital management. Majority of current literature focuses on optimal scheduling of inpatients, but largely ignores the process of accurate estimation of the trajectory of patients througho…
FONT clusters patients across health systems with privacy and efficiency.
problem Challenges in multi-site cluster analysis due to data-sharing restrictions.
method Federated One-shot Ensemble Clustering (FONT) algorithm that requires only a single round of communication and exchanges only fitted model parameters and class labels.
result FONT improves consistency of patient clusters across sites compared to locally fitted clusters.
Deep learning model creates patient representations for scalable EHR-based stratification.
problem Challenges in summarizing and representing patient data from EHRs prevent scalable stratification analysis.
method Unsupervised framework based on deep learning (ConvAE) using word embeddings, CNNs, and autoencoders.
result ConvAE significantly outperformed baselines in clustering diverse patient cohorts, identifying clinically relevant subtypes.
Paper uses NLP to cluster patient visits for diagnosis validation.
problem Validating if similar patients receive similar diagnoses.
method Representation of medical visits using word embeddings, clustering patients' visits.
result Stable and separated segments of visits positively validated against diagnoses.
Network Elastic Net identifies smoking-specific gene expression for lung cancer prognosis.
problem Identifying smoking-specific gene expression biomarkers in lung cancer prognosis.
method Introduces Network Elastic Net, a method that clusters and regresses on graphs based on smoking behavior.
result Shows efficacy of clusters in identifying cancer stages using gene expression and smoking behavior.
An unsupervised method clusters patient incident reports for content analysis.
problem Lack of methods to extract interpretable content from electronic healthcare records.
method Combines text-embedding with paragraph vectors and graph-theoretical multiscale community detection.
result Extracts high-intrinsic-consistency groups of patient incident reports.
Deep network clusters hospital patients' vital signs.
problem Sparse and irregularly collected vital sign data.
method Deep interpolation network for latent representation extraction.
result Extracted 7 distinct clusters from vital sign data.
In this paper we present a method for the unsupervised clustering of high-dimensional binary data, with a special focus on electronic healthcare records. We present a robust and efficient heuristic to face this problem using tensor decomposition. We present the reasons why this approach is preferable for tasks such as …
Due to the complexity of cancer, clustering algorithms have been used to disentangle the observed heterogeneity and identify cancer subtypes that can be treated specifically. While kernel based clustering approaches allow the use of more than one input matrix, which is an important factor when considering a multidimens…
Proposes a method to cluster fMRI data and estimate brain connectivity networks.
problem Clustering fMRI data to identify patient groups based on brain connectivity.
method Random covariance clustering model (RCCM) to cluster subjects and estimate individual and shared FC networks.
result RCCM outperforms other methods in clustering and FC network estimation, demonstrated through simulations and real data.
Paper predicts IVF pregnancy rates from basic patient info.
problem Predicting IVF pregnancy rates from patient characteristics.
method Clustering patients into groups, then SVM models for each group.
result Support vector machine models achieve best overall performance.
Study proposes a model to improve patient subtyping from EHR data.
problem Challenges in subtyping temporal EHR datasets.
method Self-supervised Mamba-based model for learning EHR representations.
result Model outperforms baseline models in EHR data subtyping.
Learning from electronic medical records (EMR) is challenging due to their relational nature and the uncertain dependence between a patient's past and future health status. Statistical relational learning is a natural fit for analyzing EMRs but is less adept at handling their inherent latent structure, such as connecti…
The paper introduces Precision Disease Networks (PDN) for predicting medical outcomes.
problem Predicting medical outcomes for patients with diseases.
method Building patient-specific disease networks, clustering, and data visualization.
result PDN improves prediction of patient outcomes compared to standard statistical analysis.
New method phenotypes sleep apnea patients using time series analysis.
problem Traditional diagnosis of sleep apnea is insufficient for capturing its multi-faceted outcomes.
method Fuzzy clustering in time and frequency domains, and persistent homology for topological analysis.
result Phenotyping patients improves understanding of sleep apnea.
Paper introduces methods to automatically generate SOAP notes from patient-physician conversations.
problem Burden of creating digital SOAP notes by physicians.
method Cluster2Sent algorithm for summarizing patient-physician conversations.
result Cluster2Sent algorithm outperforms existing methods by 8 ROUGE-1 points.
Traditional medicine typically applies one-size-fits-all treatment for the entire patient population whereas precision medicine develops tailored treatment schemes for different patient subgroups. The fact that some factors may be more significant for a specific patient subgroup motivates clinicians and medical researc…
Currently, approximately 30% of epileptic patients treated with antiepileptic drugs (AEDs) remain resistant to treatment (known as refractory patients). This project seeks to understand the underlying similarities in refractory patients vs. other epileptic patients, identify features contributing to drug resistance acr…
This paper presents an example of how demographical characteristics of patients influence their susceptibility to certain medical conditions. In this paper, we investigate the association of health conditions to age of patients in a heterogeneous population. We show that besides the symptoms a patients is having, the a…
New algorithms for clustering and synthetic data generation of heterogeneous tabular datasets.
problem Clustering and generating synthetic data from heterogeneous tabular datasets with hidden cluster structure.
method Developed MMM and MMMsynth algorithms for clustering and synthetic data generation.
result MMMsynth algorithm outperforms other literature tabular-data generators and approaches real data performance.
Study predicts heart failure patient survival using stacked ensemble ML.
problem Predicting survival of heart failure patients.
method Collect and analyze patient data, apply SMOTE, use K-Means, Fuzzy C-Means clustering, Random Forest, XGBoost, Decision Tree, and propose a stacked ensemble model.
result Supervised ML algorithms outperform unsupervised models, achieving high accuracy and F1 score.
StageNet improves health risk prediction by integrating disease stage information.
problem Improving health risk prediction for patients with chronic conditions.
method StageNet uses a stage-aware LSTM and stage-adaptive convolutional modules to extract and integrate disease stage information.
result StageNet achieves up to 12% higher AUPRC for risk prediction and over 58% higher Calinski-Harabasz score for patient subtyping compared to state-of-the-art models.
This study proposes a method to predict ICU infections from imbalanced data using clustering-based undersampling and ensemble classifiers.
problem Predicting healthcare-associated infections in ICU patients from imbalanced data.
method Clustering-based undersampling strategy combined with ensemble classifiers.
result The proposed method outperforms other resampling techniques in predicting ICU infections.
Diabetes is a major public health problem in the United States, affecting roughly 30 million people. Diabetes complications, along with the mental health comorbidities that often co-occur with them, are major drivers of high healthcare costs, poor outcomes, and reduced treatment adherence in diabetes. Here, we evaluate…
HIV RNA viral load (VL) is an important outcome variable in studies of HIV infected persons. There exists only a handful of methods which classify patients by viral load patterns. Most methods place limits on the use of viral load measurements, are often specific to a particular study design, and do not account for com…
Trans-GLMC tackles source heterogeneity in transfer learning for structured clusters.
problem Source heterogeneity makes it hard to use multiple related auxiliary sources effectively.
method Trans-GLMC constructs clusters of sources, then combines global fusion, within-cluster refinement, and target debiasing.
result Improves facility-specific prediction and identifies interpretable communities of hospitals with mutual transferability.
The paper clusters PK curves using ML, finding it useful for identifying similar patterns.
problem Improving drug development and patient outcomes through ML in pharmacogenomics.
method Unsupervised clustering of PK curves using various dissimilarity measures.
result Euclidean distance is most suitable for clustering PK curves, and clustering can validate pharmacogenomic results.
We present a novel probabilistic clustering model for objects that are represented via pairwise distances and observed at different time points. The proposed method utilizes the information given by adjacent time points to find the underlying cluster structure and obtain a smooth cluster evolution. This approach allows…
We augment linear Support Vector Machine (SVM) classifiers by adding three important features: (i) we introduce a regularization constraint to induce a sparse classifier; (ii) we devise a method that partitions the positive class into clusters and selects a sparse SVM classifier for each cluster; and (iii) we develop a…
New model captures patient-level EHR data efficiently.
problem Irregular EHR code timing and lack of temporal structure.
method Latent factor point process model with Fourier-Eigen embedding.
result Efficiently captures subgroup-specific temporal patterns.
Study develops electronic phenotypes of ICU patient acuity.
problem Limited time for patient acuity assessments and imprecise clinical trajectory prediction.
method Developed electronic phenotypes using automated variable retrieval in electronic health records.
result Identified three phenotypes: persistently stable, persistently unstable, and transitioning from unstable to stable.
Sepsis is a life-threatening disease and one of the major causes of death in hospitals. Imaging of microcirculatory dysfunction is a promising approach for automated diagnosis of sepsis. We report a machine learning classifier capable of distinguishing non-septic and septic images from dark field microcirculation video…
Method corrects bias in regression using simulated data and real-world gene expression data.
problem Bias in estimated effect parameters due to misclassification of class labels.
method Simulation and extrapolation method to correct bias.
result Corrected bias in estimated effect parameters.
Developing reliable workload predictive models can affect many aspects of clinical decision making procedure. The primary challenge in healthcare systems is handling the demand uncertainty over the time. This issue becomes more critical for the healthcare facilities that provide service for chronic disease treatment be…
The CHAMPION study clusters multi-dimensional accelerometer data to understand health links.
problem Clustering multi-dimensional data from pediatric longitudinal studies.
method Developed a finite mixture of multidimensional arrays model for clustering 4-dimensional accelerometer data.
result Demonstrated the feasibility and utility of clustering higher order data.
Paper tackles cancer mutation data challenges by creating useful low-dimensional representations.
problem Challenges in analyzing and using cancer mutation data for classification and clustering.
method Flatsomatic: variational autoencoders (VAEs) to create latent representations of somatic profiles.
result VAE embeddings perform better than PCA for clustering and equally well for classification.
MAGIC uncovers disease heterogeneity across brain scales.
problem Understanding distinct subtypes of brain diseases at different spatial scales.
method Multi-scale Heterogeneity Analysis and Clustering (MAGIC) using semi-supervised clustering.
result Two main subtypes of AD identified with distinct atrophy patterns.
latrend simplifies longitudinal clustering for numeric measurements.
problem Clustering of longitudinal data to identify common trends over time.
method Unified framework for applying various clustering methods.
result Facilitates comparison and rapid prototyping of new methods.
DPSOM combines self-organizing maps with deep learning for better data clustering.
problem Improving clustering performance in complex data.
method Integrates self-organizing maps with probabilistic clustering using a VAE.
result DPSOM outperforms current deep clustering methods in various applications.
Study developed phenotypes for ICU patients' brain dysfunction states.
problem Underdiagnosis of acute brain dysfunction in ICU patients.
method Created algorithms to quantify and cluster brain dysfunction states.
result Developed three phenotypes of ICU patients' brain dysfunction states.