Bayesian model learns cancer subtypes from diverse NGS data.
problem Overdispersed NGS count data and limited samples for specific cancer types.
method Bayesian Multi-Domain Learning (BMDL) model using hierarchical negative binomial factorization.
result BMDL achieves reproducible cancer subtyping without negative transfer effects.
A new method interprets multiple kernel learning for cancer subtypes.
problem Hard to evaluate clustering results in multidimensional diseases like cancer.
method Combines feature clustering with multiple kernel dimensionality reduction.
result Identifies integrative patient subtypes and explains feature importance.
We present a nonparametric Bayesian method for disease subtype discovery in multi-dimensional cancer data. Our method can simultaneously analyse a wide range of data types, allowing for both agreement and disagreement between their underlying clustering structure. It includes feature selection and infers the most likel…
New method clusters disease subtypes from model explanations.
problem Discovering disease subtypes in noisy, high-dimensional data.
method Train classifier, extract explanations, cluster in explanation space.
result Cluster analysis on model explanations outperforms classical methods.
Deep learning model explains breast cancer subtypes using logistic regression.
problem Clarifying the mechanisms of breast cancer subtypes for better treatment.
method Developed a PWL model that generates custom-made logistic regression for each patient.
result The PWL model reveals genes relevant to cell cycle-related pathways.
In many applications, multivariate samples may harbor previously unrecognized heterogeneity at the level of conditional independence or network structure. For example, in cancer biology, disease subtypes may differ with respect to subtype-specific interplay between molecular components. Then, both subtype discovery and…
VICatMix clusters categorical biomedical data efficiently and selects relevant variables.
problem Efficient clustering of high-dimensional categorical biomedical data.
method Variational Bayesian finite mixture model with variational inference.
result Improves clustering accuracy and variable selection on noisy, high-dimensional data.
Personalized treatment of patients based on tissue-specific cancer subtypes has strongly increased the efficacy of the chosen therapies. Even though the amount of data measured for cancer patients has increased over the last years, most cancer subtypes are still diagnosed based on individual data sources (e.g. gene exp…
New framework distinguishes lung cancer subtypes using MALDI mass spectrometry.
problem Distinguishing between adenocarcinoma and squamous cell carcinoma subtypes in lung cancer.
method Supervised topological data analysis on MALDI mass spectrometry imaging data.
result The proposed framework successfully classifies lung cancer subtypes with competitive results.
KLIC combines multiple datasets for clustering, down-weighting noisy data.
problem Robustness of COCA in noisy or conflicting datasets.
method Multiple Kernel Learning for Integrative Clustering.
result KLIC down-weights noisy datasets, improving clustering accuracy.
We release a large ECG dataset for arrhythmia subtype discovery.
problem Discovering unknown subtypes of arrhythmia from continuous raw signals.
method Unsupervised representation learning task using semi-supervised evaluation.
result Qualitative evaluations show potential for representation learning in arrhythmia sub-type discovery.
Paper proposes clustering model for ICC based on histologic patterns.
problem Challenges in grading rare cancers like ICC due to small sample sizes and difficulty in extracting patterns.
method Unsupervised deep convolutional autoencoder clustering model trained on 246 ICC digitized slides.
result Three clusters significantly associated with recurrence-free survival in Cox-proportional hazard models.
Quantum machine learning classifies lung cancer subtypes.
problem Accurately classify Adenocarcinoma vs Squamous cell carcinoma patients.
method Amalgamation of classical and quantum machine learning models, feature selection, QCrush data representation, Quantum Boltzmann Machine.
result Successfully classified 104 non-small cell lung cancer patients.
Paper proposes scalable method for analyzing multi-omic data.
problem Integrating high-dimensional multi-omic data for cancer subtyping.
method Mixed graphical model approach using Birth-Death MCMC algorithm.
result Our method outperforms LASSO and standard BDMCMC in computational efficiency and model selection accuracy.
Bayesian model clusters diverse 'omics data for disease subtyping.
problem Clustering diverse 'omics datasets conflates multiple structures.
method Multi-view Bayesian mixture model with semi-supervised learning.
result Identifies distinct clusters of patients for stratified medicine.
We present a novel method for extracting cancer signatures by applying statistical risk models (http://ssrn.com/abstract=2732453) from quantitative finance to cancer genome data. Using 1389 whole genome sequenced samples from 14 cancers, we identify an "overall" mode of somatic mutational noise. We give a prescription …
Unsupervised method selects genes for tumor subtype discovery.
problem High-dimensional tumor gene expression data with noisy variables and heterogeneity.
method Autoencoders for latent space learning, Multiple Kernel Learning for feature selection, clustering.
result Lower redundancy and better clustering performance compared to benchmarks.
Study identifies biomarkers for lung cancer in female non-smokers.
problem Identifying prognostic biomarkers for stage III NSCLC in non-smoking females.
method Gene expression profiling and XGBoost machine learning algorithm.
result Top biomarkers validated in literature, with AUC score of 0.835.
Proposes JACA for joint analysis of multi-view data with class information.
problem Finding associations between multi-view data related to class memberships.
method Joint Association and Classification Analysis (JACA) framework.
result Improved misclassification rates and stronger associations compared to existing methods.
Despite great advances, molecular cancer pathology is often limited to the use of a small number of biomarkers rather than the whole transcriptome, partly due to computational challenges. Here, we introduce a novel architecture of Deep Neural Networks (DNNs) that is capable of simultaneous inference of various properti…
The medical research facilitates to acquire a diverse type of data from the same individual for particular cancer. Recent studies show that utilizing such diverse data results in more accurate predictions. The major challenge faced is how to utilize such diverse data sets in an effective way. In this paper, we introduc…
Many researches demonstrated that the DNA methylation, which occurs in the context of a CpG, has strong correlation with diseases, including cancer. There is a strong interest in analyzing the DNA methylation data to find how to distinguish different subtypes of the tumor. However, the conventional statistical methods …
OPAL optimizes labeling strategy for precise inference from uncertain models.
problem Inference from uncertain machine learning models is brittle.
method OPAL learns a smooth policy to adaptively label data points based on model uncertainty.
result OPAL yields estimators with the lowest variance and achieves nominal coverage in finite samples.
A hybrid method clusters and characterizes cancer data efficiently.
problem Challenges in clustering high-dimensional biomedical data.
method Gaussian mixture with generalized factor analyzers for efficient estimation.
result Our approach outperforms existing methods with faster convergence and higher accuracy.
Study examines XAI methods for ECG analysis to improve model transparency.
problem Lack of transparency in deep learning models for ECG analysis.
method Investigates post-hoc XAI methods for local and global perspectives, establishes sanity checks, and demonstrates knowledge discovery.
result Quantitative evidence supports expert rules for sensible attribution methods and demonstrates XAI's utility for knowledge discovery.
The task of clustering a set of objects based on multiple sources of data arises in several modern applications. We propose an integrative statistical model that permits a separate clustering of the objects for each data source. These separate clusterings adhere loosely to an overall consensus clustering, and hence the…
MEM learns set functions from permutation-invariant data.
problem Learning from sets of instances with labels only on sets, not instances.
method Memory-based Exchangeable Model (MEM) with self-attention mechanism.
result Achieved 84.84% accuracy on lung cancer classification.
Deep learning improves tumor type classification accuracy.
problem Classifying cancer types based on DNA mutations is challenging.
method Deep transfer learning and fine-tuning of gene expression data.
result Significantly improved tumor type classification accuracy (78.3%) using DNA point mutations.
ERICA assesses replicability of cluster analysis results.
problem Lack of quantitative scrutiny for clustering results.
method ERICA: a framework to assess replicability of cluster analysis.
result Clusters are found to be replicable in synthetic data but not in real-world datasets.
Bioinformatics tools have been developed to interpret gene expression data at the gene set level, and these gene set based analyses improve the biologists' capability to discover functional relevance of their experiment design. While elucidating gene set individually, inter gene sets association is rarely taken into co…
Study compares single vs ensemble feature selection for cancer diagnosis.
problem Identifying relevant variables for cancer diagnosis and prognosis.
method Comparison of single feature selection algorithms and ensemble of diverse algorithms.
result Ensemble approach did not improve predictive performance over individual algorithms.
UCSL combines clustering with supervised learning to discover interpretable subtypes.
problem Discovering interpretable subtypes in datasets relevant to supervised tasks.
method UCSL (Unsupervised Clustering driven by Supervised Learning) framework integrating clustering and supervised learning.
result UCSL achieves +1.9 points in balanced accuracy for psychiatric diseases clustering.
We introduce a tensor-based clustering method to extract sparse, low-dimensional structure from high-dimensional, multi-indexed datasets. This framework is designed to enable detection of clusters of data in the presence of structural requirements which we encode as algebraic constraints in a linear program. Our cluste…
Model predicts anti-cancer drug responses using gene and molecular data.
problem Expensive and time-consuming cancer drug discovery and tailoring.
method Uses variational autoencoders and multi-layer perceptrons to encode gene expression and drug data.
result High average R2 of 0.83 and 0.845 in predicting drug responses for breast and pan-cancer cell lines, respectively. We introduce a new discriminant analysis method (Empirical Discriminant Analysis or EDA) for binary classification in machine learning. Given a dataset of feature vectors, this method defines an empirical feature map transforming the training and test data into new data with components having Gaussian empirical distrib…
CN-SBM clusters cancer samples and regions based on copy number variants.
problem Clonal evolution in cancer monitored by noisy copy number variants.
method Probabilistic framework using bipartite categorical block model.
result Improved model fit and clinically relevant subtypes identified.
Clustering analysis is one of the most widely used statistical tools in many emerging areas such as microarray data analysis. For microarray and other high-dimensional data, the presence of many noise variables may mask underlying clustering structures. Hence removing noise variables via variable selection is necessary…
Method uses network biology to construct gene expression models for cancer.
problem Building models for cancer phenotypes using gene expression data.
method Unsupervised construction of computational graphs based on protein-protein networks.
result The method outperforms other models in cancer phenotype analysis.
Nonparametric IPSS selects features with false discovery control.
problem Feature selection in high-dimensional data with theoretical false discovery control.
method Integrated Path Stability Selection (IPSS) applied to nonparametric feature importance scores.
result IPSS accurately controls false discovery rate and detects more true positives than existing methods.
Global optimization approach for MAP clustering under Gaussian mixtures.
problem Maximum a-posteriori clustering problem under Gaussian mixture model.
method Mixed-integer nonlinear optimization (MINLP) transformed into mixed-integer quadratic program (MIQP).
result Explicit quantification of optimality gap, leading to globally optimal solutions.
Integrative analysis of disparate data blocks measured on a common set of experimental subjects is a major challenge in modern data analysis. This data structure naturally motivates the simultaneous exploration of the joint and individual variation within each data block resulting in new insights. For instance, there i…
For mass spectra acquired from cancer patients by MALDI or SELDI techniques, automated discrimination between cancer types or stages has often been implemented by machine learnings. These techniques typically generate "black-box" classifiers, which are difficult to interpret biologically. We develop new and efficient s…
Flow cytometry is often used to characterize the malignant cells in leukemia and lymphoma patients, traced to the level of the individual cell. Typically, flow cytometric data analysis is performed through a series of 2-dimensional projections onto the axes of the data set. Through the years, clinicians have determined…
Study identifies five AD subtypes using graph diffusion and similarity learning.
problem Identifying homogeneous AD subtypes to improve diagnosis and treatment.
method Unsupervised clustering with graph diffusion and similarity learning.
result Five distinct AD subtypes identified with significant differences in biomarkers and clinical features.
Unsupervised two-view learning, or detection of dependencies between two paired data sets, is typically done by some variant of canonical correlation analysis (CCA). CCA searches for a linear projection for each view, such that the correlations between the projections are maximized. The solution is invariant to any lin…
Paper introduces tCNNS model for predicting drug cell line interactions.
problem Predicting phenotypic drug responses on cancer cell lines.
method tCNNS model using SMILES format for drugs and cancer cell lines.
result Achieves 0.84 for R2 and 0.92 for Rp. Efficient algorithm for Bayesian networks reduces marginal probability distribution computation.
problem Exact computation of marginal probability distribution is NP-hard for categorical variables in Bayesian networks.
method Divide-and-conquer approach exploiting graphical properties of Bayesian networks.
result Novel algorithm outperforms state-of-the-art methods in classification and cancer subtype identification.
Study proposes a model to improve patient subtyping from EHR data.
problem Challenges in subtyping temporal EHR datasets.
method Self-supervised Mamba-based model for learning EHR representations.
result Model outperforms baseline models in EHR data subtyping.