New method groups genetic data into coherent topics for disease insights.
problem Analyzing large, multi-dimensional genetic data sets.
method Conditional Hierarchical Bayesian Tucker Decomposition for genetic data analysis.
result Our models are more coherent than baseline models.
Binacox detects multiple cut-points in high-dimensional Cox models for genetic cancer data.
problem Detecting multiple cut-points in high-dimensional Cox models with many continuous features.
method Combines one-hot encoding with binarsity penalty for feature selection and regularization.
result Significantly outperforms state-of-the-art survival models in terms of C-index and computational speed.
Gradient boosting enhances existing Mendelian models for genetic disease risk prediction.
problem Improving existing Mendelian models for genetic disease risk prediction.
method Combining gradient boosting with existing Mendelian models.
result Improved model outperforms both original and gradient boosting-only models.
Deep learning predicts synergistic drug combinations from multi-omics data.
problem Predicting effective drug combinations to overcome cancer drug resistance.
method AuDNNsynergy model integrating gene expression, copy number, genetic mutation data and drug properties.
result AuDNNsynergy model outperforms state-of-the-art approaches.
Study improves cancer classification using gene selection and projection methods.
problem Overfitting in high-dimensional microarray datasets for cancer classification.
method FSWOR technique, random projection, Kendall test, ensemble classifiers, LDA projection, Naïve Bayes.
result Achieved a test score of 96%, significantly outperforming existing methods.
Paper presents a machine learning method to predict cancer sub-clones.
problem Predicting the growth of sub-clones in cancer tumors.
method Data-driven machine learning approach to capture sub-clonal characteristics.
result Machine learning can predict which sub-clones are more likely to grow.
New method resolves time order in genetic mutation models.
problem Underspecification in modeling genetic mutation time evolution.
method Continuous-time Markov chains with additional independent items.
result Additional items help determine time order and resolve underspecification.
Method screens weakly associated predictors in high-dimensional data.
problem Identifying weakly associated predictors in ultrahigh-dimensional data.
method Covariance-insured screening methodology.
result Validates the method through simulations and real data studies.
Paper proposes inference method for high-dimensional censored quantile regression.
problem Identifying heterogeneous effects of high-dimensional genetic biomarkers on survival outcomes.
method Combines low-dimensional model estimates based on multi-sample splittings and variable selection.
result Proposed estimator is consistent and asymptotically follows a Gaussian process.
This review chronicles AI algorithms for cervical cancer screening.
problem Automated screening of cervical cancer using AI methods.
method Analysis of various machine learning algorithms and clustering techniques.
result Holistic review of computational methods over time.
Deep autoencoder predicts cancer types from DNA methylation patterns.
problem Differentiating cancer types based on DNA methylation states.
method Deep learning system with CpG island state classification and statistical methods.
result Overall Sensitivity of 88.24%, Specificity of 83.33%, Accuracy of 84.75%.
CN-SBM clusters cancer samples and regions based on copy number variants.
problem Clonal evolution in cancer monitored by noisy copy number variants.
method Probabilistic framework using bipartite categorical block model.
result Improved model fit and clinically relevant subtypes identified.
Neural networks improve cancer risk prediction from family history data.
problem Improving cancer risk prediction from family history data using machine learning.
method Developed and trained neural network models on large pedigrees to predict hereditary cancers.
result Neural networks can achieve nearly optimal prediction performance and outperform traditional models in misreported data.
Automatically extracts phenotypes from cancer clinical notes for genetic studies.
problem Lack of structured patient representations in EHRs.
method Clustering of medical terms and sentences in clinical notes.
result 341 significant associations between clinical features and somatic mutations.
Deep learning model explains breast cancer subtypes using logistic regression.
problem Clarifying the mechanisms of breast cancer subtypes for better treatment.
method Developed a PWL model that generates custom-made logistic regression for each patient.
result The PWL model reveals genes relevant to cell cycle-related pathways.
New algorithm predicts lung cancer progression and mortality.
problem Predicting semi-competing risk outcomes in lung cancer.
method Neural Expectation-Maximization algorithm for multi-state outcomes.
result Estimates non-parametric baseline hazards and risk functions.
Dr.S recommends cancer drugs based on genomic data.
problem Personalizing cancer treatments using genomic information.
method Machine learning to identify optimal drug-gene associations.
result Developed a Drug Recommendation System (Dr.S) for cancer cell lines.
The emergence and development of cancer is a consequence of the accumulation over time of genomic mutations involving a specific set of genes, which provides the cancer clones with a functional selective advantage. In this work, we model the order of accumulation of such mutations during the progression, which eventual…
EB-VAE combines tumor growth and dropout data for personalized treatment response modeling.
problem Challenges in integrating longitudinal tumor measurements, dropout information, and genetic covariates.
method Extended EB-VAE framework to jointly model longitudinal and time-to-event data, incorporating dropout hazard and genetic covariates.
result Hybrid decoder formulation yields consistent treatment-effect parameters and prior predictive performance comparable to neural decoder.
Proposes a two-stage method for estimating heterogeneous treatment effects using gradient boosting trees.
problem Estimating heterogeneous treatment effects in randomized clinical trials with high-dimensional predictive markers.
method Two-stage statistical learning procedure using gradient boosting trees (XGBoost) to estimate main effects and HTE.
result Improves efficiency in estimating heterogeneous treatment effects through nonparametric function estimation.
We introduce a tensor-based clustering method to extract sparse, low-dimensional structure from high-dimensional, multi-indexed datasets. This framework is designed to enable detection of clusters of data in the presence of structural requirements which we encode as algebraic constraints in a linear program. Our cluste…
New method combines ensembling and regularization for genomic disease prediction.
problem Genomic diseases require accurate prediction and biomarker identification.
method Integrates regularization with ensembling techniques for high-dimensional binary classification.
result Identifies critical biomarkers overlooked by competing methods.
Selective deconfounding improves ATE estimation with less data.
problem Estimating ATE with unobserved confounders using limited data.
method Combining confounded and deconfounded observational data for ATE estimation.
result Selective deconfounding can significantly reduce the amount of deconfounded data needed.
Generative model tailors anticancer drugs based on transcriptomic data.
problem Designing effective anticancer drugs considering genetic profiles.
method RL framework using pretrained VAEs to generate compounds conditioned on transcriptomic data.
result Generative model produces molecules with high predicted inhibitory effects.
Integrative analysis of disparate data blocks measured on a common set of experimental subjects is a major challenge in modern data analysis. This data structure naturally motivates the simultaneous exploration of the joint and individual variation within each data block resulting in new insights. For instance, there i…
In this thesis we present the novel semi-supervised network-based algorithm P-Net, which is able to rank and classify patients with respect to a specific phenotype or clinical outcome under study. The peculiar and innovative characteristic of this method is that it builds a network of samples/patients, where the nodes …
Bayesian inference for factorial hidden Markov models is challenging due to the exponentially sized latent variable space. Standard Monte Carlo samplers can have difficulties effectively exploring the posterior landscape and are often restricted to exploration around localised regions that depend on initialisation. We …
The linking genotype to phenotype is the fundamental aim of modern genetics. We focus on study of links between gene expression data and phenotype data through integrative analysis. We propose three approaches. 1) The inherent complexity of phenotypes makes high-throughput phenotype profiling a very difficult and labor…
Understanding functional organization of genetic information is a major challenge in modern biology. Following the initial publication of the human genome sequence in 2001, advances in high-throughput measurement technologies and efficient sharing of research material through community databases have opened up new view…
A key goal of computational personalized medicine is to systematically utilize genomic and other molecular features of samples to predict drug responses for a previously unseen sample. Such predictions are valuable for developing hypotheses for selecting therapies tailored for individual patients. This is especially va…
Estimates cost savings from early cancer diagnosis.
problem Improving early cancer diagnosis to reduce treatment costs.
method Combining published cancer treatment cost estimates by stage with incidence rates by stage at diagnosis, and extrapolating to other cancer sites.
result Estimates U.S. national annual treatment cost-savings from early cancer diagnosis in the trillions.
Proposes a two-stage method for testing variable interactions with FDR control.
problem Testing pairwise interactions in high-dimensional data with dependence.
method Two-stage testing procedure with FDR control using Cramér type moderate deviation technique.
result The proposed method controls FDR and has comparable or improved statistical power.
ENN method uses expectile regression for genetic data analysis of complex diseases.
problem Discover additional genetic variants contributing to complex diseases.
method Developed an expectile neural network (ENN) method integrating expectile regression and neural networks.
result ENN method outperforms existing expectile regression in discovering genetic variants predisposing to sub-populations.
Semi-supervised GAN creates synthetic genetic data for disease prediction.
problem Expensive and time-consuming to build large labeled genetic databases.
method Semi-supervised Genetic Generative Adversarial Network (gGAN).
result Model achieved satisfactory results with real genetic data.
Model predicts anti-cancer drug responses using gene and molecular data.
problem Expensive and time-consuming cancer drug discovery and tailoring.
method Uses variational autoencoders and multi-layer perceptrons to encode gene expression and drug data.
result High average R2 of 0.83 and 0.845 in predicting drug responses for breast and pan-cancer cell lines, respectively. Efficiently infers graph edges from genetic similarity data in landscape genetics.
problem Inferring unknown graph edges from genetic similarity data in a heterogeneous landscape.
method Developed an efficient first-order optimization method to solve the inverse landscape genetics problem.
result Our method provides fast and reliable convergence, significantly outperforming existing heuristics.
Study uses machine learning to predict heart failure in cancer patients.
problem Early detection of cancer patients at risk for cardiotoxicity.
method Examined four machine learning algorithms on 143,199 cancer patients.
result Gradient boosting model achieved best AUC score of 0.9077.
System accurately detects lung cancer from CT images.
problem Early and accurate detection of lung cancer.
method Developed algorithms using a dataset of CT images.
result Accuracy of 72.2% on test dataset.
Method distinguishes genetic correlations from causation in GWAS.
problem Identifying causal relationships among genetically correlated traits.
method Mixed fourth moments to quantify causal relationships.
result Identified 30 putative genetically causal relationships across 52 traits.
Deep learning models improve cancer detection and typing classification from gene expression data.
problem Challenges in establishing specificity for cancer diagnosis using gene expression data.
method Developed deep learning models using mRNA datasets for cancer detection and typing classification.
result Achieved 98% accuracy in cancer detection and 18 out of 32 cancer-typing classifications over 90% accuracy.
We present a novel method for extracting cancer signatures by applying statistical risk models (http://ssrn.com/abstract=2732453) from quantitative finance to cancer genome data. Using 1389 whole genome sequenced samples from 14 cancers, we identify an "overall" mode of somatic mutational noise. We give a prescription …
Study predicts 10-year survival rates for breast cancer patients.
problem Predicting long-term survival of breast cancer patients.
method Machine learning approaches to assess survival rates.
result Improved accuracy in predicting 10-year survival.
Study uses Apple ML to accurately detect and classify lung cancer.
problem Accurate diagnosis and sub-classification of non-small cell lung cancer.
method Evaluation of Apple Create ML module on histopathological images.
result 100% detection and successful subclassification of non-small cell lung cancer.
Cancer patients admitted to ICU had improved survival over 10 years.
problem To assess changes in survival of cancer patients admitted to ICU over 10 years.
method Retrospective analysis of MIMIC-III database, adjusted for confounders using logistic regression.
result Cancer patients had significantly lower 28-day and 1-year mortality rates over 10 years.
Paper identifies key CpG methylation sites for breast cancer.
problem Early detection and treatment of breast cancer.
method Used machine learning on TCGA dataset to classify cancer vs. non-cancer samples.
result Reduced model with 25 key CpG sites achieves over 94% accuracy.
Discovering causal genetic variants from large genetic association studies poses many difficult challenges. Assessing which genetic markers are involved in determining trait status is a computationally demanding task, especially in the presence of gene-gene interactions. A non-parametric Bayesian approach in the form o…
Data mining techniques predict breast cancer types with high accuracy.
problem Early detection of breast cancer to reduce mortality rates.
method Twelve classification algorithms applied to the Breast Cancer Wisconsin dataset.
result High accuracy in predicting malignant and benign breast cancer.
GPO uses genetic algorithms for deep policy optimization in reinforcement learning.
problem Catastrophic consequence of parameter crossovers in neural networks for deep reinforcement learning.
method GPO uses imitation learning for policy crossover in the state space and applies policy gradient methods for mutation.
result GPO achieves superior performance and comparable sample efficiency compared to state-of-the-art policy gradient methods.