We present a technique to characterize differentially expressed genes in terms of their position in a high-dimensional co-expression network. The set-up of Gaussian graphical models is used to construct representations of the co-expression network in such a way that redundancy and the propagation of spurious informatio…
A new method infers neuronal cell types and their gene expression profiles from brain imaging data.
problem Lack of spatial information in single-cell RNA sequencing data.
method Spatial point process mixture model applied to in situ hybridization images.
result Inferred cell types and gene expression profiles validated with single-cell RNA sequencing data.
Machine learning improves antimicrobial resistance detection.
problem Insufficient representation of antimicrobial resistance data in the Data-for-Good community.
method Standard ensemble machine learning techniques applied to antimicrobial resistance datasets.
result Classification accuracies of mid-90% to low-80% for AMR detection.
Motivation: Cell-biological processes are regulated through a complex network of interactions between genes and their products. The processes, their activating conditions, and the associated transcriptional responses are often unknown. Organism-wide modeling of network activation can reveal unique and shared mechanisms…
With different genomes available, unsupervised learning algorithms are essential in learning genome-wide biological insights. Especially, the functional characterization of different genomes is essential for us to understand lives. In this book chapter, we review the state-of-the-art unsupervised learning algorithms fo…
BioBO optimizes gene perturbation design using Bayesian optimization with biological priors.
problem Efficient design of genomic perturbation experiments in drug discovery.
method Integrates Bayesian optimization with multimodal gene embeddings and enrichment analysis.
result Improves labeling efficiency by 25-40% and identifies top-performing perturbations more effectively.
Developed a multiway classification method for sparse data.
problem Classification of multiway arrays with sparsity.
method Extended Distance Weighted Discrimination (DWD) to multiway context, accounting for sparsity.
result Improves classification accuracy in multiway structured data.
New method improves classification of multi-way data.
problem Classifying subjects based on multi-dimensional data.
method Factorizes coefficients into specific weights for each dimension, extending SVM and DWD.
result Improves performance and simplifies interpretation over naive methods.
New datasets support supervised learning for fungal BGC discovery.
problem Lack of labeled data for fungal BGCs.
method Developed new publicly available datasets for supervised learning.
result Supervised learning outperforms data-driven methods in fungal BGC prediction.
A new method uses gene interaction networks to predict gene functions.
problem Predicting gene functions from gene interactions.
method Context graph kernel approach in a machine learning framework.
result The proposed method outperforms linkage-assumption-based methods.
Novel model predicts anticancer compound sensitivity with high accuracy and interpretability.
problem Predicting anticancer compound sensitivity with high accuracy and interpretability.
method Multimodal attention-based convolutional encoder using SMILES, gene expression profiles, and protein-protein interaction networks.
result The model significantly outperforms baseline models and demonstrates high interpretability.
VEGN uses graph neural networks to predict disease-causing mutations from genetic variants.
problem Identifying disease-causing mutations from millions of genetic variants.
method VEGN employs a graph neural network on a heterogeneous graph of genes and variants, learning gene-gene interactions.
result VEGN outperforms existing state-of-the-art models in variant effect prediction.
GSAE autoencoder models gene sets for better cancer subtype and prognosis analysis.
problem Inter-gene set associations not considered in gene set-based analyses.
method Gene superset autoencoder model incorporating prior gene sets.
result Gene supersets retain biological features and are reproducible for cancer subtype and prognosis.
Crowdsourced gene set queries reveal protein-protein interactions and gene-gene associations.
problem Lack of integrative analysis of diverse gene set queries.
method Harnessed thousands of user-submitted gene sets to construct a global gene-gene association network.
result The constructed network recapitulates known protein-protein interactions and gene-gene functional associations.
A new method for identifying significant gene subsets improves disease prediction.
problem Identifying significant subsets of genes for disease prediction.
method Kernel gene shaving using influence function of kernel CCA.
result The proposed method outperformed three popular gene selection methods.
EpiRL learns to detect gene-gene interactions.
problem Computational challenges in epistasis detection.
method Modeling epistasis as a Markov Decision Process and using reinforcement learning.
result EpiRL discovers highly interacted genes.
Bayesian model learns cell types and gene networks from two data views.
problem Estimating cell types and their regulatory networks from single-cell gene expression and epigenetic data.
method Symphony Bayesian hierarchical multi-view mixture model with Variational EM inference.
result Symphony outperforms other methods in learning cell types and regulatory networks.
New method expands seed genes to functionally related clusters.
problem Discovering functionally related genes lacking GO terms.
method Semi-supervised learning with positive and unlabeled examples.
result LPU approaches significantly outperform existing methods.
Kernel method detects higher order interactions in multi-view data for schizophrenia.
problem Detecting higher order interactions in multi-view biological data.
method Kernel method on reproducing kernel Hilbert space (RKHS) with mixed-effects linear model.
result Identified 13 triplets with significant correlations to hippocampal volume in schizophrenia.
A new method for joint eQTL mapping and gene network estimation.
problem Discovering SNP-gene relationships and gene-gene relationships in gene expression regulation.
method L1-2 regularized multi-task graphical lasso (L1-2 GLasso).
result Competitive performance on capturing true sparse structures of eQTL mapping and gene network.
New method handles correlated genes for better genomic prediction.
problem Technical issues with highly correlated genes in prediction models.
method Grouping algorithm that treats correlated genes as a group and uses their common patterns.
result Significantly outperforms standard models in prediction and feature selection.
VGAE learns gene-disease associations from networks, predicting disease-genes.
problem Predicting gene-disease associations from disease-gene networks.
method Introducing VGAE, a variational graph auto-encoder for disease-gene prediction.
result VGAE and C-VGAE outperform baseline methods in disease-gene prediction.
Elucidating the genetic basis of human diseases is a central goal of genetics and molecular biology. While traditional linkage analysis and modern high-throughput techniques often provide long lists of tens or hundreds of disease gene candidates, the identification of disease genes among the candidates remains time-con…
Proposes a method to identify relevant genes in autism-related diseases using auxiliary information.
problem Identifying relevant genes in autism-related diseases from diverse data sources.
method Uses logistic regression to filter irrelevant genes and clusters relevant genes into cohesive groups using adjacency matrix.
result Superior performance and robustness in finite samples observed in simulation studies.
Robust method detects gene-gene interactions in imaging genetics data.
problem Detecting nonlinear gene-gene interactions in imaging genetics data.
method Robust Kernel Canonical Correlation Analysis (RKCCA) with influence function variance estimation.
result The proposed robust RKCCA method outperforms state-of-the-art methods in detecting gene-gene interactions.
We present the extention and application of a new unsupervised statistical learning technique--the Partition Decoupling Method--to gene expression data. Because it has the ability to reveal non-linear and non-convex geometries present in the data, the PDM is an improvement over typical gene expression analysis algorith…
Collaborative filtering predicts drug responses from gene expression data.
problem Predicting drug responses from large gene expression datasets with limited samples.
method Low-rank matrix factorization and latent linear regression.
result The proposed method outperforms state-of-the-art methods in predicting drug-gene associations.
A novel method selects genes for high-dimensional gene expression data with class imbalance.
problem Class imbalance in gene expression datasets.
method Synthetic data balancing, greedy search, weighted robust score.
result The proposed method outperforms existing feature selection procedures.
The problem of multilabel classification when the labels are related through a hierarchical categorization scheme occurs in many application domains such as computational biology. For example, this problem arises naturally when trying to automatically assign gene function using a controlled vocabularies like Gene Ontol…
Identifying latent structure in large data matrices is essential for exploring biological processes. Here, we consider recovering gene co-expression networks from gene expression data, where each network encodes relationships between genes that are locally co-regulated by shared biological mechanisms. To do this, we de…
The method integrates survival constraints into NMF for identifying survival-associated gene clusters.
problem Understanding and interpreting high-dimensional biological data for disease markers.
method Cox proportional hazards regression integrated with NMF via proportional hazards non-negative matrix factorization.
result The method can uncover survival-associated gene clusters in cancer gene expression data.
New methods detect continuous variation in single-cell data.
problem Continuous variation within and between cell types not detected by discrete analyses.
method Three topologically motivated mathematical methods for unsupervised feature selection.
result Detect additional biologically meaningful genes with coherent expression patterns.
Microarray cancer gene expression data comprise of very high dimensions. Reducing the dimensions helps in improving the overall analysis and classification performance. We propose two hybrid techniques, Biogeography - based Optimization - Random Forests (BBO - RF) and BBO - SVM (Support Vector Machines) with gene ranki…
New model predicts traits from gene expression, accounting for heterogeneity and gene networks.
problem Predicting phenotypes from gene expression data, considering heterogeneity and gene networks.
method Developed a novel model that considers heterogeneity and gene regulatory networks.
result Model performs well on prediction and provides clusters and gene regulatory networks.
New gene selection method improves tumor classification accuracy.
problem Efficiently selecting relevant genes from high-dimensional tumor gene expression data.
method Fuzzy-Rough Set Theory for feature dependency analysis.
result The proposed method outperforms state-of-the-art techniques in tumor classification.
Various approaches to gene selection for cancer classification based on microarray data can be found in the literature and they may be grouped into two categories: univariate methods and multivariate methods. Univariate methods look at each gene in the data in isolation from others. They measure the contribution of a p…
New model generates realistic single-cell gene expression data.
problem Generating realistic single-cell gene expression profiles is challenging.
method scLDM, a latent diffusion model using Diffusion Transformers and linear interpolants.
result Superior performance in generating realistic single-cell gene expression data.
A model to fill in missing gene data from spatial studies and scRNA-seq.
problem Imputing missing gene expression measurements from spatial transcriptomics.
method A deep generative model (gimVI) for integrating spatial transcriptomic and scRNA-seq data.
result gimVI outperforms existing methods in imputing missing genes.
We address the problem of synthetic gene design using Bayesian optimization. The main issue when designing a gene is that the design space is defined in terms of long strings of characters of different lengths, which renders the optimization intractable. We propose a three-step approach to deal with this issue. First, …
NO-BEARS algorithm speeds up gene network inference from transcriptomic data.
problem Constructing accurate gene regulatory networks from transcriptomic data.
method NO-BEARS algorithm, based on NOTEARS, with new constraint and polynomial regression loss.
result Significantly reduced computational time and improved accuracy in inferring gene regulatory networks.
InfoSEM infers gene regulatory networks without GT labels, improving performance.
problem Inferring GRNs from gene expression data with high accuracy and avoiding biases.
method InfoSEM uses deep generative models with informative priors (textual gene embeddings).
result InfoSEM outperforms existing models by 38.5% across four datasets.
Stem uses diffusion models to infer gene expression from H&E images.
problem Inference of gene expression from H&E stained images is time-consuming and expensive.
method Conditional diffusion generative model to infer gene expression.
result Stem achieves state-of-the-art performance in spatial gene expression prediction.
The paper tackles controlling gene regulatory networks with noisy measurements and uncertain inputs.
problem Controlling gene regulatory networks with indirect measurements and uncertain inputs.
method Modeling GRNs with POBDS, transforming to a Markov Decision Process, using Gaussian processes for cost function, and applying reinforcement learning and sparsification.
result Near-optimal control strategy for infinite-horizon control of GRNs is found.
In most gene expression data, the number of training samples is very small compared to the large number of genes involved in the experiments. However, among the large amount of genes, only a small fraction is effective for performing a certain task. Furthermore, a small subset of genes is desirable in developing gene e…
Unified framework improves gene prioritization in disease studies.
problem Identifying genes involved in diseases using heterogeneous biological data.
method Network propagation-based gene prioritization with integrated biological information.
result Significant improvements in prioritizing genes not identified by traditional methods.
Bayesian method discovers local causal relationships among genes from gene expression data.
problem Discovering gene regulatory relationships from gene expression data.
method Bayesian approach scoring covariance structures for triplets of normally distributed variables, incorporating background knowledge as priors.
result Stable and conservative posterior probability estimates of local causal structures.
Machine learning predicts gene involvement in axon regeneration.
problem Predicting gene involvement in specific biological processes.
method Extracted 31 features from databases, trained five machine learning models (Random Forest Classifier with 50 submodels), achieved 85.71% test score.
result Models have some predictive capability for gene involvement in axon regeneration.
Graphs improve deep learning for gene expression data.
problem Challenges in applying deep learning to gene expression data due to non-linear signal and low sample sizes.
method Graph Convolutional Neural Networks (GCNN) with dropout and gene embeddings.
result GCNN with graph information provides an advantage in a low data regime.