ABC uses GPs to speed up estimating bacterial gene transfers.
problem Intractable likelihood in bacterial gene transfer models.
method Gaussian process modeling to approximate discrepancies.
result GP choice significantly affects ABC posterior accuracy.
Develops probabilistic models for gene regulatory network inference.
problem Challenges in reconstructing gene regulatory networks from genome-wide data.
method Two complementary frameworks: PMF-GRN and GLM-Prior.
result Probabilistic inference refines regulatory estimates with quantified uncertainty.
Paper proposes dp-VAE for preserving spatial context in gene expression data.
problem Inaccessibility of spatial context in single-cell gene expression data.
method Generic representation learning and transfer learning framework with a distance-preserving regularizer.
result dp-VAE effectively reconstructs and imputes spatial context from gene expression data.
Estimates target GGM using auxiliary studies with false discovery rate control.
problem Estimating high-dimensional GGMs from related studies.
method Transfer learning with Trans-CLIME and debiased Trans-CLIME estimators.
result Debiased Trans-CLIME estimator provides element-wise asymptotic normality and false discovery rate control.
The paper improves high-dimensional linear regression prediction and estimation using auxiliary samples.
problem Estimating and predicting high-dimensional linear regression models with auxiliary samples.
method Proposes Trans-Lasso for data-driven transfer learning, establishing optimality for prediction and estimation.
result Knowledge from auxiliary samples can improve learning performance in target problems.
Trans-Glasso uses transfer learning to estimate precision matrices from related studies.
problem Challenges in precision matrix estimation with limited target samples.
method Two-step transfer learning: multi-task learning followed by differential network estimation.
result Trans-Glasso achieves minimax optimality under certain conditions and outperforms baseline methods in simulations and real-world applications.
The paper develops new data organization methods for non-regular data.
problem Organizing high-dimensional data with arbitrary feature order and non-regular observations.
method Data-driven tree transforms and metrics based on iterative refinement.
result Improved clustering of tumor samples from multiple gene expression cohorts.
A new method uses gene interaction networks to predict gene functions.
problem Predicting gene functions from gene interactions.
method Context graph kernel approach in a machine learning framework.
result The proposed method outperforms linkage-assumption-based methods.
Model learns to select relevant clinical variables for disease subtype prediction from small data.
problem Few-shot disease subtype prediction from small genomic data.
method Meta learning Prototypical Network with feature selection and sample reweighting.
result Superior performance in predicting disease subtypes and identifying genes.
VEGN uses graph neural networks to predict disease-causing mutations from genetic variants.
problem Identifying disease-causing mutations from millions of genetic variants.
method VEGN employs a graph neural network on a heterogeneous graph of genes and variants, learning gene-gene interactions.
result VEGN outperforms existing state-of-the-art models in variant effect prediction.
New method combines knowledge from related tasks to improve test task performance.
problem Improving performance on a test task using knowledge from related tasks.
method Relaxing covariate shift assumption, focusing on invariant subset of predictors.
result Optimal prediction in Domain Generalization using the invariant subset.
GSAE autoencoder models gene sets for better cancer subtype and prognosis analysis.
problem Inter-gene set associations not considered in gene set-based analyses.
method Gene superset autoencoder model incorporating prior gene sets.
result Gene supersets retain biological features and are reproducible for cancer subtype and prognosis.
Crowdsourced gene set queries reveal protein-protein interactions and gene-gene associations.
problem Lack of integrative analysis of diverse gene set queries.
method Harnessed thousands of user-submitted gene sets to construct a global gene-gene association network.
result The constructed network recapitulates known protein-protein interactions and gene-gene functional associations.
A new method for identifying significant gene subsets improves disease prediction.
problem Identifying significant subsets of genes for disease prediction.
method Kernel gene shaving using influence function of kernel CCA.
result The proposed method outperformed three popular gene selection methods.
EpiRL learns to detect gene-gene interactions.
problem Computational challenges in epistasis detection.
method Modeling epistasis as a Markov Decision Process and using reinforcement learning.
result EpiRL discovers highly interacted genes.
Bayesian model learns cell types and gene networks from two data views.
problem Estimating cell types and their regulatory networks from single-cell gene expression and epigenetic data.
method Symphony Bayesian hierarchical multi-view mixture model with Variational EM inference.
result Symphony outperforms other methods in learning cell types and regulatory networks.
New method expands seed genes to functionally related clusters.
problem Discovering functionally related genes lacking GO terms.
method Semi-supervised learning with positive and unlabeled examples.
result LPU approaches significantly outperform existing methods.
A new method for joint eQTL mapping and gene network estimation.
problem Discovering SNP-gene relationships and gene-gene relationships in gene expression regulation.
method L1-2 regularized multi-task graphical lasso (L1-2 GLasso).
result Competitive performance on capturing true sparse structures of eQTL mapping and gene network.
New method handles correlated genes for better genomic prediction.
problem Technical issues with highly correlated genes in prediction models.
method Grouping algorithm that treats correlated genes as a group and uses their common patterns.
result Significantly outperforms standard models in prediction and feature selection.
VGAE learns gene-disease associations from networks, predicting disease-genes.
problem Predicting gene-disease associations from disease-gene networks.
method Introducing VGAE, a variational graph auto-encoder for disease-gene prediction.
result VGAE and C-VGAE outperform baseline methods in disease-gene prediction.
Elucidating the genetic basis of human diseases is a central goal of genetics and molecular biology. While traditional linkage analysis and modern high-throughput techniques often provide long lists of tens or hundreds of disease gene candidates, the identification of disease genes among the candidates remains time-con…
Proposes a method to identify relevant genes in autism-related diseases using auxiliary information.
problem Identifying relevant genes in autism-related diseases from diverse data sources.
method Uses logistic regression to filter irrelevant genes and clusters relevant genes into cohesive groups using adjacency matrix.
result Superior performance and robustness in finite samples observed in simulation studies.
Improved differentially private drug sensitivity prediction using compact representations.
problem Challenges in differentially private machine learning with genomic data.
method Representation learning using variational autoencoders, PCA, and random projection.
result Variational autoencoders provide the most accurate predictions for differentially private drug sensitivity prediction.
Bayesian optimization tackles gene design by modeling cell behavior and optimizing multiple aspects.
problem Designing genes in long string spaces is intractable.
method Three-step approach: Gaussian process model, multi-task acquisition function, evaluation function.
result Illustrated the approach's performance in mammalian cells.
Robust method detects gene-gene interactions in imaging genetics data.
problem Detecting nonlinear gene-gene interactions in imaging genetics data.
method Robust Kernel Canonical Correlation Analysis (RKCCA) with influence function variance estimation.
result The proposed robust RKCCA method outperforms state-of-the-art methods in detecting gene-gene interactions.
We present the extention and application of a new unsupervised statistical learning technique--the Partition Decoupling Method--to gene expression data. Because it has the ability to reveal non-linear and non-convex geometries present in the data, the PDM is an improvement over typical gene expression analysis algorith…
SDSR reconstructs species trees from genetic markers efficiently.
problem Challenges in reconstructing species trees from genetic data.
method Spectral divide-and-conquer approach based on graph theory.
result SDSR achieves up to 10-fold faster runtime with comparable accuracy.
Collaborative filtering predicts drug responses from gene expression data.
problem Predicting drug responses from large gene expression datasets with limited samples.
method Low-rank matrix factorization and latent linear regression.
result The proposed method outperforms state-of-the-art methods in predicting drug-gene associations.
Hybrid method selects fewer genes for cancer classification.
problem Selecting genes for cancer classification from microarray data.
method Hybrid of univariate (LIK) and multivariate (RFE) feature selection methods.
result Hybrid method selects fewer genes with similar or better accuracy.
A novel method selects genes for high-dimensional gene expression data with class imbalance.
problem Class imbalance in gene expression datasets.
method Synthetic data balancing, greedy search, weighted robust score.
result The proposed method outperforms existing feature selection procedures.
The problem of multilabel classification when the labels are related through a hierarchical categorization scheme occurs in many application domains such as computational biology. For example, this problem arises naturally when trying to automatically assign gene function using a controlled vocabularies like Gene Ontol…
Identifying latent structure in large data matrices is essential for exploring biological processes. Here, we consider recovering gene co-expression networks from gene expression data, where each network encodes relationships between genes that are locally co-regulated by shared biological mechanisms. To do this, we de…
The method integrates survival constraints into NMF for identifying survival-associated gene clusters.
problem Understanding and interpreting high-dimensional biological data for disease markers.
method Cox proportional hazards regression integrated with NMF via proportional hazards non-negative matrix factorization.
result The method can uncover survival-associated gene clusters in cancer gene expression data.
New methods detect continuous variation in single-cell data.
problem Continuous variation within and between cell types not detected by discrete analyses.
method Three topologically motivated mathematical methods for unsupervised feature selection.
result Detect additional biologically meaningful genes with coherent expression patterns.
Microarray cancer gene expression data comprise of very high dimensions. Reducing the dimensions helps in improving the overall analysis and classification performance. We propose two hybrid techniques, Biogeography - based Optimization - Random Forests (BBO - RF) and BBO - SVM (Support Vector Machines) with gene ranki…
New model predicts traits from gene expression, accounting for heterogeneity and gene networks.
problem Predicting phenotypes from gene expression data, considering heterogeneity and gene networks.
method Developed a novel model that considers heterogeneity and gene regulatory networks.
result Model performs well on prediction and provides clusters and gene regulatory networks.
New gene selection method improves tumor classification accuracy.
problem Efficiently selecting relevant genes from high-dimensional tumor gene expression data.
method Fuzzy-Rough Set Theory for feature dependency analysis.
result The proposed method outperforms state-of-the-art techniques in tumor classification.
Integrative analysis connects gene expression to phenotypes.
problem Linking genotype to phenotype through complex phenotypes.
method Automated multi-dimensional profiling, sequence of gene cluster combinations, integrative approach to gene network modules.
result Robust profiling and classification performance in phenotypes.
New model generates realistic single-cell gene expression data.
problem Generating realistic single-cell gene expression profiles is challenging.
method scLDM, a latent diffusion model using Diffusion Transformers and linear interpolants.
result Superior performance in generating realistic single-cell gene expression data.
A model to fill in missing gene data from spatial studies and scRNA-seq.
problem Imputing missing gene expression measurements from spatial transcriptomics.
method A deep generative model (gimVI) for integrating spatial transcriptomic and scRNA-seq data.
result gimVI outperforms existing methods in imputing missing genes.
NO-BEARS algorithm speeds up gene network inference from transcriptomic data.
problem Constructing accurate gene regulatory networks from transcriptomic data.
method NO-BEARS algorithm, based on NOTEARS, with new constraint and polynomial regression loss.
result Significantly reduced computational time and improved accuracy in inferring gene regulatory networks.
InfoSEM infers gene regulatory networks without GT labels, improving performance.
problem Inferring GRNs from gene expression data with high accuracy and avoiding biases.
method InfoSEM uses deep generative models with informative priors (textual gene embeddings).
result InfoSEM outperforms existing models by 38.5% across four datasets.
Stem uses diffusion models to infer gene expression from H&E images.
problem Inference of gene expression from H&E stained images is time-consuming and expensive.
method Conditional diffusion generative model to infer gene expression.
result Stem achieves state-of-the-art performance in spatial gene expression prediction.
The paper tackles controlling gene regulatory networks with noisy measurements and uncertain inputs.
problem Controlling gene regulatory networks with indirect measurements and uncertain inputs.
method Modeling GRNs with POBDS, transforming to a Markov Decision Process, using Gaussian processes for cost function, and applying reinforcement learning and sparsification.
result Near-optimal control strategy for infinite-horizon control of GRNs is found.
In most gene expression data, the number of training samples is very small compared to the large number of genes involved in the experiments. However, among the large amount of genes, only a small fraction is effective for performing a certain task. Furthermore, a small subset of genes is desirable in developing gene e…
Unified framework improves gene prioritization in disease studies.
problem Identifying genes involved in diseases using heterogeneous biological data.
method Network propagation-based gene prioritization with integrated biological information.
result Significant improvements in prioritizing genes not identified by traditional methods.
Bayesian method discovers local causal relationships among genes from gene expression data.
problem Discovering gene regulatory relationships from gene expression data.
method Bayesian approach scoring covariance structures for triplets of normally distributed variables, incorporating background knowledge as priors.
result Stable and conservative posterior probability estimates of local causal structures.
Machine learning predicts gene involvement in axon regeneration.
problem Predicting gene involvement in specific biological processes.
method Extracted 31 features from databases, trained five machine learning models (Random Forest Classifier with 50 submodels), achieved 85.71% test score.
result Models have some predictive capability for gene involvement in axon regeneration.