A new method for identifying significant gene subsets improves disease prediction.
problem Identifying significant subsets of genes for disease prediction.
method Kernel gene shaving using influence function of kernel CCA.
result The proposed method outperformed three popular gene selection methods.
In this work a new way to calculate the multivariate joint entropy is presented. This measure is the basis for a fast information-theoretic based evaluation of gene relevance in a Microarray Gene Expression data context. Its low complexity is based on the reuse of previous computations to calculate current feature rele…
Identifying latent structure in large data matrices is essential for exploring biological processes. Here, we consider recovering gene co-expression networks from gene expression data, where each network encodes relationships between genes that are locally co-regulated by shared biological mechanisms. To do this, we de…
Microarray cancer gene expression data comprise of very high dimensions. Reducing the dimensions helps in improving the overall analysis and classification performance. We propose two hybrid techniques, Biogeography - based Optimization - Random Forests (BBO - RF) and BBO - SVM (Support Vector Machines) with gene ranki…
A novel method selects genes for high-dimensional gene expression data with class imbalance.
problem Class imbalance in gene expression datasets.
method Synthetic data balancing, greedy search, weighted robust score.
result The proposed method outperforms existing feature selection procedures.
In most gene expression data, the number of training samples is very small compared to the large number of genes involved in the experiments. However, among the large amount of genes, only a small fraction is effective for performing a certain task. Furthermore, a small subset of genes is desirable in developing gene e…
New gene selection method improves tumor classification accuracy.
problem Efficiently selecting relevant genes from high-dimensional tumor gene expression data.
method Fuzzy-Rough Set Theory for feature dependency analysis.
result The proposed method outperforms state-of-the-art techniques in tumor classification.
Concrete autoencoder selects key features for efficient data reconstruction.
problem Efficiently identifying and selecting important features for data reconstruction.
method Concrete selector layer with temperature-controlled selection during training, followed by reconstruction using a standard neural network.
result Concrete autoencoder selects a small subset of genes that can reconstruct the remaining gene expression levels, improving on existing methods.
Motivation: Modelling methods that find structure in data are necessary with the current large volumes of genomic data, and there have been various efforts to find subsets of genes exhibiting consistent patterns over subsets of treatments. These biclustering techniques have focused on one data source, often gene expres…
MarkerMap selects key genes for cell type analysis in single-cell RNA-seq.
problem Selecting informative genes from large single-cell RNA-seq datasets is challenging and computationally intensive.
method MarkerMap is a generative model that identifies minimal gene sets explaining cell type variability.
result MarkerMap outperforms existing methods in both supervised and unsupervised marker selection.
A method selects key genes from tumor transcriptomics data using kernel methods and improves classification performance.
problem Feature selection for tumor classification using gene expression data.
method Multiple Kernel Learning with latent regularization and non-linear dimensionality reduction.
result Improved tumor classification performance on unseen test samples.
Identifies a sub-matrix with maximal sum in large data matrices.
problem Finding a rectangular sub-matrix with the highest sum of entries.
method Proposes two algorithms: CP-GC and MILP, leveraging problem characteristics.
result CPGC approach tends to be the fastest to produce a good solution.
The linking genotype to phenotype is the fundamental aim of modern genetics. We focus on study of links between gene expression data and phenotype data through integrative analysis. We propose three approaches. 1) The inherent complexity of phenotypes makes high-throughput phenotype profiling a very difficult and labor…
Exclusive row biclustering for gene expression data.
problem Identifying groups of cancer patients with unique types of cancer.
method Combination of biclustering algorithms and combinatorial auction techniques.
result Identification of large span non-overlapping row submatrices.
Unsupervised method selects genes for tumor subtype discovery.
problem High-dimensional tumor gene expression data with noisy variables and heterogeneity.
method Autoencoders for latent space learning, Multiple Kernel Learning for feature selection, clustering.
result Lower redundancy and better clustering performance compared to benchmarks.
New method interprets neural networks at multiple scales.
problem Interpreting neural networks at various scales and identifying important subsets of inputs.
method Rank Projection Trees framework using any scoring function.
result Successfully identifies biologically important genes and gene sets.
Method learns shared and specific factors in multi-study gene expression data.
problem Understanding shared and specific factors in high-dimensional multi-study data.
method Nonlinear multi-study factor model with sparse variational autoencoder.
result Method recovers meaningful shared and specific factors in platelet gene expression data.
Motivation: Cell-biological processes are regulated through a complex network of interactions between genes and their products. The processes, their activating conditions, and the associated transcriptional responses are often unknown. Organism-wide modeling of network activation can reveal unique and shared mechanisms…
MCPCA analyzes shared factors across multiple data contexts.
problem No tools to recover shared factors across multiple contexts.
method Developed a theoretical and algorithmic framework (MCPCA).
result Reveals shared axes of variation across subsets of contexts.
A comprehensive benchmark of 15 scRNA-seq imputation methods across various datasets and analyses.
problem Imputation of single-cell RNA sequencing data to recover latent transcriptional signals.
method Evaluation of 15 imputation methods across 30 datasets and 6 downstream analyses.
result Traditional methods generally outperform DL-based methods in scRNA-seq data analysis.
A new classification method using disjoint centroids and normalized distance.
problem Improving classification accuracy and feature selection.
method Nearest disjoint centroid classifier with normalized distance.
result Our method outperforms other classifiers in terms of misclassification rates and feature usage.
SDSR reconstructs species trees from genetic markers efficiently.
problem Challenges in reconstructing species trees from genetic data.
method Spectral divide-and-conquer approach based on graph theory.
result SDSR achieves up to 10-fold faster runtime with comparable accuracy.
The paper proposes a unified taxonomy for biclustering methods.
problem Lack of a unified taxonomy for biclustering methods.
method Using concept lattices and attribute exploration to build a taxonomy.
result A unified taxonomy for biclustering methods.
Improved drug response prediction using ensemble learning and gene expression signatures.
problem Predicting chemotherapeutic response of cancer cells to drugs.
method Combining machine learning methods and drug-induced gene expression signatures for improved performance.
result Ensemble method improves drug activity prediction accuracy.
A new method uses gene interaction networks to predict gene functions.
problem Predicting gene functions from gene interactions.
method Context graph kernel approach in a machine learning framework.
result The proposed method outperforms linkage-assumption-based methods.
Proposes a new clustering algorithm for high-dimensional data.
problem Challenges of feature selection in high-dimensional clustering.
method An EM algorithm with lasso-type constraints on cluster pairs.
result Identifies informative features and cluster separability.
VEGN uses graph neural networks to predict disease-causing mutations from genetic variants.
problem Identifying disease-causing mutations from millions of genetic variants.
method VEGN employs a graph neural network on a heterogeneous graph of genes and variants, learning gene-gene interactions.
result VEGN outperforms existing state-of-the-art models in variant effect prediction.
The paper proposes a method to infer differentiation trees from RNA velocity data.
problem Reconstructing dynamic cellular processes from sequencing data.
method Defining varifold distances between RNA velocity curves to approximate shortest-path distances in a tree.
result The varifold distance method approximates the shortest-path distance in a tree isomorphic to the target differentiation tree.
GSAE autoencoder models gene sets for better cancer subtype and prognosis analysis.
problem Inter-gene set associations not considered in gene set-based analyses.
method Gene superset autoencoder model incorporating prior gene sets.
result Gene supersets retain biological features and are reproducible for cancer subtype and prognosis.
GSPPCA selects relevant features in high-dimensional data.
problem Difficulty in interpreting sparse principal components.
method Bayesian probabilistic PCA with a relaxation for model selection.
result GSPPCA identifies relevant variables with the same sparsity pattern.
There is no known efficient method for selecting k Gaussian features from n which achieve the lowest Bayesian classification error. We show an example of how greedy algorithms faced with this task are led to give results that are not optimal. This motivates us to propose a more robust approach. We present a Branch and …
EpiRL learns to detect gene-gene interactions.
problem Computational challenges in epistasis detection.
method Modeling epistasis as a Markov Decision Process and using reinforcement learning.
result EpiRL discovers highly interacted genes.
Bayesian model learns cell types and gene networks from two data views.
problem Estimating cell types and their regulatory networks from single-cell gene expression and epigenetic data.
method Symphony Bayesian hierarchical multi-view mixture model with Variational EM inference.
result Symphony outperforms other methods in learning cell types and regulatory networks.
New method expands seed genes to functionally related clusters.
problem Discovering functionally related genes lacking GO terms.
method Semi-supervised learning with positive and unlabeled examples.
result LPU approaches significantly outperform existing methods.
One of the objectives of designing feature selection learning algorithms is to obtain classifiers that depend on a small number of attributes and have verifiable future performance guarantees. There are few, if any, approaches that successfully address the two goals simultaneously. Performance guarantees become crucial…
A new method for joint eQTL mapping and gene network estimation.
problem Discovering SNP-gene relationships and gene-gene relationships in gene expression regulation.
method L1-2 regularized multi-task graphical lasso (L1-2 GLasso).
result Competitive performance on capturing true sparse structures of eQTL mapping and gene network.
We introduce a graph-theoretic approach to extract clusters and hierarchies in complex data-sets in an unsupervised and deterministic manner, without the use of any prior information. This is achieved by building topologically embedded networks containing the subset of most significant links and analyzing the network s…
New method handles correlated genes for better genomic prediction.
problem Technical issues with highly correlated genes in prediction models.
method Grouping algorithm that treats correlated genes as a group and uses their common patterns.
result Significantly outperforms standard models in prediction and feature selection.
Popular online enrichment analysis tools from the field of molecular systems biology provide users with the ability to submit their experimental results as gene sets for individual analysis. Such queries are kept private, and have never before been considered as a resource for integrative analysis. By harnessing gene s…
VGAE learns gene-disease associations from networks, predicting disease-genes.
problem Predicting gene-disease associations from disease-gene networks.
method Introducing VGAE, a variational graph auto-encoder for disease-gene prediction.
result VGAE and C-VGAE outperform baseline methods in disease-gene prediction.
Elucidating the genetic basis of human diseases is a central goal of genetics and molecular biology. While traditional linkage analysis and modern high-throughput techniques often provide long lists of tens or hundreds of disease gene candidates, the identification of disease genes among the candidates remains time-con…
gOMP algorithm selects features for various types of data.
problem Feature selection for scalable molecular data.
method Generalized Orthogonal Matching Pursuit algorithm for multiple types of data.
result gOMP performs similarly or better than LASSO on various datasets.
Proposes a method to identify relevant genes in autism-related diseases using auxiliary information.
problem Identifying relevant genes in autism-related diseases from diverse data sources.
method Uses logistic regression to filter irrelevant genes and clusters relevant genes into cohesive groups using adjacency matrix.
result Superior performance and robustness in finite samples observed in simulation studies.
Robust method detects gene-gene interactions in imaging genetics data.
problem Detecting nonlinear gene-gene interactions in imaging genetics data.
method Robust Kernel Canonical Correlation Analysis (RKCCA) with influence function variance estimation.
result The proposed robust RKCCA method outperforms state-of-the-art methods in detecting gene-gene interactions.
We present the extention and application of a new unsupervised statistical learning technique--the Partition Decoupling Method--to gene expression data. Because it has the ability to reveal non-linear and non-convex geometries present in the data, the PDM is an improvement over typical gene expression analysis algorith…
We define the beta diffusion tree, a random tree structure with a set of leaves that defines a collection of overlapping subsets of objects, known as a feature allocation. A generative process for the tree structure is defined in terms of particles (representing the objects) diffusing in some continuous space, analogou…
Improves convex biclustering for high-dimensional data.
problem Discovering meaningful biclusters in high-dimensional data.
method Biconvex modification with adaptive feature weighting.
result Consistently recovers biclusters and selects features appropriately.
Collaborative filtering predicts drug responses from gene expression data.
problem Predicting drug responses from large gene expression datasets with limited samples.
method Low-rank matrix factorization and latent linear regression.
result The proposed method outperforms state-of-the-art methods in predicting drug-gene associations.