Deep model interprets gene expression from single-cell RNA sequencing.
problem Interpreting gene expression levels from single-cell RNA sequencing data.
method Probabilistic model with neural network conditional distributions, variational inference, stochastic optimization.
result The model outperforms state-of-the-art methods for differential expression analysis.
This study reviews and evaluates clustering methods for single-cell RNA-seq data.
problem Identifying and characterizing novel cell types from single-cell RNA-seq data.
method Review and performance comparison of clustering methods.
result Performance comparison experiments on two datasets.
A deep model detects differentially expressed genes from single-cell RNA seq data.
problem Detecting differentially expressed genes from single-cell RNA sequencing data.
method Probabilistic model with neural network conditional distributions, variational inference, stochastic optimization.
result The model outperforms state-of-the-art methods for differential expression detection.
MarkerMap selects key genes for cell type analysis in single-cell RNA-seq.
problem Selecting informative genes from large single-cell RNA-seq datasets is challenging and computationally intensive.
method MarkerMap is a generative model that identifies minimal gene sets explaining cell type variability.
result MarkerMap outperforms existing methods in both supervised and unsupervised marker selection.
New model reconstructs cell differentiation paths from single-cell RNA data.
problem Reconstructing dynamic biological phenomena from noisy, heterogeneous, and sparse single-cell RNA-seq data.
method Developed a generative model using Dirichlet diffusion tree and Markov chain Monte Carlo sampler.
result Recovered latent trajectories from simulated single-cell transcriptomes.
Recent advances in high-throughput cDNA sequencing (RNA-Seq) technology have revolutionized transcriptome studies. A major motivation for RNA-Seq is to map the structure of expressed transcripts at nucleotide resolution. With accurate computational tools for transcript reconstruction, this technology may also become us…
A comprehensive benchmark of 15 scRNA-seq imputation methods across various datasets and analyses.
problem Imputation of single-cell RNA sequencing data to recover latent transcriptional signals.
method Evaluation of 15 imputation methods across 30 datasets and 6 downstream analyses.
result Traditional methods generally outperform DL-based methods in scRNA-seq data analysis.
Develops a model for RNA-seq data clustering.
problem Challenges in clustering RNA-seq count data with variable selection.
method Sparse negative binomial mixture model with lasso or fused lasso regularization.
result Superior performance in clustering accuracy, feature selection, and biological interpretation.
The paper develops methods for causal inference from single-cell RNA sequencing data with multiple outcomes.
problem Causal inference from single-cell RNA sequencing data with multiple heterogeneous outcomes.
method Generic semiparametric inference framework for doubly robust estimation with multiple derived outcomes.
result Demonstrates the use of semiparametric inferential results for estimating causal effects in genomics.
Paper reviews machine learning in RNASeq DE analysis.
problem Improving accuracy and efficiency in RNASeq DE analysis.
method Illustrates workflow, performs DE analysis, presents machine learning improvements.
result Demonstrates machine learning's role and capabilities in RNASeq DE analysis.
Discovering genomic structure through learned neural architectures.
problem Decoding the complex, unknown structure of human genomes using deep learning.
method Developed a novel search algorithm to learn optimal architectures for genomic data.
result Architectures learned from RNA expression data predict gene regulatory structure and identify key sequence motifs.
sgdGMF efficiently estimates generalized matrix factorization models for single-cell RNA sequencing data.
problem Challenges in dimensionality reduction for large single-cell RNA sequencing datasets.
method Scalable adaptive stochastic gradient descent algorithm for generalized matrix factorization models.
result sgdGMF outperforms existing methods in scalability and accuracy for large datasets.
Develops GMNB model for better NGS data analysis.
problem Lack of tools for analyzing temporal NGS data.
method Integrates gamma Markov chain into negative binomial distribution model.
result GMNB outperforms existing methods in differential expression analysis.
The paper proposes a method to infer differentiation trees from RNA velocity data.
problem Reconstructing dynamic cellular processes from sequencing data.
method Defining varifold distances between RNA velocity curves to approximate shortest-path distances in a tree.
result The varifold distance method approximates the shortest-path distance in a tree isomorphic to the target differentiation tree.
A new method infers neuronal cell types and their gene expression profiles from brain imaging data.
problem Lack of spatial information in single-cell RNA sequencing data.
method Spatial point process mixture model applied to in situ hybridization images.
result Inferred cell types and gene expression profiles validated with single-cell RNA sequencing data.
iDeepA predicts RNA-protein binding sites from RNA sequences using a CNN with attention.
problem Predicting RNA-protein binding sites from raw RNA sequences efficiently.
method Attention based convolutional neural network (iDeepA) encoding RNA sequences into one-hot encoding, followed by a CNN with an attention mechanism.
result iDeepA achieves comparable performance to state-of-the-art methods on CLIP-seq data.
Next-generation sequencing technologies provide a revolutionary tool for generating gene expression data. Starting with a fixed RNA sample, they construct a library of millions of differentially abundant short sequence tags or "reads", which constitute a fundamentally discrete measure of the level of gene expression. A…
Machine learning improves RNA secondary structure prediction.
problem Stagnant performance of RNA secondary structure prediction methods.
method Machine learning, especially deep learning, is used to predict RNA secondary structures.
result Machine learning methods have improved the prediction of RNA secondary structures.
LEARNA uses deep reinforcement learning to design RNA sequences.
problem Designing RNA molecules to satisfy structural constraints.
method LEARNA employs deep reinforcement learning to train a policy network for RNA design.
result LEARNA achieves new state-of-the-art performance in RNA Design benchmarks.
JojoSCL improves scRNA-seq clustering by reducing intra-cluster dispersion.
problem High dimensionality and sparsity of scRNA-seq data challenge clustering models.
method Integrates shrinkage estimator and contrastive learning for improved clustering.
result JojoSCL outperforms existing methods on ten scRNA-seq datasets.
Machine learning models improve cancer type classification accuracy.
problem Early and accurate cancer diagnosis is challenging due to high costs and biological marker limitations.
method Assessed five machine learning algorithms for 17 cancer types using RNA-seq data.
result Ensemble algorithms achieve 100% accuracy in 14 out of 17 cancer types.
The paper improves Fisher-Pitman tests for Poisson mixtures, detecting autism-related genes.
problem Detecting differentially expressed genes between autism and control subjects.
method Nonparametric Poisson mixtures and Fisher-Pitman permutation tests.
result The tests reveal genes missed by common methods, demonstrating rate optimality.
SentRNA improves RNA design by integrating human strategies.
problem Designing sequences for large or complex RNA targets.
method SentRNA uses a neural network trained on human-designed RNA sequences.
result SentRNA solves complex targets previously unsolvable by machines.
Pipeline predicts cancer survival using multi-modal gene expression data.
problem Predicting cancer survival from gene expression data is challenging due to limited labeled samples and high dimensionality.
method Graph-based semi-supervised learning (GSSL) with manifold learning and stacked generalization.
result Fusing different modalities of gene expression data improves survival prediction accuracy.
Deep learning predicts RNA degradation from crowdsourced data.
problem Predicting RNA degradation to improve thermostability.
method Crowdsourced machine learning competition on Kaggle.
result 41% of predictions matched experimental data, and models generalized to longer RNA molecules.
Generalized quandle polynomial used for stuquandles, stuck links, and RNA folding.
problem Defining polynomial invariants for stuquandles, stuck links, and RNA foldings.
method Introduced a generalized quandle polynomial and proved its invariance for stuquandles. Used this invariant to define polynomials for stuck links and RNA foldings.
result Polynomial invariants for stuquandles, stuck links, and RNA foldings.
Predicting RNA base distances using a large language model.
problem Accurately predicting RNA structural information, especially distance maps.
method Using a large pretrained RNA language model coupled with a transformer.
result The model can accurately infer RNA base distances from sequence data.
PLIT identifies plant lncRNAs from RNA-seq data with high accuracy.
problem Inaccurate identification of lncRNAs in plant transcriptomic datasets.
method PLIT uses L1 regularization and iRF classification to select optimal features from sequence and codon-bias data.
result PLIT outperforms existing CPC tools in identifying lncRNAs in plant RNA-seq datasets.
The paper tackles extrapolation of gene knockouts effects on RNA counts.
problem Modeling effects of gene knockouts on RNA counts for new perturbations.
method Formulated as a latent variable model with additive perturbation effects, proved identifiability, proposed PDAE for estimation.
result PDAE can accurately predict effects of unseen but identifiable perturbations.
RNA accelerates CNNs for image recognition.
problem Improving the optimization process of CNNs for image recognition.
method Regularized Nonlinear Acceleration (RNA) applied to neural networks.
result RNA improves the optimization process of CNNs slightly.
A deep learning model organizes RNA graphs to reveal folding patterns and properties.
problem Organizing and understanding the complex folding patterns of RNA secondary structures.
method Geometric scattering autoencoder (GSAE) network for learning graph embeddings.
result GSAE accurately reflects bistable RNA structures and can sample new folding trajectories.
New benchmarks for RNA 3D structure-function modeling.
problem Lack of standardized benchmarks for RNA deep learning.
method Developed seven benchmark datasets, provided tools for data handling, and offered a user-friendly environment for model comparison.
result Demonstrated utility with baseline results using a relational graph neural network.
ContrastiveVI+ models CRISPR screens with noisy guide efficiency.
problem Noisy guide efficiency in CRISPR screens.
method Generative modeling framework that disentangles perturbation-induced from shared variations.
result ContrastiveVI+ better recovers perturbation-induced variations and identifies cells without edits.
E2Efold predicts RNA secondary structures better than previous methods.
problem RNA secondary structure prediction with constraints.
method End-to-end deep learning model using unrolled algorithms to enforce constraints.
result E2Efold predicts significantly better structures, especially for pseudoknotted structures.
New method detects RNA modifications without prior training, revealing novel sites.
problem Detecting RNA modifications with high accuracy and sensitivity.
method Anomaly detection using nanopore raw ionic current signals and nearest neighbor comparison.
result Detects diverse RNA modifications without prior training, including a novel 2'-O-methylated site in DENV.
New invariants for RNA foldings and stuck links defined.
problem Defining invariants for RNA foldings and stuck links.
method Assigning Boltzmann weights at classical and stuck crossings.
result Explicit computations of new invariants provided.
Unsupervised learning helps understand genome-wide patterns and miRNA regulation.
problem Understanding genome-wide biological insights and miRNA regulation.
method Review of unsupervised learning algorithms for genome informatics and miRNA regulation.
result Reviewed several unsupervised learning methods for genome and miRNA analysis.
Study uses knot theory to model RNA foldings, emphasizing both entanglement and intrachain interactions.
problem Modeling RNA foldings considering both entanglement and intrachain interactions.
method Combines knot theory with embedded rigid vertex graphs to emphasize both entanglement and intrachain interactions of RNA foldings.
result Defines and computes a coloring counting invariant for stuck links, providing explicit computations for arc diagrams of RNA foldings.
Graph ConvNet improves ncRNA classification accuracy.
problem Classifying non-coding RNA sequences into families.
method Graph Convolutional Network model trained on raw RNA graphs.
result 85.73% accuracy and 85.61% F1-score over 13 classes.
Method computes embeddings for RNA-seq data without genome alignment.
problem No need for genome alignment for RNA-seq data analysis.
method RNN transforms kmers into 2D latent space for transcriptomic analysis.
result Captures DNA sequence similarity and abundance in latent space.
Graph Canonical Correlation Analysis improves CCA for multiomics datasets.
problem Limited ability of conventional CCA methods to incorporate structured patterns in cross-correlation matrices.
method Graph Canonical Correlation Analysis (gCCA) calculates canonical correlations based on the graph structure of cross-correlation matrices.
result gCCA outperforms competing CCA methods in simulations and multiomics dataset analysis.
Improved GPLVM model for single-cell RNA-seq data.
problem Lack of effective scalable models for clustering cell types in large-scale single-cell RNA-seq data.
method Introduces amortized stochastic variational Bayesian GPLVM (BGPLVM) tailored for single-cell RNA-seq.
result Matches the performance of scVI on synthetic and real-world datasets and reveals more interpretable latent structures.
Machine learning accurately diagnoses cancer from whole genome sequencing data.
problem Accurate cancer diagnosis at all stages.
method Novel MLAC (Machine Learning Against Cancer) method using next-gen RNA sequencing.
result Perfect precision, sensitivity, and specificity achieved for most tumor types.
HSSE framework embeds single-cell RNA-seq data at multiple scales.
problem Capturing heterogeneous local structure in single-cell RNA-seq data.
method Hierarchical sheaf spectral embedding (HSSE) framework.
result HSSE achieves competitive or improved performance in single-cell RNA-seq data representation learning.
Method identifies key features for clustering in high-dimensional data.
problem Understanding hidden patterns in high-dimensional data.
method Unsupervised feature selection based on discriminative power.
result 27 key transcription factors identified, 18 known to define cell states.
New model clusters cells and individuals, revealing genetic influences on cell types.
problem Clustering nested data with group-level and observation-level variables.
method Nested Atoms Model (NAM), Bayesian nonparametric approach.
result Identifies clusters of genetically similar individuals with homogeneous cell-type profiles.
RNA structures show that a significant portion of bases do not form hydrogen bonds.
problem Understanding the unpaired bases in RNA secondary structures.
method Comparing random words in free groups to RNA sequences, analyzing word lengths.
result The expected fraction of unpaired bases converges to a constant λ2. New methods improve analysis of single cell RNA sequencing data.
problem High dimensionality and complexity of scRNA-seq data.
method Topological Nonnegative Matrix Factorization (TNMF) and Robust Topological NMF (rTNMF).
result TNMF and rTNMF significantly outperform other NMF-based methods.