RNA-binding proteins (RBPs) play crucial roles in many biological processes, e.g. gene regulation. Computational identification of RBP binding sites on RNAs are urgently needed. In particular, RBPs bind to RNAs by recognizing sequence motifs. Thus, fast locating those motifs on RNA sequences is crucial and time-efficie…
New method detects RNA modifications without prior training, revealing novel sites.
problem Detecting RNA modifications with high accuracy and sensitivity.
method Anomaly detection using nanopore raw ionic current signals and nearest neighbor comparison.
result Detects diverse RNA modifications without prior training, including a novel 2'-O-methylated site in DENV.
Solving the RNA inverse folding problem is a critical prerequisite to RNA design, an emerging field in bioengineering with a broad range of applications from reaction catalysis to cancer therapy. Although significant progress has been made in developing machine-based inverse RNA folding algorithms, current approaches s…
Designing RNA molecules has garnered recent interest in medicine, synthetic biology, biotechnology and bioinformatics since many functional RNA molecules were shown to be involved in regulatory processes for transcription, epigenetics and translation. Since an RNA's function depends on its structural properties, the RN…
New methods improve analysis of single cell RNA sequencing data.
problem High dimensionality and complexity of scRNA-seq data.
method Topological Nonnegative Matrix Factorization (TNMF) and Robust Topological NMF (rTNMF).
result TNMF and rTNMF significantly outperform other NMF-based methods.
Long non-coding RNAs (lncRNAs) are a class of non-coding RNAs which play a significant role in several biological processes. RNA-seq based transcriptome sequencing has been extensively used for identification of lncRNAs. However, accurate identification of lncRNAs in RNA-seq datasets is crucial for exploring their char…
Predicting RNA base distances using a large language model.
problem Accurately predicting RNA structural information, especially distance maps.
method Using a large pretrained RNA language model coupled with a transformer.
result The model can accurately infer RNA base distances from sequence data.
Non-coding RNA (ncRNA) are RNA sequences which don't code for a gene but instead carry important biological functions. The task of ncRNA classification consists in classifying a given ncRNA sequence into its family. While it has been shown that the graph structure of an ncRNA sequence folding is of great importance for…
Recent advances in high-throughput cDNA sequencing (RNA-Seq) technology have revolutionized transcriptome studies. A major motivation for RNA-Seq is to map the structure of expressed transcripts at nucleotide resolution. With accurate computational tools for transcript reconstruction, this technology may also become us…
Flexible models cluster RNA sequencing data.
problem Clustering discrete data from RNA sequencing studies.
method Finite mixtures of multivariate Poisson-log normal factor analyzers with constraints.
result Models give favorable clustering performance on real and simulated data.
With ongoing developments and innovations in single-cell RNA sequencing methods, advancements in sequencing performance could empower significant discoveries as well as new emerging possibilities to address biological and medical investigations. In the study, we will be using the dataset collected by the authors of Sys…
Develops a method to infer cell trajectories from RNA sequencing data.
problem Inferring cell trajectories from single cell RNA-sequencing data.
method Entropy-regularized optimal transport for global optimization.
result Proves and implements a method to recover ground truth trajectories from limited samples.
In this work we propose a method to compute continuous embeddings for kmers from raw RNA-seq data, without the need for alignment to a reference genome. The approach uses an RNN to transform kmers of the RNA-seq reads into a 2 dimensional representation that is used to predict abundance of each kmer. We report that our…
Machine learning accurately diagnoses cancer from whole genome sequencing data.
problem Accurate cancer diagnosis at all stages.
method Novel MLAC (Machine Learning Against Cancer) method using next-gen RNA sequencing.
result Perfect precision, sensitivity, and specificity achieved for most tumor types.
sgdGMF efficiently estimates generalized matrix factorization models for single-cell RNA sequencing data.
problem Challenges in dimensionality reduction for large single-cell RNA sequencing datasets.
method Scalable adaptive stochastic gradient descent algorithm for generalized matrix factorization models.
result sgdGMF outperforms existing methods in scalability and accuracy for large datasets.
Paper proposes a new method for sparse spectral clustering on Stiefel manifold.
problem Sparse spectral clustering on Stiefel manifold with nonsmooth and nonconvex objective.
method Proposes a manifold proximal linear method (ManPL) to solve the original SSC formulation.
result Demonstrates the advantage of ManPL over existing methods on single-cell RNA sequencing data.
Symmetric CNNs improve sequential recommendation and protein structure prediction.
problem Improving prediction accuracy in sequential recommendation and protein structure inference.
method Developed a CNN architecture that preserves symmetry in convolutional layers, using parameterized convolutional kernels.
result Symmetric structured CNNs achieve better performance with fewer parameters.
The paper develops methods for causal inference from single-cell RNA sequencing data with multiple outcomes.
problem Causal inference from single-cell RNA sequencing data with multiple heterogeneous outcomes.
method Generic semiparametric inference framework for doubly robust estimation with multiple derived outcomes.
result Demonstrates the use of semiparametric inferential results for estimating causal effects in genomics.
In this paper, we explore the limitations of PCA as a dimension reduction technique and study its extension, projection pursuit (PP), which is a broad class of linear dimension reduction methods. We first discuss the relevant concepts and theorems and then apply PCA and PP (with negative standardized Shannon's entropy …
RNA structures show that a significant portion of bases do not form hydrogen bonds.
problem Understanding the unpaired bases in RNA secondary structures.
method Comparing random words in free groups to RNA sequences, analyzing word lengths.
result The expected fraction of unpaired bases converges to a constant λ2. We propose a probabilistic model for interpreting gene expression levels that are observed through single-cell RNA sequencing. In the model, each cell has a low-dimensional latent representation. Additional latent variables account for technical effects that may erroneously set some observations of gene expression leve…
Improved neural transducer model outperforms attention model on longer sequences.
problem Improving performance of neural transducer models.
method Comparison of training criteria (marginalization vs. maximum approximation), model generalization, and output label topology.
result Final transducer model outperforms attention model by over 6% relative WER on Switchboard 300h.
In biological research machine learning algorithms are part of nearly every analytical process. They are used to identify new insights into biological phenomena, interpret data, provide molecular diagnosis for diseases and develop personalized medicine that will enable future treatments of diseases. In this paper we (1…
We propose a probabilistic model for interpreting gene expression levels that are observed through single-cell RNA sequencing. In the model, each cell has a low-dimensional latent representation. Additional latent variables account for technical effects that may erroneously set some observations of gene expression leve…
The paper proposes a method to infer differentiation trees from RNA velocity data.
problem Reconstructing dynamic cellular processes from sequencing data.
method Defining varifold distances between RNA velocity curves to approximate shortest-path distances in a tree.
result The varifold distance method approximates the shortest-path distance in a tree isomorphic to the target differentiation tree.
Generative Distribution Embeddings learn multiscale representations of distributions.
problem Learning representations of entire distributions for multiscale reasoning.
method Introducing GDE framework that lifts autoencoders to the space of distributions, using conditional generative models and distributional invariance.
result GDEs learn predictive sufficient statistics embedded in Wasserstein space, recovering distances and trajectories for Gaussian and Gaussian mixture distributions.
Machine learning improves RNA secondary structure prediction.
problem Stagnant performance of RNA secondary structure prediction methods.
method Machine learning, especially deep learning, is used to predict RNA secondary structures.
result Machine learning methods have improved the prediction of RNA secondary structures.
A comprehensive benchmark of 15 scRNA-seq imputation methods across various datasets and analyses.
problem Imputation of single-cell RNA sequencing data to recover latent transcriptional signals.
method Evaluation of 15 imputation methods across 30 datasets and 6 downstream analyses.
result Traditional methods generally outperform DL-based methods in scRNA-seq data analysis.
Proposes selective inference for testing differences in means between clusters.
problem Inflated type I error rate when testing differences in means between clusters.
method Selective inference approach to control selective type I error rate.
result Controls selective type I error rate by accounting for data-driven cluster definition.
New method uses dendrograms for better mixture model selection and clustering.
problem Selecting the correct number of components in finite mixture models.
method Hierarchical clustering tree derived from overfitted latent mixing measures.
result Consistently selects the true number of mixing components and optimal convergence rate for parameter estimation.
MSBM extends SB for multi-marginal trajectory inference.
problem Trajectory inference from multiple discrete snapshots.
method Multi-Marginal Schrödinger Bridge Matching (MSBM) using iterative Markovian fitting (IMF).
result MSBM effectively captures complex trajectories and respects intermediate distributions.
JojoSCL improves scRNA-seq clustering by reducing intra-cluster dispersion.
problem High dimensionality and sparsity of scRNA-seq data challenge clustering models.
method Integrates shrinkage estimator and contrastive learning for improved clustering.
result JojoSCL outperforms existing methods on ten scRNA-seq datasets.
Data thinning splits observations into independent parts for convolution-closed distributions.
problem Validation of unsupervised learning results in settings with limited data.
method Data thinning, splitting observations into independent parts following the same distribution.
result Data thinning provides an attractive alternative to cross-validation in settings with limited sample splitting.
New method learns flows between multiple distributions efficiently.
problem Learning dynamic transport maps between multiple empirical distributions.
method Combining flow matching and dynamic optimal transport with potential terms.
result OTP-FM achieves state-of-the-art performance on various datasets.
Motivation: Deep learning architectures have recently demonstrated their power in predicting DNA- and RNA-binding specificities. Existing methods fall into three classes: Some are based on Convolutional Neural Networks (CNNs), others use Recurrent Neural Networks (RNNs), and others rely on hybrid architectures combinin…
Generalized quandle polynomial used for stuquandles, stuck links, and RNA folding.
problem Defining polynomial invariants for stuquandles, stuck links, and RNA foldings.
method Introduced a generalized quandle polynomial and proved its invariance for stuquandles. Used this invariant to define polynomials for stuck links and RNA foldings.
result Polynomial invariants for stuquandles, stuck links, and RNA foldings.
Deep learning predicts RNA degradation from crowdsourced data.
problem Predicting RNA degradation to improve thermostability.
method Crowdsourced machine learning competition on Kaggle.
result 41% of predictions matched experimental data, and models generalized to longer RNA molecules.
TrajectoryNet models dynamic cellular trajectories using optimal transport.
problem Modeling continuous and non-linear paths in dynamic processes.
method Continuous normalizing flows linked to dynamic optimal transport.
result TrajectoryNet improves interpolation of cellular distributions.
A new method improves data representation for diverse tasks.
problem Learning meaningful representations for tasks like batch correction and counterfactual inference.
method Contrastive Mixture of Posteriors (CoMP) method using misalignment penalties.
result CoMP achieves state-of-the-art performance on challenging tasks.
Proposes CCCVAE for better single-cell clustering with cell-cell communication.
problem Improving single-cell RNA sequencing clustering by incorporating cell-cell communication.
method Integrates cell-cell communication into a variational autoencoder framework.
result Empirical results show CCCVAE outperforms standard VAEs in clustering performance.
Nonparametric IPSS selects features with false discovery control.
problem Feature selection in high-dimensional data with theoretical false discovery control.
method Integrated Path Stability Selection (IPSS) applied to nonparametric feature importance scores.
result IPSS accurately controls false discovery rate and detects more true positives than existing methods.
A deep learning model organizes RNA graphs to reveal folding patterns and properties.
problem Organizing and understanding the complex folding patterns of RNA secondary structures.
method Geometric scattering autoencoder (GSAE) network for learning graph embeddings.
result GSAE accurately reflects bistable RNA structures and can sample new folding trajectories.
New benchmarks for RNA 3D structure-function modeling.
problem Lack of standardized benchmarks for RNA deep learning.
method Developed seven benchmark datasets, provided tools for data handling, and offered a user-friendly environment for model comparison.
result Demonstrated utility with baseline results using a relational graph neural network.
Single-cell RNA sequencing (scRNA-seq) is a fast growing approach to measure the genome-wide transcriptome of many individual cells in parallel, but results in noisy data with many dropout events. Existing methods to learn molecular signatures from bulk transcriptomic data may therefore not be adapted to scRNA-seq data…
In recent years, the advances in single-cell RNA-seq techniques have enabled us to perform large-scale transcriptomic profiling at single-cell resolution in a high-throughput manner. Unsupervised learning such as data clustering has become the central component to identify and characterize novel cell types and gene exp…
New framework quantifies uncertainty in flexible density-based clustering.
problem Uncertainty quantification in clustering with non-parametric density estimation.
method Martingale posterior distributions and density-based clustering.
result Efficient GPU-compatible inference on clustering structures with uncertainty.
Estimates change point in high-dimensional dynamic graphical models.
problem Detecting change points in high-dimensional graphical models.
method Developed an estimator with Op(ψ−2) rate of convergence, established asymptotic distribution under high-dimensional scaling. result Asymptotic distribution characterized under vanishing and non-vanishing jump size regimes.
E2Efold predicts RNA secondary structures better than previous methods.
problem RNA secondary structure prediction with constraints.
method End-to-end deep learning model using unrolled algorithms to enforce constraints.
result E2Efold predicts significantly better structures, especially for pseudoknotted structures.