DeepRAM evaluates and selects the best deep learning architecture for DNA/RNA binding specificity prediction.
problem Selecting the best deep learning architecture for predicting DNA/RNA binding specificity.
method Systematic exploration of various deep learning architectures using deepRAM, an end-to-end deep learning tool.
result A k-mer embedding convolutional layer and recurrent layer architecture outperforms other methods.
RNA-binding proteins (RBPs) play crucial roles in many biological processes, e.g. gene regulation. Computational identification of RBP binding sites on RNAs are urgently needed. In particular, RBPs bind to RNAs by recognizing sequence motifs. Thus, fast locating those motifs on RNA sequences is crucial and time-efficie…
With different genomes available, unsupervised learning algorithms are essential in learning genome-wide biological insights. Especially, the functional characterization of different genomes is essential for us to understand lives. In this book chapter, we review the state-of-the-art unsupervised learning algorithms fo…
When analyzing the genome, researchers have discovered that proteins bind to DNA based on certain patterns of the DNA sequence known as "motifs". However, it is difficult to manually construct motifs due to their complexity. Recently, externally learned memory models have proven to be effective methods for reasoning ov…
In this work we propose a method to compute continuous embeddings for kmers from raw RNA-seq data, without the need for alignment to a reference genome. The approach uses an RNN to transform kmers of the RNA-seq reads into a 2 dimensional representation that is used to predict abundance of each kmer. We report that our…
We propose generative neural network methods to generate DNA sequences and tune them to have desired properties. We present three approaches: creating synthetic DNA sequences using a generative adversarial network; a DNA-based variant of the activation maximization ("deep dream") design method; and a joint procedure wh…
Graph Canonical Correlation Analysis improves CCA for multiomics datasets.
problem Limited ability of conventional CCA methods to incorporate structured patterns in cross-correlation matrices.
method Graph Canonical Correlation Analysis (gCCA) calculates canonical correlations based on the graph structure of cross-correlation matrices.
result gCCA outperforms competing CCA methods in simulations and multiomics dataset analysis.
We consider the task of detecting regulatory elements in the human genome directly from raw DNA. Past work has focused on small snippets of DNA, making it difficult to model long-distance dependencies that arise from DNA's 3-dimensional conformation. In order to study long-distance dependencies, we develop and release …
New method optimizes diffusion models without fine-tuning, integrating soft value functions.
problem Optimizing natural design spaces of images, molecules, DNA, RNA, and protein sequences.
method Iterative sampling method integrating soft value functions into diffusion model inference.
result Directly utilizes non-differentiable features/reward feedback, applies to discrete diffusion models.
We develop an algorithm for minimizing a function using n batched function value measurements at each of T rounds by using classifiers to identify a function's sublevel set. We show that sufficiently accurate classifiers can achieve linear convergence rates, and show that the convergence rate is tied to the difficu…
Algorithm optimizes biological sequences using bootstrapped training with a score-conditioned generator.
problem Optimizing biological sequences for a black-box score function.
method Bootstrapped training of score-conditioned generator (BootGen) algorithm.
result Our method outperforms competitive baselines on biological sequential design tasks.
We present a sparse knowledge gradient (SpKG) algorithm for adaptively selecting the targeted regions within a large RNA molecule to identify which regions are most amenable to interactions with other molecules. Experimentally, such regions can be inferred from fluorescence measurements obtained by binding a complement…
Proposes a copula-based model for multi-view clustering with directional dependency.
problem Challenges in integrating multi-source datasets with directional dependency.
method Copula-based multi-view clustering model accounting for directional dependence.
result Ignoring directional dependence negatively impacts clustering performance.
Motivation: A major challenge in the development of machine learning based methods in computational biology is that data may not be accurately labeled due to the time and resources required for experimentally annotating properties of proteins and DNA sequences. Standard supervised learning algorithms assume accurate in…
VSD efficiently learns conditional distributions for combinatorial designs.
problem Learning conditional distributions for rare combinatorial designs.
method Variational Search Distributions (VSD) using variational inference.
result VSD outperforms existing methods on real sequence-design problems.
Machine learning improves RNA secondary structure prediction.
problem Stagnant performance of RNA secondary structure prediction methods.
method Machine learning, especially deep learning, is used to predict RNA secondary structures.
result Machine learning methods have improved the prediction of RNA secondary structures.
LEARNA uses deep reinforcement learning to design RNA sequences.
problem Designing RNA molecules to satisfy structural constraints.
method LEARNA employs deep reinforcement learning to train a policy network for RNA design.
result LEARNA achieves new state-of-the-art performance in RNA Design benchmarks.
DNA Methylation has been the most extensively studied epigenetic mark. Usually a change in the genotype, DNA sequence, leads to a change in the phenotype, observable characteristics of the individual. But DNA methylation, which happens in the context of CpG (cytosine and guanine bases linked by phosphate backbone) dinu…
An evolutionary algorithm separates mixed DNA profiles in forensic genetics.
problem Deconvolving mixed DNA profiles from crime samples.
method Multiple population evolutionary algorithm (MEA) with guided mutation.
result The MEA successfully deconvoluted DNA profiles from crime samples.
The folding structure of the DNA molecule combined with helper molecules, also referred to as the chromatin, is highly relevant for the functional properties of DNA. The chromatin structure is largely determined by the underlying primary DNA sequence, though the interaction is not yet fully understood. In this paper we…
Deep learning predicts RNA degradation from crowdsourced data.
problem Predicting RNA degradation to improve thermostability.
method Crowdsourced machine learning competition on Kaggle.
result 41% of predictions matched experimental data, and models generalized to longer RNA molecules.
Many researches demonstrated that the DNA methylation, which occurs in the context of a CpG, has strong correlation with diseases, including cancer. There is a strong interest in analyzing the DNA methylation data to find how to distinguish different subtypes of the tumor. However, the conventional statistical methods …
MOD improves ensemble-based uncertainty estimates by encouraging larger diversity.
problem Improving model uncertainty estimates for inputs not seen during training.
method Maximize Overall Diversity (MOD) approach to encourage larger diversity in ensemble predictions.
result Significantly improves predictive performance for out-of-distribution test examples.
Generalized quandle polynomial used for stuquandles, stuck links, and RNA folding.
problem Defining polynomial invariants for stuquandles, stuck links, and RNA foldings.
method Introduced a generalized quandle polynomial and proved its invariance for stuquandles. Used this invariant to define polynomials for stuck links and RNA foldings.
result Polynomial invariants for stuquandles, stuck links, and RNA foldings.
This study reviews and evaluates clustering methods for single-cell RNA-seq data.
problem Identifying and characterizing novel cell types from single-cell RNA-seq data.
method Review and performance comparison of clustering methods.
result Performance comparison experiments on two datasets.
Predicting RNA base distances using a large language model.
problem Accurately predicting RNA structural information, especially distance maps.
method Using a large pretrained RNA language model coupled with a transformer.
result The model can accurately infer RNA base distances from sequence data.
Solving the RNA inverse folding problem is a critical prerequisite to RNA design, an emerging field in bioengineering with a broad range of applications from reaction catalysis to cancer therapy. Although significant progress has been made in developing machine-based inverse RNA folding algorithms, current approaches s…
This abstract reviews recent methods for predicting protein-ligand binding affinity.
problem Predicting protein-ligand binding affinity for various applications in life sciences.
method Traditional and deep learning models for binding affinity prediction.
result Improved predictive performance of AI-driven models.
PLIT identifies plant lncRNAs from RNA-seq data with high accuracy.
problem Inaccurate identification of lncRNAs in plant transcriptomic datasets.
method PLIT uses L1 regularization and iRF classification to select optimal features from sequence and codon-bias data.
result PLIT outperforms existing CPC tools in identifying lncRNAs in plant RNA-seq datasets.
A deep learning model organizes RNA graphs to reveal folding patterns and properties.
problem Organizing and understanding the complex folding patterns of RNA secondary structures.
method Geometric scattering autoencoder (GSAE) network for learning graph embeddings.
result GSAE accurately reflects bistable RNA structures and can sample new folding trajectories.
New benchmarks for RNA 3D structure-function modeling.
problem Lack of standardized benchmarks for RNA deep learning.
method Developed seven benchmark datasets, provided tools for data handling, and offered a user-friendly environment for model comparison.
result Demonstrated utility with baseline results using a relational graph neural network.
Study on binding numbers of tight contact structures on lens spaces L(n,1).
problem Determining the minimum number of binding components for tight contact structures on lens spaces.
method Using the d3-invariant, restrictions on planar monodromy factorizations, and the Durst-Kegel algorithm. result The binding number of universally tight contact structures on L(n,1) is equal to n. E2Efold predicts RNA secondary structures better than previous methods.
problem RNA secondary structure prediction with constraints.
method End-to-end deep learning model using unrolled algorithms to enforce constraints.
result E2Efold predicts significantly better structures, especially for pseudoknotted structures.
The Regularized Nonlinear Acceleration (RNA) algorithm is an acceleration method capable of improving the rate of convergence of many optimization schemes such as gradient descend, SAGA or SVRG. Until now, its analysis is limited to convex problems, but empirical observations shows that RNA may be extended to wider set…
New method detects RNA modifications without prior training, revealing novel sites.
problem Detecting RNA modifications with high accuracy and sensitivity.
method Anomaly detection using nanopore raw ionic current signals and nearest neighbor comparison.
result Detects diverse RNA modifications without prior training, including a novel 2'-O-methylated site in DENV.
New invariants for RNA foldings and stuck links defined.
problem Defining invariants for RNA foldings and stuck links.
method Assigning Boltzmann weights at classical and stuck crossings.
result Explicit computations of new invariants provided.
This research adapts superpixels for Shapley value computation in DNA profile classification.
problem Efficiently computing Shapley values for large, multidimensional time-series data.
method Adapting the concept of superpixels to streamline Shapley value computation for time-series-like data.
result Realistic, accurate, and fast computation of Shapley values for DNA profile classification.
Graph DNA uses Bloom filters to efficiently encode deep graph neighborhoods for better collaborative filtering.
problem Collaborative filtering struggles with exploiting deeper graph neighborhoods due to high time and space complexity.
method Graph DNA employs Bloom filters to compute approximate deep neighborhood information in linear time, enabling efficient encoding and utilization in collaborative filtering.
result Graph DNA significantly improves collaborative filtering performance with minimal computational and memory overhead.
We develop topological methods for analyzing difference topology experiments involving 3-string tangles. Difference topology is a novel technique used to unveil the structure of stable protein-DNA complexes involving two or more DNA segments. We analyze such experiments for the Mu protein-DNA complex. We characterize t…
Foundation models fail to preserve continuous geometry, identified as the Geometric Alignment Tax.
problem Continuous geometry is lost in foundation models due to discrete categorical bottlenecks.
method Controlled ablations on synthetic systems and evaluation of 14 biological models using rate-distortion theory and MINE.
result Replacing cross-entropy with a continuous head reduces geometric distortion by up to 8.5x.
New model improves DNA methylation data analysis.
problem Analyzing DNA methylation data with complex distributions.
method Doubly non-central beta (DNCB) distribution for non-negative matrix factorization.
result Improves predictive performance and yields meaningful latent representations.
The protein recombinase can change the knot type of circular DNA. The action of a recombinase converting one knot into another knot is normally mathematically modeled by band surgery. Band surgeries on a 2-bridge knot N((4mn-1)/(2m)) yielding a (2,2k)-torus link are characterized. We apply this and other rational tangl…
The study shows examples of contact 3-manifold binding sums that fail to preserve certain properties.
problem Examples of contact 3-manifold binding sums that fail to preserve properties like tightness or symplectic fillability.
method Examples and proofs of vanishing Heegaard Floer contact invariant for Stein fillable manifolds.
result Binding sums of contact 3-manifolds do not preserve properties such as tightness or symplectic fillability.
Study uses knot theory to model RNA foldings, emphasizing both entanglement and intrachain interactions.
problem Modeling RNA foldings considering both entanglement and intrachain interactions.
method Combines knot theory with embedded rigid vertex graphs to emphasize both entanglement and intrachain interactions of RNA foldings.
result Defines and computes a coloring counting invariant for stuck links, providing explicit computations for arc diagrams of RNA foldings.
DNAS disentangles neural architecture search for better interpretability and performance.
problem Lack of interpretability in existing neural architecture search methods.
method DNAS disentangles the hidden representation of the controller into semantically meaningful concepts.
result DNAS achieves state-of-the-art performance and competitive architectures.
New method shows links can be braided open book bindings.
problem Understanding fibered links and their bindings.
method Mutual arc presentations and braided open books.
result Every fibered link is the binding of a braided open book.
WideDTA predicts drug-target binding affinity using text-based information.
problem Predicting drug-target binding affinity is a major challenge in drug discovery.
method WideDTA uses chemical and biological textual sequence information, including protein sequence, ligand SMILES, protein domains and motifs, and maximum common substructure words.
result WideDTA outperformed DeepDTA on the KIBA dataset, indicating the word-based sequence representation is a promising alternative.
Graph ConvNet improves ncRNA classification accuracy.
problem Classifying non-coding RNA sequences into families.
method Graph Convolutional Network model trained on raw RNA graphs.
result 85.73% accuracy and 85.61% F1-score over 13 classes.