Convolutional network predicts DNA chromatin structure from sequence images.
problem Predicting chromatin structure from DNA sequences.
method Developed a convolutional neural network using image-representation of DNA sequences.
result The method outperforms existing methods in prediction accuracy and training time.
Memory Matching Networks classify DNA sequences for protein binding sites.
problem Manual construction of DNA motifs is difficult due to their complexity.
method Memory Matching Networks (MMN) learn a dynamic memory bank of encoded motifs and match them to new sequences.
result MMN effectively classifies DNA sequences as protein binding or nonbinding sites.
Deep models generate and optimize DNA sequences for protein binding.
problem Designing DNA sequences with desired properties.
method Three approaches: GAN for synthetic sequences, activation maximization for design, and a combined method.
result Generated DNA sequences have superior properties to those in training data.
Genomic models learn DNA sequences to predict functions.
problem Understanding complex genetic interactions.
method Training LLMs on DNA sequences to predict functions.
result gLMs can predict functions of DNA elements.
CNNs classify cancer types based on DNA methylation patterns.
problem Classifying cancer types using DNA methylation data.
method Convolutional Neural Networks (CNNs) for classification.
result CNNs can classify cancer types from DNA methylation profiles.
A faster method for optimizing DNA and protein sequences using machine learning.
problem Designing DNA and protein sequences with improved function.
method Activation maximization with a straight-through approximation and adaptive entropy variable.
result Fast SeqProp achieves up to 100-fold faster convergence and improved fitness optima.
New method combines personal and reference genomes for better machine learning in DNA sequencing.
problem Improving accuracy of genetic variant calls in sequencing data.
method Interlaces personal and reference genomes to generate images for machine learning.
result Significant improvement in germline variant calling and somatic variant calling across tumor/normal data.
AI4AI uses machine learning to classify avian influenza host species from DNA sequences.
problem Classifying avian influenza host species from DNA sequences to reduce emergency response time.
method Quantitative methods using machine learning and deep learning.
result Best deep learning models achieve top-1 classification accuracy of 47%, and top-3 classification accuracy of 82%.
Paper uses transfer learning and Bayesian optimization to reduce DNA sequence design experiments.
problem Designing many similar DNA sequences for specific applications is expensive and time-consuming.
method Combines transfer learning with Bayesian optimization to reduce experiment count.
result Total number of experiments can be significantly reduced by sharing information between tasks.
Study on the structure of classifier boundaries in DNA sequencing.
problem Understanding the structure of boundaries in a Bayes classifier for DNA sequencing.
method Examined the structure of the boundary in a Bayes classifier applied to DNA sequencing data. Introduced a new measure of uncertainty, Neighbor Similarity.
result The boundary is large and complex, and Neighbor Similarity effectively measures classifier uncertainty.
Measures DNA quality degradation effects.
problem Identifying degraded DNA sequence data.
method Novel quality quantification based on intentional degradation effects.
result Quantified measures of degradation can be used for multiple purposes.
New method embeds DNA sequences for faster, more informative gene comparison.
problem Slow and costly sequence comparison methods for genes without exact matches.
method Recurrent neural networks to embed sequences in a low-dimensional space.
result Embedding allows for better comparison of genes without exact matches.
New spectral method learns DNA methylation models efficiently.
problem Learning parameters of Binomial HMMs for DNA methylation data.
method Feature-map based approach exploiting Binomial HMM properties.
result The new algorithm provides theoretical guarantees and performs well on real data.
We study two systems of tangle equations that arise when modeling the action of the Integrase family of proteins on DNA. These two systems--direct and inverted repeats--correspond to two different possibilities for the initial DNA sequence. We present one new class of solutions to the tangle equations. In the case of i…
Introduces BWMD, a new distance measure for DNA and malware clustering.
problem Shortcomings of previous compression-based distance metrics.
method Embeds sequences into a fixed-length feature vector.
result Significantly improved clustering performance on larger malware corpora.
DeepRAM evaluates and selects the best deep learning architecture for DNA/RNA binding specificity prediction.
problem Selecting the best deep learning architecture for predicting DNA/RNA binding specificity.
method Systematic exploration of various deep learning architectures using deepRAM, an end-to-end deep learning tool.
result A k-mer embedding convolutional layer and recurrent layer architecture outperforms other methods.
New approach speeds up DNA sequence alignment.
problem Efficiently estimating alignment scores for large sets of reads.
method Rank-one crowdsourcing models and multi-armed bandit algorithm.
result Adaptive algorithm identifies pairs with large alignment scores.
dna2vec creates consistent vectors from DNA sequences, addressing sequence analysis challenges.
problem Inequivalent distances between one-hot vectors of k-mers and limitations of machine learning on long DNA sequences.
method Proposes a neural network-based approach to train distributed representations of variable-length k-mers.
result Summing dna2vec vectors is equivalent to nucleotide concatenation and correlates with sequence similarity.
Algorithm optimizes biological sequences using bootstrapped training with a score-conditioned generator.
problem Optimizing biological sequences for a black-box score function.
method Bootstrapped training of score-conditioned generator (BootGen) algorithm.
result Our method outperforms competitive baselines on biological sequential design tasks.
Paper solves NP-hard haplotyping problem using matrix completion.
problem Reconstructing inherited genetic variations from DNA sequencing data.
method Binary matrix factorization and alternating minimization.
result The proposed technique achieves lower haplotype reconstruction error.
Integrase proteins acting on circular double-stranded DNA often change its topology by transforming unknotted circles into torus knots and links. Two systems of tangle equations--corresponding to the two initial DNA sequences--arise when modelling this transformation: direct and inverted. With no a priori assumptions o…
With different genomes available, unsupervised learning algorithms are essential in learning genome-wide biological insights. Especially, the functional characterization of different genomes is essential for us to understand lives. In this book chapter, we review the state-of-the-art unsupervised learning algorithms fo…
Method computes embeddings for RNA-seq data without genome alignment.
problem No need for genome alignment for RNA-seq data analysis.
method RNN transforms kmers into 2D latent space for transcriptomic analysis.
result Captures DNA sequence similarity and abundance in latent space.
An evolutionary algorithm separates mixed DNA profiles in forensic genetics.
problem Deconvolving mixed DNA profiles from crime samples.
method Multiple population evolutionary algorithm (MEA) with guided mutation.
result The MEA successfully deconvoluted DNA profiles from crime samples.
Deep neural network improves DNA methylation data analysis.
problem Analyzing highly dimensional DNA methylation data with bounded support.
method Designing a deep neural network composed of stacked binary restricted Boltzmann machines.
result Deep features learned by the neural network perform best in cluster analysis of breast cancer DNA methylation data.
Metagenomics characterizes the taxonomic diversity of microbial communities by sequencing DNA directly from an environmental sample. One of the main challenges in metagenomics data analysis is the binning step, where each sequenced read is assigned to a taxonomic clade. Due to the large volume of metagenomics datasets,…
Robust machine learning models improve DNA regulatory sequence prediction under various shifts.
problem Real-world applications of DNA regulatory sequence prediction involve shifts not captured by standard i.i.d. assumptions.
method Introduces a robustness framework combining simulation benchmarks and real data analysis.
result Models remain accurate and calibrated under mild shifts but show higher error and miscalibration under strong shifts.
Many practical modeling problems involve discrete data that are best represented as draws from multinomial or categorical distributions. For example, nucleotides in a DNA sequence, children's names in a given state and year, and text documents are all commonly modeled with multinomial distributions. In all of these cas…
A new framework scales active search for large datasets.
problem Scaling active search for large, high-dimensional data sets.
method Hierarchical Batch Bandit Search (HBBS) framework.
result HBBS improves performance and scalability for batch search.
GeNet classifies metagenomic sequences with less memory and comparable recall to state-of-the-art methods.
problem Classifying metagenomic sequences from raw DNA sequences.
method Exploits hierarchical structure between labels for training, using deep representations.
result GeNet achieves competitive precision and good recall with less memory requirements.
This research adapts superpixels for Shapley value computation in DNA profile classification.
problem Efficiently computing Shapley values for large, multidimensional time-series data.
method Adapting the concept of superpixels to streamline Shapley value computation for time-series-like data.
result Realistic, accurate, and fast computation of Shapley values for DNA profile classification.
Graph DNA uses Bloom filters to efficiently encode deep graph neighborhoods for better collaborative filtering.
problem Collaborative filtering struggles with exploiting deeper graph neighborhoods due to high time and space complexity.
method Graph DNA employs Bloom filters to compute approximate deep neighborhood information in linear time, enabling efficient encoding and utilization in collaborative filtering.
result Graph DNA significantly improves collaborative filtering performance with minimal computational and memory overhead.
We develop topological methods for analyzing difference topology experiments involving 3-string tangles. Difference topology is a novel technique used to unveil the structure of stable protein-DNA complexes involving two or more DNA segments. We analyze such experiments for the Mu protein-DNA complex. We characterize t…
New model improves DNA methylation data analysis.
problem Analyzing DNA methylation data with complex distributions.
method Doubly non-central beta (DNCB) distribution for non-negative matrix factorization.
result Improves predictive performance and yields meaningful latent representations.
The protein recombinase can change the knot type of circular DNA. The action of a recombinase converting one knot into another knot is normally mathematically modeled by band surgery. Band surgeries on a 2-bridge knot N((4mn-1)/(2m)) yielding a (2,2k)-torus link are characterized. We apply this and other rational tangl…
DNAS disentangles neural architecture search for better interpretability and performance.
problem Lack of interpretability in existing neural architecture search methods.
method DNAS disentangles the hidden representation of the controller into semantically meaningful concepts.
result DNAS achieves state-of-the-art performance and competitive architectures.
A new method selects optimal PHMM models for sequence alignment, improving accuracy.
problem Improving sequence alignment accuracy using PHMMs with optimal hidden states.
method Factorized Asymptotic Bayesian algorithm (FIC) for model selection.
result Improved alignment accuracy with more complex models than previous studies.
META2 improves taxonomic classification and abundance estimation in metagenomics with deep learning and memory efficiency.
problem Memory constraints and inefficiencies in taxonomic classification and abundance estimation for metagenomics.
method Developed a novel memory-efficient read classification technique combining deep learning and locality-sensitive hashing, and formulated abundance estimation as a Multiple Instance Learning problem.
result Our approach outperforms conventional methods in both single-read taxonomic classification and abundance estimation, especially when memory is limited.
New algorithm estimates intrinsic dimension of discrete datasets.
problem Inaccuracies in using continuous methods for discrete datasets.
method Introduced an algorithm to infer intrinsic dimension of discrete spaces.
result Demonstrated accuracy on benchmark datasets and found a small intrinsic dimension in a metagenomic dataset.
Site-specific recombination on supercoiled circular DNA molecules can yield a variety of knots and catenanes. Twist knots are some of the most common conformations of these products and they can act as substrates for further rounds of site-specific recombination. They are also one of the simplest families of knots and …
New method optimizes diffusion models without fine-tuning, integrating soft value functions.
problem Optimizing natural design spaces of images, molecules, DNA, RNA, and protein sequences.
method Iterative sampling method integrating soft value functions into diffusion model inference.
result Directly utilizes non-differentiable features/reward feedback, applies to discrete diffusion models.
Study uses DNA methylation data to predict suicidal and non-suicidal deaths.
problem Predicting suicidal and non-suicidal deaths from DNA methylation data.
method Support Vector Machines (SVM) with dimensionality reduction using PCA and t-SNE.
result t-SNE outperforms PCA in reducing data dimensionality.
This paper is an introduction to rational tangles, rational knots and links and their applications to DNA. The paper can be read as an introduction to our more technical papers on rational tangles (math.GT/0311499) and on rational knots (math.GT/0212011). The present paper includes a self-contained account of the tangl…
Graphs represent gene segment organization, revealing complex interrelationships in a scrambled genome.
problem Understanding gene segment organization and interrelationships in a scrambled genome.
method Directed graphs representing gene segments and their relationships, with graph properties mapped to higher-dimensional space for analysis.
result Emerging star-like structures indicate complex interrelationships, including segments from multiple genes interleaving or overlapping.
A deep probabilistic model analyzes DNA-encoded library data for efficient screening.
problem Complex data from DNA-encoded library experiments mask underlying signals.
method Compositional deep probabilistic model of DEL data, modeling latent reactions between synthons.
result DEL-Compose model demonstrates strong performance and valuable insights.
In this paper we propose network methodology to infer prognostic cancer biomarkers based on the epigenetic pattern DNA methylation. Epigenetic processes such as DNA methylation reflect environmental risk factors, and are increasingly recognised for their fundamental role in diseases such as cancer. DNA methylation is a…
The paper introduces a new method to infer phylogenetic trees without bifurcations.
problem Inferring phylogenetic trees with zero-length branches and polytomies.
method Adaptive LASSO-type regularization estimators for phylogenetics.
result Regularization is a practical approach for phylogenetics, revealing zero-length branches.
Develops geometric causal models for causal inference from dependent data.
problem Causal inference from structured, dependent data (e.g., spatial, network, molecular).
method Geometric causal models (GCMs) exploiting symmetries of data generating process, combining group theory, ergodic theory, and Bayesian inference.
result Establishes identification and estimation of causal effects from dependent data.