Convolutional network predicts DNA chromatin structure from sequence images.
problem Predicting chromatin structure from DNA sequences.
method Developed a convolutional neural network using image-representation of DNA sequences.
result The method outperforms existing methods in prediction accuracy and training time.
CNNs classify cancer types based on DNA methylation patterns.
problem Classifying cancer types using DNA methylation data.
method Convolutional Neural Networks (CNNs) for classification.
result CNNs can classify cancer types from DNA methylation profiles.
Deep neural network improves DNA methylation data analysis.
problem Analyzing highly dimensional DNA methylation data with bounded support.
method Designing a deep neural network composed of stacked binary restricted Boltzmann machines.
result Deep features learned by the neural network perform best in cluster analysis of breast cancer DNA methylation data.
Deep models generate and optimize DNA sequences for protein binding.
problem Designing DNA sequences with desired properties.
method Three approaches: GAN for synthetic sequences, activation maximization for design, and a combined method.
result Generated DNA sequences have superior properties to those in training data.
Memory Matching Networks classify DNA sequences for protein binding sites.
problem Manual construction of DNA motifs is difficult due to their complexity.
method Memory Matching Networks (MMN) learn a dynamic memory bank of encoded motifs and match them to new sequences.
result MMN effectively classifies DNA sequences as protein binding or nonbinding sites.
In this paper we propose network methodology to infer prognostic cancer biomarkers based on the epigenetic pattern DNA methylation. Epigenetic processes such as DNA methylation reflect environmental risk factors, and are increasingly recognised for their fundamental role in diseases such as cancer. DNA methylation is a…
This research adapts superpixels for Shapley value computation in DNA profile classification.
problem Efficiently computing Shapley values for large, multidimensional time-series data.
method Adapting the concept of superpixels to streamline Shapley value computation for time-series-like data.
result Realistic, accurate, and fast computation of Shapley values for DNA profile classification.
DNA improves graph neural networks by selectively aggregating node embeddings.
problem Static neighborhood aggregation limits graph neural networks' performance.
method Dynamic neighborhood aggregation guided by attention and controlled channel connections.
result DNA outperforms current methods in transductive node classification.
Dilated convolutions model long-distance genomic dependencies effectively.
problem Detecting regulatory elements from raw DNA with long-distance dependencies.
method Developed and used a novel dataset for dilated convolutional neural networks.
result Dilated convolutions are effective at modeling regulatory elements in the human genome.
DeepRAM evaluates and selects the best deep learning architecture for DNA/RNA binding specificity prediction.
problem Selecting the best deep learning architecture for predicting DNA/RNA binding specificity.
method Systematic exploration of various deep learning architectures using deepRAM, an end-to-end deep learning tool.
result A k-mer embedding convolutional layer and recurrent layer architecture outperforms other methods.
DNAS disentangles neural architecture search for better interpretability and performance.
problem Lack of interpretability in existing neural architecture search methods.
method DNAS disentangles the hidden representation of the controller into semantically meaningful concepts.
result DNAS achieves state-of-the-art performance and competitive architectures.
New method combines personal and reference genomes for better machine learning in DNA sequencing.
problem Improving accuracy of genetic variant calls in sequencing data.
method Interlaces personal and reference genomes to generate images for machine learning.
result Significant improvement in germline variant calling and somatic variant calling across tumor/normal data.
Chirality affects the curvature of molecular networks, influencing their shape and stability.
problem Understanding how chirality influences the curvature of molecular networks.
method Langevin dynamics simulations and constrained gradient optimization of square lattice networks.
result Linking chirality dictates the sign of Gaussian curvature in molecular chainmail networks.
An evolutionary algorithm separates mixed DNA profiles in forensic genetics.
problem Deconvolving mixed DNA profiles from crime samples.
method Multiple population evolutionary algorithm (MEA) with guided mutation.
result The MEA successfully deconvoluted DNA profiles from crime samples.
Study on stability of distributed filtering over DNA networks for Gaussian-HMMs.
problem Stability of distributed nonlinear filtering over DNA networks in partially observed Gaussian noise.
method Distributed evaluation of likelihood using ADMM for consensus, with constraints on consensus steps.
result Uniform ε-stability depends loglinearly on T and network size, independent of HMM structure.
With different genomes available, unsupervised learning algorithms are essential in learning genome-wide biological insights. Especially, the functional characterization of different genomes is essential for us to understand lives. In this book chapter, we review the state-of-the-art unsupervised learning algorithms fo…
Graph DNA uses Bloom filters to efficiently encode deep graph neighborhoods for better collaborative filtering.
problem Collaborative filtering struggles with exploiting deeper graph neighborhoods due to high time and space complexity.
method Graph DNA employs Bloom filters to compute approximate deep neighborhood information in linear time, enabling efficient encoding and utilization in collaborative filtering.
result Graph DNA significantly improves collaborative filtering performance with minimal computational and memory overhead.
We develop topological methods for analyzing difference topology experiments involving 3-string tangles. Difference topology is a novel technique used to unveil the structure of stable protein-DNA complexes involving two or more DNA segments. We analyze such experiments for the Mu protein-DNA complex. We characterize t…
New spectral method learns DNA methylation models efficiently.
problem Learning parameters of Binomial HMMs for DNA methylation data.
method Feature-map based approach exploiting Binomial HMM properties.
result The new algorithm provides theoretical guarantees and performs well on real data.
New model improves DNA methylation data analysis.
problem Analyzing DNA methylation data with complex distributions.
method Doubly non-central beta (DNCB) distribution for non-negative matrix factorization.
result Improves predictive performance and yields meaningful latent representations.
The protein recombinase can change the knot type of circular DNA. The action of a recombinase converting one knot into another knot is normally mathematically modeled by band surgery. Band surgeries on a 2-bridge knot N((4mn-1)/(2m)) yielding a (2,2k)-torus link are characterized. We apply this and other rational tangl…
Novel U-learning method for predicting continuous outcomes from high-dimensional data.
problem Challenges in making valid inferences on predictions from high-dimensional inputs.
method U-learning via combinatory multi-subsampling for ensemble predictions and confidence intervals.
result Valid inferences on predictions from Lasso and neural networks.
New method embeds DNA sequences for faster, more informative gene comparison.
problem Slow and costly sequence comparison methods for genes without exact matches.
method Recurrent neural networks to embed sequences in a low-dimensional space.
result Embedding allows for better comparison of genes without exact matches.
UNAS combines DNAS and RL for efficient architecture search.
problem Discovering high accuracy or low latency neural architectures.
method Unified framework combining differentiable and reinforcement learning approaches.
result UNAS achieves state-of-the-art accuracy on CIFAR-10, CIFAR-100, and ImageNet datasets.
Genomic models learn DNA sequences to predict functions.
problem Understanding complex genetic interactions.
method Training LLMs on DNA sequences to predict functions.
result gLMs can predict functions of DNA elements.
Study uses DNA methylation data to predict suicidal and non-suicidal deaths.
problem Predicting suicidal and non-suicidal deaths from DNA methylation data.
method Support Vector Machines (SVM) with dimensionality reduction using PCA and t-SNE.
result t-SNE outperforms PCA in reducing data dimensionality.
This paper is an introduction to rational tangles, rational knots and links and their applications to DNA. The paper can be read as an introduction to our more technical papers on rational tangles (math.GT/0311499) and on rational knots (math.GT/0212011). The present paper includes a self-contained account of the tangl…
AI4AI uses machine learning to classify avian influenza host species from DNA sequences.
problem Classifying avian influenza host species from DNA sequences to reduce emergency response time.
method Quantitative methods using machine learning and deep learning.
result Best deep learning models achieve top-1 classification accuracy of 47%, and top-3 classification accuracy of 82%.
We study two systems of tangle equations that arise when modeling the action of the Integrase family of proteins on DNA. These two systems--direct and inverted repeats--correspond to two different possibilities for the initial DNA sequence. We present one new class of solutions to the tangle equations. In the case of i…
A faster method for optimizing DNA and protein sequences using machine learning.
problem Designing DNA and protein sequences with improved function.
method Activation maximization with a straight-through approximation and adaptive entropy variable.
result Fast SeqProp achieves up to 100-fold faster convergence and improved fitness optima.
A deep probabilistic model analyzes DNA-encoded library data for efficient screening.
problem Complex data from DNA-encoded library experiments mask underlying signals.
method Compositional deep probabilistic model of DEL data, modeling latent reactions between synthons.
result DEL-Compose model demonstrates strong performance and valuable insights.
Measures DNA quality degradation effects.
problem Identifying degraded DNA sequence data.
method Novel quality quantification based on intentional degradation effects.
result Quantified measures of degradation can be used for multiple purposes.
dna2vec creates consistent vectors from DNA sequences, addressing sequence analysis challenges.
problem Inequivalent distances between one-hot vectors of k-mers and limitations of machine learning on long DNA sequences.
method Proposes a neural network-based approach to train distributed representations of variable-length k-mers.
result Summing dna2vec vectors is equivalent to nucleotide concatenation and correlates with sequence similarity.
Adaptive model improves age prediction from DNA methylation data.
problem Improving age prediction from DNA methylation data with heterogeneity.
method Adaptive model for feature selection based on individual sample distributions.
result Substantial improvement in age prediction for one tissue type.
Paper uses transfer learning and Bayesian optimization to reduce DNA sequence design experiments.
problem Designing many similar DNA sequences for specific applications is expensive and time-consuming.
method Combines transfer learning with Bayesian optimization to reduce experiment count.
result Total number of experiments can be significantly reduced by sharing information between tasks.
Study on the structure of classifier boundaries in DNA sequencing.
problem Understanding the structure of boundaries in a Bayes classifier for DNA sequencing.
method Examined the structure of the boundary in a Bayes classifier applied to DNA sequencing data. Introduced a new measure of uncertainty, Neighbor Similarity.
result The boundary is large and complex, and Neighbor Similarity effectively measures classifier uncertainty.
Study of Betti numbers in prodsimplicial complexes for directed graphs, focusing on DNA recombination.
problem Analyzing Betti numbers in directed graphs for DNA recombination.
method Custom prodsimplicial complexes for acyclic directed graphs, investigating Betti numbers.
result Investigated Betti numbers and cycles in prodsimplicial complexes for DNA recombination.
Deep autoencoder predicts cancer types from DNA methylation patterns.
problem Differentiating cancer types based on DNA methylation states.
method Deep learning system with CpG island state classification and statistical methods.
result Overall Sensitivity of 88.24%, Specificity of 83.33%, Accuracy of 84.75%.
Introduces BWMD, a new distance measure for DNA and malware clustering.
problem Shortcomings of previous compression-based distance metrics.
method Embeds sequences into a fixed-length feature vector.
result Significantly improved clustering performance on larger malware corpora.
Integrase proteins acting on circular double-stranded DNA often change its topology by transforming unknotted circles into torus knots and links. Two systems of tangle equations--corresponding to the two initial DNA sequences--arise when modelling this transformation: direct and inverted. With no a priori assumptions o…
We categorise coherent band (aka nullification) pathways between knots and 2-component links. Additionally, we characterise the minimal coherent band pathways (with intermediates) between any two knots or 2-component links with small crossing number. We demonstrate these band surgeries for knots and links with small cr…
Deep learning improves tumor type classification accuracy.
problem Classifying cancer types based on DNA mutations is challenging.
method Deep transfer learning and fine-tuning of gene expression data.
result Significantly improved tumor type classification accuracy (78.3%) using DNA point mutations.
A novel framework refines diffusion models iteratively for better downstream reward optimization.
problem Optimizing reward functions during inference of diffusion models.
method Iterative refinement process with noising and reward-guided denoising steps.
result Superior empirical performance in protein and DNA design.
Kernel and MKCCA classify schizophrenia patients from imaging and genetic data.
problem Classifying schizophrenia patients from imaging and genetic data.
method Employed Kernel and Multiple Kernel Canonical Correlation Analysis (CCA) for classification.
result Kernel and Multiple Kernel CCA significantly outperform regularized linear CCA in classification accuracy.
Graph Canonical Correlation Analysis improves CCA for multiomics datasets.
problem Limited ability of conventional CCA methods to incorporate structured patterns in cross-correlation matrices.
method Graph Canonical Correlation Analysis (gCCA) calculates canonical correlations based on the graph structure of cross-correlation matrices.
result gCCA outperforms competing CCA methods in simulations and multiomics dataset analysis.
In this paper, we study a geometric/topological measure of knots and links called the nullification number. The nullification of knots/links is believed to be biologically relevant. For example, in DNA topology, one can intuitively regard it as a way to measure how easily a knotted circular DNA can unknot itself throug…
Automated image segmentation distinguishes overlapping human chromosomes.
problem Distinguishing overlapping human chromosomes for medical diagnostics.
method Customized convolutional neural network for image segmentation.
result IOU scores of 94.7% for overlapping regions, 88-94% for non-overlapping regions.
A cheap CRC screening method combining FIT, BMI, smoking, and diabetes.
problem Inadequate sensitivity and specificity of current CRC screening methods.
method Ensemble-based classification algorithm considering FIT, BMI, smoking, and diabetes.
result 92% specificity, 89% sensitivity, 90% precision, AUC of .95.