A faster method for optimizing DNA and protein sequences using machine learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper uses transfer learning and Bayesian optimization to reduce DNA sequence design experiments.
DNA Methylation has been the most extensively studied epigenetic mark. Usually a change in the genotype, DNA sequence, leads to a change in the phenotype, observable characteristics of the individual. But DNA methylation, which happens in the context of CpG (cytosine and guanine bases linked by phosphate backbone) dinu…
An evolutionary algorithm separates mixed DNA profiles in forensic genetics.
A novel framework refines diffusion models iteratively for better downstream reward optimization.
The folding structure of the DNA molecule combined with helper molecules, also referred to as the chromatin, is highly relevant for the functional properties of DNA. The chromatin structure is largely determined by the underlying primary DNA sequence, though the interaction is not yet fully understood. In this paper we…
Many researches demonstrated that the DNA methylation, which occurs in the context of a CpG, has strong correlation with diseases, including cancer. There is a strong interest in analyzing the DNA methylation data to find how to distinguish different subtypes of the tumor. However, the conventional statistical methods …
This research adapts superpixels for Shapley value computation in DNA profile classification.
We develop topological methods for analyzing difference topology experiments involving 3-string tangles. Difference topology is a novel technique used to unveil the structure of stable protein-DNA complexes involving two or more DNA segments. We analyze such experiments for the Mu protein-DNA complex. We characterize t…
Chirality affects the curvature of molecular networks, influencing their shape and stability.
We propose generative neural network methods to generate DNA sequences and tune them to have desired properties. We present three approaches: creating synthetic DNA sequences using a generative adversarial network; a DNA-based variant of the activation maximization ("deep dream") design method; and a joint procedure wh…
New model improves DNA methylation data analysis.
The protein recombinase can change the knot type of circular DNA. The action of a recombinase converting one knot into another knot is normally mathematically modeled by band surgery. Band surgeries on a 2-bridge knot N((4mn-1)/(2m)) yielding a (2,2k)-torus link are characterized. We apply this and other rational tangl…
DNA-SE uses deep learning to solve semiparametric problems efficiently.
DNAS disentangles neural architecture search for better interpretability and performance.
Genomic models learn DNA sequences to predict functions.
In this paper, we consider recommender systems with side information in the form of graphs. Existing collaborative filtering algorithms mainly utilize only immediate neighborhood information and have a hard time taking advantage of deeper neighborhoods beyond 1-2 hops. The main caveat of exploiting deeper graph informa…
This paper is an introduction to rational tangles, rational knots and links and their applications to DNA. The paper can be read as an introduction to our more technical papers on rational tangles (math.GT/0311499) and on rational knots (math.GT/0212011). The present paper includes a self-contained account of the tangl…
Study uses DNA methylation data to predict suicidal and non-suicidal deaths.
We study two systems of tangle equations that arise when modeling the action of the Integrase family of proteins on DNA. These two systems--direct and inverted repeats--correspond to two different possibilities for the initial DNA sequence. We present one new class of solutions to the tangle equations. In the case of i…
A deep probabilistic model analyzes DNA-encoded library data for efficient screening.
In this paper we propose network methodology to infer prognostic cancer biomarkers based on the epigenetic pattern DNA methylation. Epigenetic processes such as DNA methylation reflect environmental risk factors, and are increasingly recognised for their fundamental role in diseases such as cancer. DNA methylation is a…
Measures DNA quality degradation effects.
When analyzing the genome, researchers have discovered that proteins bind to DNA based on certain patterns of the DNA sequence known as "motifs". However, it is difficult to manually construct motifs due to their complexity. Recently, externally learned memory models have proven to be effective methods for reasoning ov…
We consider the task of detecting regulatory elements in the human genome directly from raw DNA. Past work has focused on small snippets of DNA, making it difficult to model long-distance dependencies that arise from DNA's 3-dimensional conformation. In order to study long-distance dependencies, we develop and release …
We consider learning parameters of Binomial Hidden Markov Models, which may be used to model DNA methylation data. The standard algorithm for the problem is EM, which is computationally expensive for sequences of the scale of the mammalian genome. Recently developed spectral algorithms can learn parameters of latent va…
With different genomes available, unsupervised learning algorithms are essential in learning genome-wide biological insights. Especially, the functional characterization of different genomes is essential for us to understand lives. In this book chapter, we review the state-of-the-art unsupervised learning algorithms fo…
Study on the structure of classifier boundaries in DNA sequencing.
Neural architecture search (NAS) aims to discover network architectures with desired properties such as high accuracy or low latency. Recently, differentiable NAS (DNAS) has demonstrated promising results while maintaining a search cost orders of magnitude lower than reinforcement learning (RL) based NAS. However, DNAS…
Study of Betti numbers in prodsimplicial complexes for directed graphs, focusing on DNA recombination.
We study optimal double helices with straight axes (or the fattest tubes around them) computationally using three kinds of functionals; ideal ones using ropelength, best volume packing ones, and energy minimizers using two one-parameter families of interaction energies between two strands of types and $\frac1r…
Integrase proteins acting on circular double-stranded DNA often change its topology by transforming unknotted circles into torus knots and links. Two systems of tangle equations--corresponding to the two initial DNA sequences--arise when modelling this transformation: direct and inverted. With no a priori assumptions o…
We categorise coherent band (aka nullification) pathways between knots and 2-component links. Additionally, we characterise the minimal coherent band pathways (with intermediates) between any two knots or 2-component links with small crossing number. We demonstrate these band surgeries for knots and links with small cr…
Algorithm optimizes biological sequences using bootstrapped training with a score-conditioned generator.
Over 150,000 new people in the United States are diagnosed with colorectal cancer each year. Nearly a third die from it (American Cancer Society). The only approved noninvasive diagnosis tools currently involve fecal blood count tests (FOBTs) or stool DNA tests. Fecal blood count tests take only five minutes and are av…
We propose a dynamic neighborhood aggregation (DNA) procedure guided by (multi-head) attention for representation learning on graphs. In contrast to current graph neural networks which follow a simple neighborhood aggregation scheme, our DNA procedure allows for a selective and node-adaptive aggregation of neighboring …
Graph Canonical Correlation Analysis improves CCA for multiomics datasets.
In this paper, we study a geometric/topological measure of knots and links called the nullification number. The nullification of knots/links is believed to be biologically relevant. For example, in DNA topology, one can intuitively regard it as a way to measure how easily a knotted circular DNA can unknot itself throug…
We introduce a method to learn a hierarchy of successively more abstract representations of complex data based on optimizing an information-theoretic objective. Intuitively, the optimization searches for a set of latent factors that best explain the correlations in the data as measured by multivariate mutual informatio…
Study surgeries between lens spaces using Heegaard Floer d-invariant.
The research reported in this paper identifies the epigenetic biomarker (methylation beta pattern) of breast cancer. Many cancers are triggered by abnormal gene expression levels caused by aberrant methylation of CpG sites in the DNA. In order to develop early diagnostics of cancer-causing methylations and to develop a…
The paper studies pseudo links in genus g handlebodies, generalizing knot theory.
Novel U-learning method for predicting continuous outcomes from high-dimensional data.
Over the last years, huge resources of biological and medical data have become available for research. This data offers great chances for machine learning applications in health care, e.g. for precision medicine, but is also challenging to analyze. Typical challenges include a large number of possibly correlated featur…
We develop an algorithm for minimizing a function using batched function value measurements at each of rounds by using classifiers to identify a function's sublevel set. We show that sufficiently accurate classifiers can achieve linear convergence rates, and show that the convergence rate is tied to the difficu…
Avian Influenza breakouts cause millions of dollars in damage each year globally, especially in Asian countries such as China and South Korea. The impact magnitude of a breakout directly correlates to time required to fully understand the influenza virus, particularly the interspecies pathogenicity. The procedure requi…
A natural generalization of a crossing change is a rational subtangle replacement (RSR). We characterize the fundamental situation of the rational tangles obtained from a given rational tangle via RSR, building on work of Berge and Gabai, and determine the sites where these RSR may occur. In addition we also determine …
New approach speeds up DNA sequence alignment.