Introduces BWMD, a new distance measure for DNA and malware clustering.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Many researches demonstrated that the DNA methylation, which occurs in the context of a CpG, has strong correlation with diseases, including cancer. There is a strong interest in analyzing the DNA methylation data to find how to distinguish different subtypes of the tumor. However, the conventional statistical methods …
This research adapts superpixels for Shapley value computation in DNA profile classification.
Proposes a copula-based model for multi-view clustering with directional dependency.
Techniques for data-mining, latent semantic analysis, contextual search of databases, etc. have long ago been developed by computer scientists working on information retrieval (IR). Experimental scientists, from all disciplines, having to analyse large collections of raw experimental data (astronomical, physical, biolo…
DNA Methylation has been the most extensively studied epigenetic mark. Usually a change in the genotype, DNA sequence, leads to a change in the phenotype, observable characteristics of the individual. But DNA methylation, which happens in the context of CpG (cytosine and guanine bases linked by phosphate backbone) dinu…
An evolutionary algorithm separates mixed DNA profiles in forensic genetics.
The folding structure of the DNA molecule combined with helper molecules, also referred to as the chromatin, is highly relevant for the functional properties of DNA. The chromatin structure is largely determined by the underlying primary DNA sequence, though the interaction is not yet fully understood. In this paper we…
Subspace clustering aims to find groups of similar objects (clusters) that exist in lower dimensional subspaces from a high dimensional dataset. It has a wide range of applications, such as analysing high dimensional sensor data or DNA sequences. However, existing algorithms have limitations in finding clusters in non-…
Kernel and Multiple Kernel Canonical Correlation Analysis (CCA) are employed to classify schizophrenic and healthy patients based on their SNPs, DNA Methylation and fMRI data. Kernel and Multiple Kernel CCA are popular methods for finding nonlinear correlations between high-dimensional datasets. Data was gathered from …
We develop topological methods for analyzing difference topology experiments involving 3-string tangles. Difference topology is a novel technique used to unveil the structure of stable protein-DNA complexes involving two or more DNA segments. We analyze such experiments for the Mu protein-DNA complex. We characterize t…
We propose generative neural network methods to generate DNA sequences and tune them to have desired properties. We present three approaches: creating synthetic DNA sequences using a generative adversarial network; a DNA-based variant of the activation maximization ("deep dream") design method; and a joint procedure wh…
New model improves DNA methylation data analysis.
The protein recombinase can change the knot type of circular DNA. The action of a recombinase converting one knot into another knot is normally mathematically modeled by band surgery. Band surgeries on a 2-bridge knot N((4mn-1)/(2m)) yielding a (2,2k)-torus link are characterized. We apply this and other rational tangl…
DNAS disentangles neural architecture search for better interpretability and performance.
Gene annotation has traditionally required direct comparison of DNA sequences between an unknown gene and a database of known ones using string comparison methods. However, these methods do not provide useful information when a gene does not have a close match in the database. In addition, each comparison can be costly…
Genomic models learn DNA sequences to predict functions.
In this paper, we consider recommender systems with side information in the form of graphs. Existing collaborative filtering algorithms mainly utilize only immediate neighborhood information and have a hard time taking advantage of deeper neighborhoods beyond 1-2 hops. The main caveat of exploiting deeper graph informa…
This paper is an introduction to rational tangles, rational knots and links and their applications to DNA. The paper can be read as an introduction to our more technical papers on rational tangles (math.GT/0311499) and on rational knots (math.GT/0212011). The present paper includes a self-contained account of the tangl…
Study uses DNA methylation data to predict suicidal and non-suicidal deaths.
We study two systems of tangle equations that arise when modeling the action of the Integrase family of proteins on DNA. These two systems--direct and inverted repeats--correspond to two different possibilities for the initial DNA sequence. We present one new class of solutions to the tangle equations. In the case of i…
A faster method for optimizing DNA and protein sequences using machine learning.
A deep probabilistic model analyzes DNA-encoded library data for efficient screening.
In this paper we propose network methodology to infer prognostic cancer biomarkers based on the epigenetic pattern DNA methylation. Epigenetic processes such as DNA methylation reflect environmental risk factors, and are increasingly recognised for their fundamental role in diseases such as cancer. DNA methylation is a…
Deep generative model for healthcare data identifies coherent substructures and mutational clusters.
Measures DNA quality degradation effects.
When analyzing the genome, researchers have discovered that proteins bind to DNA based on certain patterns of the DNA sequence known as "motifs". However, it is difficult to manually construct motifs due to their complexity. Recently, externally learned memory models have proven to be effective methods for reasoning ov…
Paper uses transfer learning and Bayesian optimization to reduce DNA sequence design experiments.
Genie clusters faster and resists outliers.
We consider the task of detecting regulatory elements in the human genome directly from raw DNA. Past work has focused on small snippets of DNA, making it difficult to model long-distance dependencies that arise from DNA's 3-dimensional conformation. In order to study long-distance dependencies, we develop and release …
We consider learning parameters of Binomial Hidden Markov Models, which may be used to model DNA methylation data. The standard algorithm for the problem is EM, which is computationally expensive for sequences of the scale of the mammalian genome. Recently developed spectral algorithms can learn parameters of latent va…
With different genomes available, unsupervised learning algorithms are essential in learning genome-wide biological insights. Especially, the functional characterization of different genomes is essential for us to understand lives. In this book chapter, we review the state-of-the-art unsupervised learning algorithms fo…
Study on the structure of classifier boundaries in DNA sequencing.
In many real-world problems, we are dealing with collections of high-dimensional data, such as images, videos, text and web documents, DNA microarray data, and more. Often, high-dimensional data lie close to low-dimensional structures corresponding to several classes or categories the data belongs to. In this paper, we…
Study of Betti numbers in prodsimplicial complexes for directed graphs, focusing on DNA recombination.
Method cleans covariance matrices for better statistical inference.
Integrase proteins acting on circular double-stranded DNA often change its topology by transforming unknotted circles into torus knots and links. Two systems of tangle equations--corresponding to the two initial DNA sequences--arise when modelling this transformation: direct and inverted. With no a priori assumptions o…
We categorise coherent band (aka nullification) pathways between knots and 2-component links. Additionally, we characterise the minimal coherent band pathways (with intermediates) between any two knots or 2-component links with small crossing number. We demonstrate these band surgeries for knots and links with small cr…
In recent years, ideas from statistics and scientific computing have begun to interact in increasingly sophisticated and fruitful ways with ideas from computer science and the theory of algorithms to aid in the development of improved worst-case algorithms that are useful for large-scale scientific and Internet data an…
GIDS reduces high-dimensional response and predictor spaces, improving interpretability and computational efficiency.
Analysis of somatic mutation profiles from cancer patients is essential in the development of cancer research. However, the low frequency of most mutations and the varying rates of mutations across patients makes the data extremely challenging to statistically analyze as well as difficult to use in classification probl…
We propose a dynamic neighborhood aggregation (DNA) procedure guided by (multi-head) attention for representation learning on graphs. In contrast to current graph neural networks which follow a simple neighborhood aggregation scheme, our DNA procedure allows for a selective and node-adaptive aggregation of neighboring …
UNAS combines DNAS and RL for efficient architecture search.
A novel framework refines diffusion models iteratively for better downstream reward optimization.
Motivation: In this paper we present the latest release of EBIC, a next-generation biclustering algorithm for mining genetic data. The major contribution of this paper is adding support for big data, making it possible to efficiently run large genomic data mining analyses. Additional enhancements include integration wi…
Graph Canonical Correlation Analysis improves CCA for multiomics datasets.
Chirality affects the curvature of molecular networks, influencing their shape and stability.
In this paper, we study a geometric/topological measure of knots and links called the nullification number. The nullification of knots/links is believed to be biologically relevant. For example, in DNA topology, one can intuitively regard it as a way to measure how easily a knotted circular DNA can unknot itself throug…