New statistical model improves protein alignment accuracy.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new method estimates protein evolutionary fields and couplings from alignments.
Algorithm aligns 3D density maps using Wasserstein distance.
Bayesian HMM for protein alignment state estimation.
Rapid progress in deep learning has spurred its application to bioinformatics problems including protein structure prediction and design. In classic machine learning problems like computer vision, progress has been driven by standardized data sets that facilitate fair assessment of new methods and lower the barrier to …
Inferring the structural properties of a protein from its amino acid sequence is a challenging yet important problem in biology. Structures are not known for the vast majority of protein sequences, but structure is critical for understanding function. Existing approaches for detecting structural similarity between prot…
A new framework uses text descriptions to improve protein design.
A new method for aligning datasets without known correspondences.
The application of machine learning to bioinformatics problems is well established. Less well understood is the application of bioinformatics techniques to machine learning and, in particular, the representation of non-biological data as biosequences. The aim of this paper is to explore the effects of giving amino acid…
Proteins are the major building blocks of life, and actuators of almost all chemical and biophysical events in living organisms. Their native structures in turn enable their biological functions which have a fundamental role in drug design. This motivates predicting the structure of a protein from its sequence of amino…
As high-throughput biological sequencing becomes faster and cheaper, the need to extract useful information from sequencing becomes ever more paramount, often limited by low-throughput experimental characterizations. For proteins, accurate prediction of their functions directly from their primary amino-acid sequences h…
New method learns diverse protein scaffolds for motif design.
Unified model learns from proteins and ligands for drug design.
The inverse Potts problem to infer a Boltzmann distribution for homologous protein sequences from their single-site and pairwise amino acid frequencies recently attracts a great deal of attention in the studies of protein structure and evolution. We study regularization and learning methods and how to tune regularizati…
NeuralMD accelerates protein-ligand binding simulations 1Kx faster.
AbDiffuser generates full-atom antibodies with sequence and structure fidelity.
Tutorial on optimizing diffusion model samples for specific metrics.
FDBM models use fractional Brownian motion to model complex stochastic processes.
The effective representation of proteins is a crucial task that directly affects the performance of many bioinformatics problems. Related proteins usually bind to similar ligands. Chemical characteristics of ligands are known to capture the functional and mechanistic properties of proteins suggesting that a ligand base…
Motivation: Prediction of ligands for proteins of known 3D structure is important to understand structure-function relationship, predict molecular function, or design new drugs. Results: We explore a new approach for ligand prediction in which binding pockets are represented by atom clouds. Each target pocket is compar…
Deep learning models optimize protein sequences.
Mathematical pipeline identifies structural homology of knotted proteins.
Proposes a thermodynamic work minimization framework for guiding generative models.
ProGen models protein sequences for synthetic biology.
New 3D protein analysis methods improve accuracy.
New method detects and compares folding pathways of knotted proteins.
Proteins are commonly used by biochemical industry for numerous processes. Refining these proteins' properties via mutations causes stability effects as well. Accurate computational method to predict how mutations affect protein stability are necessary to facilitate efficient protein design. However, accuracy of predic…
PANDA predicts protein binding affinity changes from sequences, outperforming existing methods.
A new model explains protein interactions via electron delocalization.
EBM predicts protein conformations at atomic scale using crystallized data.
Experimental determination of protein function is resource-consuming. As an alternative, computational prediction of protein function has received attention. In this context, protein structural classification (PSC) can help, by allowing for determining structural classes of currently unclassified proteins based on thei…
Protein interactions constitute the fundamental building block of almost every life activity. Identifying protein communities from Protein-Protein Interaction (PPI) networks is essential to understand the principles of cellular organization and explore the causes of various diseases. It is critical to integrate multipl…
Mathematician summarizes protein geometry and mutation effects.
Algorithm optimizes biological sequences using bootstrapped training with a score-conditioned generator.
Many complex systems can be represented as networks, and the problem of network comparison is becoming increasingly relevant. There are many techniques for network comparison, from simply comparing network summary statistics to sophisticated but computationally costly alignment-based approaches. Yet it remains challeng…
EGR refines and assesses protein complex structures.
We introduce a new model of proteins, which extends and enhances the traditional graphical representation by associating a combinatorial object called a fatgraph to any protein based upon its intrinsic geometry. Fatgraphs can easily be stored and manipulated as triples of permutations, and these methods are therefore a…
Two proteins are homologous if they have a common evolutionary origin, and the binary classification problem is to identify proteins in a candidate set that are homologous to a particular native protein. The feature (explanatory) variables available for classification are various measures of similarity of proteins. The…
As proteins with similar structures often have similar functions, analysis of protein structures can help predict protein functions and is thus important. We consider the problem of protein structure classification, which computationally classifies the structures of proteins into pre-defined groups. We develop a weight…
Flexible Kernels for Protein Property Prediction
Training-free source selection for LLM families with shared vocabularies
New method steers protein design towards desired properties.
A new diffusion model generates novel protein backbones without relying on pretrained networks.
Proteins are linear molecular chains that often fold to function. The topology of folding is widely believed to define its properties and function, and knot theory has been applied to study protein structure and its implications. More that 97% of proteins are, however, classified as unknots when intra-chain interaction…
Many machine learning tasks can be expressed as the transformation---or \emph{transduction}---of input sequences into output sequences: speech recognition, machine translation, protein secondary structure prediction and text-to-speech to name but a few. One of the key challenges in sequence transduction is learning to …
ProtTrans models predict protein features without evolutionary info.
Study improves LLMs for PPI analysis by addressing uncertainty.
Few-step protein backbone generators reduce sampling time by over 20x.