New method detects and compares folding pathways of knotted proteins.
problem Understanding the function of knots in protein folding.
method Topological analysis of protein knotoid distributions and entanglement.
result Reveals unique folding pathway for shallow knotted Carbonic Anhydrases.
Machine learning predicts protein structures and simulates dynamics.
problem Understanding and predicting protein folding and dynamics.
method Machine learning techniques for structure prediction and simulation.
result Machine learning enhances protein simulation and structure prediction.
Focusing on a small set of proteins that i) fold in a concerted, all-or-none fashion and ii) do not contain knots or slipknots, we show that the Gauss linking integral, the torsion and the number of sequence-distant contacts provide information regarding the folding rate. Our results suggest that the global topology/ge…
Recently exciting progress has been made on protein contact prediction, but the predicted contacts for proteins without many sequence homologs is still of low quality and not very useful for de novo structure prediction. This paper presents a new deep learning method that predicts contacts by integrating both evolution…
Knot theory applied to proteins, distinguishing folded linear chains.
problem Classifying proteins as unknots when intra-chain interactions are ignored.
method Developing knot theory for folded linear molecular chains, considering self-bonding, and using Gauss codes and quandles.
result Extended knot theory to distinguish topologies of proteins with intra-chain bonds.
Paper improves zero-shot protein stability prediction by clarifying free-energy foundations.
problem Improving zero-shot protein stability prediction using inverse folding models.
method Clarifying the free-energy foundations of inverse folding models and proposing better estimates of relative stability.
result Significant gains in zero-shot performance can be achieved with simple methods.
Non-autoregressive method speeds up protein folding prediction 23 times.
problem Generating protein sequences with higher order interactions.
method Discrete diffusion conditioned on 3D structure using ProteinMPNN.
result 23 times speed up in inference without performance loss.
Recent developments in specialized computer hardware have greatly accelerated atomic level Molecular Dynamics (MD) simulations. A single GPU-attached cluster is capable of producing microsecond-length trajectories in reasonable amounts of time. Multiple protein states and a large number of microstates associated with f…
Enhanced coloring invariant distinguishes folded molecular chain topologies.
problem Apparent indistinguishability of folded chain topologies using current coloring invariants.
method Introduced Boltzmann weights to improve the resolving power of quandle colorings.
result Improved resolution in distinguishing folded chain topologies.
Continuous-depth Evoformer reduces protein folding prediction time and resource usage.
problem Efficient protein structure prediction with reduced computational costs.
method Continuous-depth formulation of Evoformer using Neural Ordinary Differential Equations (Neural ODEs).
result The continuous-time Evoformer achieves constant memory cost and improved efficiency.
A new machine-learned CG model predicts protein structures efficiently.
problem Developing a universal, computationally efficient protein simulation model.
method Combining deep learning with all-atom protein simulations to create a transferable CG force field.
result The model predicts protein structures, intermediates, and fluctuations efficiently.
A geometric analysis of protein folding, which complements many of the models in the literature, is presented. We examine the process from unfolded strand to the point where the strand becomes self-interacting. A central question is how it is possible that so many initial configurations proceed to fold to a unique fina…
The backbone of most proteins forms an open curve. To study their entanglement, a common strategy consists in searching for the presence of knots in their backbones using topological invariants. However, this approach requires to close the curve into a loop, which alters the geometry of curve. Knoto-ID allows evaluatin…
Motivation: Drug discovery demands rapid quantification of compound-protein interaction (CPI). However, there is a lack of methods that can predict compound-protein affinity from sequences alone with high applicability, accuracy, and interpretability. Results: We present a seamless integration of domain knowledges and …
DiAMoNDBack models protein backmapping from coarse-grained Cα traces.
problem Restoring all-atom details from coarse-grained protein representations.
method Autoregressive denoising diffusion model for residue-by-residue backmapping.
result Achieves state-of-the-art reconstruction performance in diverse applications.
New 3D protein analysis methods improve accuracy.
problem Lack of suitable learning algorithms for protein data.
method Intrinsic-Extrinsic Convolution and Pooling for 3D protein structures.
result Outperforms state-of-the-art methods on protein analysis tasks.
Polynomial invariants classify molecular chains based on their contact arrangements.
problem No established invariants for molecular chains with both hard and soft contacts.
method Developed polynomial invariants for circuit topology of molecular chains.
result Polynomial invariants efficiently classify chains with various contact types.
This thesis improves protein contact prediction using unsupervised and supervised methods.
problem Improving accuracy of protein contact prediction.
method Unsupervised and supervised deep learning methods.
result A scoring system called diversity score for measuring contact novelty.
We present a machine learning framework for modeling protein dynamics. Our approach uses L1-regularized, reversible hidden Markov models to understand large protein datasets generated via molecular dynamics simulations. Our model is motivated by three design principles: (1) the requirement of massive scalability; (2) t…
Smoothed fitness landscape improves protein optimization.
problem Infeasibility of combinatorially large protein sequence space.
method Formulate protein fitness as a graph signal, smooth using Tikunov regularization, and optimize with Gibbs sampling.
result 2.5 fold fitness improvement over training set.
A new approach to protein language models combines latent space prediction with masked language modeling.
problem Improving protein language models by predicting amino acid identities at masked positions.
method A variant of masked language modeling that predicts latent targets only at masked positions, retaining the MLM cross-entropy.
result The new approach outperforms pure masked language modeling on 11 out of 16 downstream tasks.
Enhanced diffusion sampling improves rare event sampling in biomolecular simulations.
problem Efficiently sampling rare transition events in biomolecular systems.
method Quantitative steering protocols to generate biased ensembles and exact reweighting.
result Fast, accurate, and scalable estimation of equilibrium properties.
Enhanced diffusion sampling tackles rare event sampling in biomolecular simulations.
problem Efficiently sampling rare transition events in biomolecular simulations.
method Quantitative steering protocols to generate biased ensembles, followed by exact reweighting.
result Fast, accurate, and scalable estimation of equilibrium properties for folding free energies.
A deep neural network based architecture was constructed to predict amino acid side chain conformation with unprecedented accuracy. Amino acid side chain conformation prediction is essential for protein homology modeling and protein design. Current widely-adopted methods use physics-based energy functions to evaluate s…
Persistent homology provides a new, efficient molecular descriptor for protein dynamics.
problem Designing effective molecular descriptors for high-dimensional MD trajectories.
method Introduced masked Flood complex, a protein-tailored modification of simplicial complexes, for persistent homology.
result Persistent homology-based descriptors are competitive across protein dynamics tasks, including frame-level observable regression and MSM estimation.
Few-step protein backbone generators reduce sampling time by over 20x.
problem Computational bottleneck in diffusion-based protein generation models.
method Score distillation adapted for protein backbone generation, combined with inference time noise modulation.
result Significant reduction in sampling time (20+ fold) while maintaining comparable performance.
During the past decade, with the significant progress of computational power as well as ever-rising data availability, deep learning techniques became increasingly popular due to their excellent performance on computer vision problems. The size of the Protein Data Bank has increased more than 15 fold since 1999, which …
A faster method for optimizing DNA and protein sequences using machine learning.
problem Designing DNA and protein sequences with improved function.
method Activation maximization with a straight-through approximation and adaptive entropy variable.
result Fast SeqProp achieves up to 100-fold faster convergence and improved fitness optima.
The inverse Potts problem to infer a Boltzmann distribution for homologous protein sequences from their single-site and pairwise amino acid frequencies recently attracts a great deal of attention in the studies of protein structure and evolution. We study regularization and learning methods and how to tune regularizati…
We present a new method for design problems wherein the goal is to maximize or specify the value of one or more properties of interest. For example, in protein design, one may wish to find the protein sequence that maximizes fluorescence. We assume access to one or more, potentially black box, stochastic "oracle" predi…
Paper improves Tm prediction of protein fragments using sparsity and probabilistic models.
problem Improving accuracy of melting temperature prediction for protein fragments.
method Promoting sparsity in pre-trained transformer models and adopting probabilistic frameworks.
result Mean absolute error of 0.23C for predicting melting temperature.
New methods avoid spectral pollution in transfer operators for accurate analysis.
problem Spectral pollution in finite-dimensional approximations of transfer operators.
method Algorithms for computing spectral properties of transfer operators without spectral pollution.
result Accurate spectral estimation across various applications, including protein folding models.
A framework for multilayer networks predicts links without shared structures.
problem Link prediction in multilayer networks without shared structures.
method Bi-level model averaging with K-fold cross-validation. result Framework outperforms existing methods in predictive accuracy and robustness.
Protein contacts contain important information for protein structure and functional study, but contact prediction from sequence remains very challenging. Both evolutionary coupling (EC) analysis and supervised machine learning methods are developed to predict contacts, making use of different types of information, resp…
Solvents can induce helical knots in simulated biopolymer tubes.
problem Understanding how solvents influence the folding of biopolymers.
method Computer simulations using morphometric solvation techniques.
result Solvents can induce complex helical structures, including knots.
RaNNDy uses randomized neural networks to learn transfer operators efficiently.
problem Efficiently learning transfer operators from data.
method Randomized neural network approach with randomly initialized hidden layers and trained output layer.
result Significant reduction in training time and resources with improved stability.
KCUSUM detects abrupt changes in real-time data streams efficiently.
problem Detecting abrupt changes in high-volume scientific data streams.
method Kernel-based Cumulative Sum (KCUSUM) algorithm using Maximum Mean Discrepancy (MMD).
result KCUSUM outperforms traditional CUSUM in online change point detection.
Optimal transport embedding learns feature sets efficiently.
problem Learning on sets of features with long-range dependencies and few labeled data.
method Parametrized fixed-size embedding that aggregates features according to optimal transport plan.
result Achieves state-of-the-art results on protein fold recognition and chromatin profiles.
Molecular simulations produce very high-dimensional data-sets with millions of data points. As analysis methods are often unable to cope with so many dimensions, it is common to use dimensionality reduction and clustering methods to reach a reduced representation of the data. Yet these methods often fail to capture the…
Motivation: A major challenge in the development of machine learning based methods in computational biology is that data may not be accurately labeled due to the time and resources required for experimentally annotating properties of proteins and DNA sequences. Standard supervised learning algorithms assume accurate in…
Proposes a continuous relaxation for discrete Bayesian optimization.
problem Efficiently optimizing discrete data with limited target observations.
method Continuous relaxation of objective function, incorporating prior knowledge.
result Optimization can be computationally tractable with few observations.
Review of mathematical representations for biomolecular data.
problem Complexity and high dimensionality of biomolecular datasets hinder ML applications.
method Developed low-dimensional and scalable mathematical representations using algebraic topology, differential geometry, and graph theory.
result Mathematical representations improve protein-ligand binding predictions and other biomolecular applications.
The modeling of atomistic biomolecular simulations using kinetic models such as Markov state models (MSMs) has had many notable algorithmic advances in recent years. The variational principle has opened the door for a nearly fully automated toolkit for selecting models that predict the long-time kinetics from molecular…
Proposes a thermodynamic work minimization framework for guiding generative models.
problem Guiding generative models in sparse-data regimes with limited target samples or constraints.
method Regularization framework inspired by thermodynamic work, introducing Path Guidance and Observable Guidance.
result Improves sample efficiency and reduces bias in molecular simulations.
Branching Flows generates sequences of varying lengths using binary trees.
problem Generating sequences of unknown lengths or fixed elements.
method A generative modeling framework that evolves states over binary trees, controlling sequence length.
result Branching Flows can generate sequences of varying lengths and mix different types of state spaces.
The effective representation of proteins is a crucial task that directly affects the performance of many bioinformatics problems. Related proteins usually bind to similar ligands. Chemical characteristics of ligands are known to capture the functional and mechanistic properties of proteins suggesting that a ligand base…
A new framework uses text descriptions to improve protein design.
problem Lack of effective methods to incorporate textual descriptions in protein design.
method ProteinDT framework that combines text and protein structural information.
result ProteinDT significantly improves protein design accuracy and performance.
Deep learning models optimize protein sequences.
problem Optimizing protein properties through sequence design.
method Deep generative models guided by machine learning.
result Improved protein sequence generation from prior knowledge.