New method detects and compares folding pathways of knotted proteins.
problem Understanding the function of knots in protein folding.
method Topological analysis of protein knotoid distributions and entanglement.
result Reveals unique folding pathway for shallow knotted Carbonic Anhydrases.
Study shows protein folding rate linked to topological changes.
problem Understanding protein folding kinetics and topology.
method Gauss linking integral, torsion, and sequence-distant contacts.
result Protein topology shifts from right-handed to left-handed with decreasing folding rate, associated with more sequence-distant contacts.
Machine learning predicts protein structures and simulates dynamics.
problem Understanding and predicting protein folding and dynamics.
method Machine learning techniques for structure prediction and simulation.
result Machine learning enhances protein simulation and structure prediction.
A novel method clusters protein conformations from MD simulations.
problem Clustering long MD protein dynamics for identifying states and behavior.
method Adversarial Autoencoder (AAE) for conformation clustering.
result Identifies many salient features of the folding process.
Recently exciting progress has been made on protein contact prediction, but the predicted contacts for proteins without many sequence homologs is still of low quality and not very useful for de novo structure prediction. This paper presents a new deep learning method that predicts contacts by integrating both evolution…
Knot theory applied to proteins, distinguishing folded linear chains.
problem Classifying proteins as unknots when intra-chain interactions are ignored.
method Developing knot theory for folded linear molecular chains, considering self-bonding, and using Gauss codes and quandles.
result Extended knot theory to distinguish topologies of proteins with intra-chain bonds.
Knoto-ID studies the entanglement of open protein chains without closing them.
problem Analyzing the entanglement of open protein chains without altering their geometry.
method Using knotoids, a generalization of knot theory for open curves, to evaluate entanglement without closing the curve.
result Knoto-ID can analyze both global and local topologies of protein chains, identifying non-trivial folds.
Paper improves zero-shot protein stability prediction by clarifying free-energy foundations.
problem Improving zero-shot protein stability prediction using inverse folding models.
method Clarifying the free-energy foundations of inverse folding models and proposing better estimates of relative stability.
result Significant gains in zero-shot performance can be achieved with simple methods.
Non-autoregressive method speeds up protein folding prediction 23 times.
problem Generating protein sequences with higher order interactions.
method Discrete diffusion conditioned on 3D structure using ProteinMPNN.
result 23 times speed up in inference without performance loss.
Continuous-depth Evoformer reduces protein folding prediction time and resource usage.
problem Efficient protein structure prediction with reduced computational costs.
method Continuous-depth formulation of Evoformer using Neural Ordinary Differential Equations (Neural ODEs).
result The continuous-time Evoformer achieves constant memory cost and improved efficiency.
Enhanced coloring invariant distinguishes folded molecular chain topologies.
problem Apparent indistinguishability of folded chain topologies using current coloring invariants.
method Introduced Boltzmann weights to improve the resolving power of quandle colorings.
result Improved resolution in distinguishing folded chain topologies.
A new machine-learned CG model predicts protein structures efficiently.
problem Developing a universal, computationally efficient protein simulation model.
method Combining deep learning with all-atom protein simulations to create a transferable CG force field.
result The model predicts protein structures, intermediates, and fluctuations efficiently.
A geometric analysis of protein folding, which complements many of the models in the literature, is presented. We examine the process from unfolded strand to the point where the strand becomes self-interacting. A central question is how it is possible that so many initial configurations proceed to fold to a unique fina…
DeepAffinity predicts compound-protein affinity from sequences, outperforming existing methods.
problem Lack of methods to predict compound-protein affinity from sequences alone.
method Unified RNN/GCNN-CNN model that unifies recurrent and convolutional neural networks.
result Model outperforms conventional options in predicting affinities with high accuracy.
DiAMoNDBack models protein backmapping from coarse-grained Cα traces.
problem Restoring all-atom details from coarse-grained protein representations.
method Autoregressive denoising diffusion model for residue-by-residue backmapping.
result Achieves state-of-the-art reconstruction performance in diverse applications.
Polynomial invariants classify molecular chains based on their contact arrangements.
problem No established invariants for molecular chains with both hard and soft contacts.
method Developed polynomial invariants for circuit topology of molecular chains.
result Polynomial invariants efficiently classify chains with various contact types.
New 3D protein analysis methods improve accuracy.
problem Lack of suitable learning algorithms for protein data.
method Intrinsic-Extrinsic Convolution and Pooling for 3D protein structures.
result Outperforms state-of-the-art methods on protein analysis tasks.
This thesis improves protein contact prediction using unsupervised and supervised methods.
problem Improving accuracy of protein contact prediction.
method Unsupervised and supervised deep learning methods.
result A scoring system called diversity score for measuring contact novelty.
We present a machine learning framework for modeling protein dynamics. Our approach uses L1-regularized, reversible hidden Markov models to understand large protein datasets generated via molecular dynamics simulations. Our model is motivated by three design principles: (1) the requirement of massive scalability; (2) t…
Smoothed fitness landscape improves protein optimization.
problem Infeasibility of combinatorially large protein sequence space.
method Formulate protein fitness as a graph signal, smooth using Tikunov regularization, and optimize with Gibbs sampling.
result 2.5 fold fitness improvement over training set.
A new approach to protein language models combines latent space prediction with masked language modeling.
problem Improving protein language models by predicting amino acid identities at masked positions.
method A variant of masked language modeling that predicts latent targets only at masked positions, retaining the MLM cross-entropy.
result The new approach outperforms pure masked language modeling on 11 out of 16 downstream tasks.
Enhanced diffusion sampling improves rare event sampling in biomolecular simulations.
problem Efficiently sampling rare transition events in biomolecular systems.
method Quantitative steering protocols to generate biased ensembles and exact reweighting.
result Fast, accurate, and scalable estimation of equilibrium properties.
Enhanced diffusion sampling tackles rare event sampling in biomolecular simulations.
problem Efficiently sampling rare transition events in biomolecular simulations.
method Quantitative steering protocols to generate biased ensembles, followed by exact reweighting.
result Fast, accurate, and scalable estimation of equilibrium properties for folding free energies.
A deep neural network based architecture was constructed to predict amino acid side chain conformation with unprecedented accuracy. Amino acid side chain conformation prediction is essential for protein homology modeling and protein design. Current widely-adopted methods use physics-based energy functions to evaluate s…
Persistent homology provides a new, efficient molecular descriptor for protein dynamics.
problem Designing effective molecular descriptors for high-dimensional MD trajectories.
method Introduced masked Flood complex, a protein-tailored modification of simplicial complexes, for persistent homology.
result Persistent homology-based descriptors are competitive across protein dynamics tasks, including frame-level observable regression and MSM estimation.
Few-step protein backbone generators reduce sampling time by over 20x.
problem Computational bottleneck in diffusion-based protein generation models.
method Score distillation adapted for protein backbone generation, combined with inference time noise modulation.
result Significant reduction in sampling time (20+ fold) while maintaining comparable performance.
Study infers evolutionary interactions from protein sequences using regularization methods.
problem Inferring evolutionary interactions from protein sequences.
method Regularization methods, including L2 for fields and group L1 for couplings, with parameter tuning. result Effective regularization parameters for sparse couplings improve accuracy.
During the past decade, with the significant progress of computational power as well as ever-rising data availability, deep learning techniques became increasingly popular due to their excellent performance on computer vision problems. The size of the Protein Data Bank has increased more than 15 fold since 1999, which …
New method optimizes protein design by sampling from realistic inputs.
problem Optimizing properties of interest in design problems, especially with black box predictive models.
method Conditioning by Adaptive Sampling, using model-based adaptive sampling to estimate conditional input distributions.
result Achieves state-of-the-art results on protein fluorescence problem.
A faster method for optimizing DNA and protein sequences using machine learning.
problem Designing DNA and protein sequences with improved function.
method Activation maximization with a straight-through approximation and adaptive entropy variable.
result Fast SeqProp achieves up to 100-fold faster convergence and improved fitness optima.
Paper improves Tm prediction of protein fragments using sparsity and probabilistic models.
problem Improving accuracy of melting temperature prediction for protein fragments.
method Promoting sparsity in pre-trained transformer models and adopting probabilistic frameworks.
result Mean absolute error of 0.23C for predicting melting temperature.
New methods avoid spectral pollution in transfer operators for accurate analysis.
problem Spectral pollution in finite-dimensional approximations of transfer operators.
method Algorithms for computing spectral properties of transfer operators without spectral pollution.
result Accurate spectral estimation across various applications, including protein folding models.
A framework for multilayer networks predicts links without shared structures.
problem Link prediction in multilayer networks without shared structures.
method Bi-level model averaging with K-fold cross-validation. result Framework outperforms existing methods in predictive accuracy and robustness.
Protein contacts contain important information for protein structure and functional study, but contact prediction from sequence remains very challenging. Both evolutionary coupling (EC) analysis and supervised machine learning methods are developed to predict contacts, making use of different types of information, resp…
RaNNDy uses randomized neural networks to learn transfer operators efficiently.
problem Efficiently learning transfer operators from data.
method Randomized neural network approach with randomly initialized hidden layers and trained output layer.
result Significant reduction in training time and resources with improved stability.
Solvents can induce helical knots in simulated biopolymer tubes.
problem Understanding how solvents influence the folding of biopolymers.
method Computer simulations using morphometric solvation techniques.
result Solvents can induce complex helical structures, including knots.
KCUSUM detects abrupt changes in real-time data streams efficiently.
problem Detecting abrupt changes in high-volume scientific data streams.
method Kernel-based Cumulative Sum (KCUSUM) algorithm using Maximum Mean Discrepancy (MMD).
result KCUSUM outperforms traditional CUSUM in online change point detection.
Deep learning shows neural networks can be trained with limited data, revealing a low-dimensional manifold of optimal configurations.
problem How neural networks can be trained with limited data despite having billions of potential configurations.
method Using mutual information between layers of a deep neural network to speed up training and find optimal configurations.
result Adding structure to neural networks that enforces higher mutual information between layers speeds training and leads to more accurate results.
SRVs improve MSMs for Trp-cage miniprotein, revealing new folding states.
problem Constructing high-resolution MSMs for complex protein dynamics.
method Employing SRVs as feature set for MSM construction, leveraging slowest modes identified by SRVs.
result SRV-MSMs reveal new folding states and faster convergence.
Optimal transport embedding learns feature sets efficiently.
problem Learning on sets of features with long-range dependencies and few labeled data.
method Parametrized fixed-size embedding that aggregates features according to optimal transport plan.
result Achieves state-of-the-art results on protein fold recognition and chromatin profiles.
Molecular simulations produce very high-dimensional data-sets with millions of data points. As analysis methods are often unable to cope with so many dimensions, it is common to use dimensionality reduction and clustering methods to reach a reduced representation of the data. Yet these methods often fail to capture the…
Motivation: A major challenge in the development of machine learning based methods in computational biology is that data may not be accurately labeled due to the time and resources required for experimentally annotating properties of proteins and DNA sequences. Standard supervised learning algorithms assume accurate in…
Proposes a continuous relaxation for discrete Bayesian optimization.
problem Efficiently optimizing discrete data with limited target observations.
method Continuous relaxation of objective function, incorporating prior knowledge.
result Optimization can be computationally tractable with few observations.
sCSC clusters data without prior assumptions, revealing natural groupings.
problem Clustering data without prior knowledge of its structure.
method sCSC performs binary splittings maximizing dissimilarity, producing a binary tree.
result Clusters emerge naturally from the binary tree, revealing data structure.
Review of mathematical representations for biomolecular data.
problem Complexity and high dimensionality of biomolecular datasets hinder ML applications.
method Developed low-dimensional and scalable mathematical representations using algebraic topology, differential geometry, and graph theory.
result Mathematical representations improve protein-ligand binding predictions and other biomolecular applications.
Proposes a thermodynamic work minimization framework for guiding generative models.
problem Guiding generative models in sparse-data regimes with limited target samples or constraints.
method Regularization framework inspired by thermodynamic work, introducing Path Guidance and Observable Guidance.
result Improves sample efficiency and reduces bias in molecular simulations.
Branching Flows generates sequences of varying lengths using binary trees.
problem Generating sequences of unknown lengths or fixed elements.
method A generative modeling framework that evolves states over binary trees, controlling sequence length.
result Branching Flows can generate sequences of varying lengths and mix different types of state spaces.
The effective representation of proteins is a crucial task that directly affects the performance of many bioinformatics problems. Related proteins usually bind to similar ligands. Chemical characteristics of ligands are known to capture the functional and mechanistic properties of proteins suggesting that a ligand base…