Mathematical pipeline identifies structural homology of knotted proteins.
problem Quantification and classification of protein structures, especially knotted proteins, require noise-free and complete data.
method Developed a geometric framework using persistent homology to analyze protein structures.
result Persistent homology accurately represents structural homology of knotted proteins and identifies geometric features of protein entanglement.
Ensemble method ranks homologous proteins robustly across various similarity metrics.
problem Ranking homologous proteins in a candidate set with high accuracy and robustness.
method Ensemble of models and assessment metrics, phalanxes, and aggregation of diverse metrics.
result Ensemble of phalanxes identifies strong and diverse subsets of feature variables for robust ranking.
Persistent homology can recognize knotting in curves.
problem Recognizing knotting in curves
method Compute one-dimensional persistent homology, extract cycle representatives, and assign a hypergraph curvature-based score.
result Systematic differences between knotted and unknotted structures are revealed.
Persistent homology provides a new, efficient molecular descriptor for protein dynamics.
problem Designing effective molecular descriptors for high-dimensional MD trajectories.
method Introduced masked Flood complex, a protein-tailored modification of simplicial complexes, for persistent homology.
result Persistent homology-based descriptors are competitive across protein dynamics tasks, including frame-level observable regression and MSM estimation.
New lattice path method for statistical inference of persistent diagrams.
problem Statistical inference on persistent diagrams.
method Lattice path representation and combinatorial enumerations.
result Topological changes observed in spike proteins of COVID-19 virus.
Motivation. Protein contact map describes the pairwise spatial and functional relationship of residues in a protein and contains key information for protein 3D structure prediction. Although studied extensively, it remains very challenging to predict contact map using only sequence information. Most existing methods pr…
Deep learning predicts protein contacts with high accuracy.
problem Low quality contact predictions for proteins without homologs.
method Integrates evolutionary coupling and sequence conservation through an ultra-deep neural network.
result Significantly outperforms existing methods in contact prediction and ab initio folding.
New method predicts protein features using statistical relational learning.
problem Challenges in automatic protein feature annotation due to limited homology data.
method Introduces Semantically Based Regularization to incorporate prior knowledge.
result Improved overall prediction quality with constraints.
ProteinNet provides a standardized data set for protein structure prediction.
problem Lack of standardized data sets for protein structure prediction.
method Created high-quality sequence alignments, multiple data splits, and validation sets.
result Facilitates fair assessment of machine learning models for protein structure.
Deep neural network improves amino acid side chain prediction accuracy.
problem Predicting amino acid side chain conformation for protein modeling and design.
method Deep neural network architecture without physics-based assumptions.
result Improved accuracy by more than 25% for aromatic residues.
Study infers evolutionary interactions from protein sequences using regularization methods.
problem Inferring evolutionary interactions from protein sequences.
method Regularization methods, including L2 for fields and group L1 for couplings, with parameter tuning. result Effective regularization parameters for sparse couplings improve accuracy.
A new method estimates protein evolutionary fields and couplings from alignments.
problem Estimating evolutionary fields and couplings from protein sequence alignments.
method Boltzmann machine with parallel, persistent Markov chain Monte Carlo method.
result Improved precision in predicting contact residue pairs.
Bayesian Active Learning improves protein docking accuracy and uncertainty quantification.
problem Uncertainty quantification in protein docking optimization.
method Bayesian Active Learning (BAL) for optimization and uncertainty quantification of protein docking.
result BAL significantly improves docking accuracy and provides tight confidence intervals.
New method analyzes knots and links using multiscale Gauss link integral.
problem Lack of localization and quantization in knot theory applications.
method Integrates curve segmentation and multiscale analysis into the Gauss link integral.
result Significantly outperforms other methods in protein flexibility analysis.
A new approach to protein language models combines latent space prediction with masked language modeling.
problem Improving protein language models by predicting amino acid identities at masked positions.
method A variant of masked language modeling that predicts latent targets only at masked positions, retaining the MLM cross-entropy.
result The new approach outperforms pure masked language modeling on 11 out of 16 downstream tasks.
New algorithm improves model generalization in structured biomedical domains.
problem Improving model generalization in structured biomedical domains.
method Proposes a new regret minimization (RGM) algorithm and its structured extension for better performance in diverse environments.
result Significantly outperforms previous state-of-the-art baselines on molecular property prediction, protein homology, and stability prediction.
New method for manifold topological learning avoids remeshing issues.
problem Persistent homology on manifolds is numerically inconsistent.
method Persistent de Rham-Hodge Laplacians in Eulerian representation.
result Avoids numerical inconsistency over multiscale manifolds.
Novel ligand-based method improves protein representation performance.
problem Improving protein representation for bioinformatics tasks.
method Proposes SMILESVec method to represent ligands and compute protein similarity.
result Ligand-based protein representation performs as well as sequence-based methods.
A new framework uses text descriptions to improve protein design.
problem Lack of effective methods to incorporate textual descriptions in protein design.
method ProteinDT framework that combines text and protein structural information.
result ProteinDT significantly improves protein design accuracy and performance.
Deep learning models optimize protein sequences.
problem Optimizing protein properties through sequence design.
method Deep generative models guided by machine learning.
result Improved protein sequence generation from prior knowledge.
New method classifies protein structures using network features.
problem Efficiently predicting protein function from structural data.
method Modelled protein structures as PSNs, used graphlets and deep learning for features.
result Proposed methods outperform existing PSC approaches in accuracy.
mGPfusion predicts protein stability changes using a novel Gaussian process method.
problem Limited experimental data for predicting protein stability changes.
method Bayesian data fusion model combining experimental and molecular simulation data.
result mGPfusion outperforms state-of-the-art methods in predicting protein stability.
ProGen models protein sequences for synthetic biology.
problem Generating proteins without structural annotations.
method Trained a 1.2B-parameter language model on 280M protein sequences.
result ProGen generates proteins with fine-grained control and accuracy.
Paper proposes MLPCD for protein community detection in large PPI networks.
problem Identifying reliable protein communities from large-scale PPI networks.
method Integrates Gene Expression Data and uses Multi-source Learning with cloud computing.
result Demonstrates superior performance compared to existing methods.
New 3D protein analysis methods improve accuracy.
problem Lack of suitable learning algorithms for protein data.
method Intrinsic-Extrinsic Convolution and Pooling for 3D protein structures.
result Outperforms state-of-the-art methods on protein analysis tasks.
New method detects and compares folding pathways of knotted proteins.
problem Understanding the function of knots in protein folding.
method Topological analysis of protein knotoid distributions and entanglement.
result Reveals unique folding pathway for shallow knotted Carbonic Anhydrases.
Fast and efficient homology algorithms are in demand in the applied sciences for analyzing solid materials and proteins, processing digital imaging data, or pattern classification among others. Recent advances employ discrete Morse theory as a preprocessor. Research in this area has lead to the need to find complicated…
Improved protein structure classification using weighted graphlets and deep neural networks.
problem Protein structure classification for function prediction.
method Developed a weighted network and graphlet-based measure, combined with a deep neural network.
result Significantly improved performance on 36 real datasets compared to existing methods.
PANDA predicts protein binding affinity changes from sequences, outperforming existing methods.
problem Accurately predicting changes in protein binding affinity due to mutations.
method Sequence-based machine learning approach using protein sequence information.
result PANDA achieves higher Pearson correlation coefficients than existing methods.
Sequence-based model predicts protein-protein interactions with high accuracy.
problem Predicting protein-protein interactions for alternative treatment options.
method Sequence clustering, discrete cosine transform, supervised machine learning, SVM with RBF.
result Mesh model achieved an average AUC of 0.84.
A new model explains protein interactions via electron delocalization.
problem Understanding how protein interactions affect each other.
method Quantized discrete differential geometry of n-simplices.
result Allosteric regulation follows from the model of interactions.
EBM predicts protein conformations at atomic scale using crystallized data.
problem Predicting the conformation of a side chain from its context within a protein structure.
method Energy-based model trained on crystallized protein data, evaluating performance on rotamer recovery task.
result EBM achieves performance close to state-of-the-art methods, including Rosetta energy function.
Knot theory applied to proteins, distinguishing folded linear chains.
problem Classifying proteins as unknots when intra-chain interactions are ignored.
method Developing knot theory for folded linear molecular chains, considering self-bonding, and using Gauss codes and quandles.
result Extended knot theory to distinguish topologies of proteins with intra-chain bonds.
Mathematician summarizes protein geometry and mutation effects.
problem Understanding how proteins mutate and their structure-function relationship.
method Mathematical analysis of protein structures and functions, focusing on hydrogen bonds and secondary structure.
result Protein secondary structure regulates mutation by stabilizing or destabilizing regions.
Mapper tool preserves graph structures for better visualization.
problem Graphs can be hard to visualize for large datasets.
method Developed a variation of mapper for weighted, undirected graphs.
result Homology-preserving skeletons enable multi-scale visualization.
New multitask algorithm separates rare from frequent protein functions.
problem Challenging automated protein function prediction with unbalanced data.
method Uses dissimilarity information to separate rare class labels, unlike similarity-based approaches.
result Multitask label propagation algorithm performs best with dissimilarity matrix.
Machine learning predicts protein structures and simulates dynamics.
problem Understanding and predicting protein folding and dynamics.
method Machine learning techniques for structure prediction and simulation.
result Machine learning enhances protein simulation and structure prediction.
A new method predicts protein functions using variable-length sequences.
problem Computational methods for protein function prediction are slow and inaccurate for long sequences.
method Two feature sets: single fixed-sized segments and multi-sized segments, using bi-directional LSTM. Combined with MLDA features.
result Significant improvement in accuracy for long protein sequences.
EGR refines and assesses protein complex structures.
problem Improving the accuracy of protein complex 3D structures for drug discovery.
method E(3)-equivariant graph neural network (GNN) for multi-task refinement and assessment.
result EGR achieves state-of-the-art performance in refining and assessing protein complexes.
We introduce a new model of proteins, which extends and enhances the traditional graphical representation by associating a combinatorial object called a fatgraph to any protein based upon its intrinsic geometry. Fatgraphs can easily be stored and manipulated as triples of permutations, and these methods are therefore a…
Deep neural networks predict protein functions from sequences.
problem Accurately predicting protein functions from amino acid sequences.
method Artificial recurrent neural networks (RNN) with LSTM units trained on annotated datasets.
result RNN models achieved high performance for in-class and out-of-class protein function predictions.
Flexible Kernels for Protein Property Prediction
problem Predicting protein properties from sparse experimental data
method Sequence kernels using evolutionary substitution matrices and local linearity
result Data-efficient models of protein property landscapes
New method maps protein sequences to embeddings encoding structural information.
problem Inferring structural properties from amino acid sequences when structures are unknown.
method Representation learning using bidirectional LSTM models with structural similarity and residue contact maps.
result Trained embeddings improve structural similarity prediction and transfer to other tasks.
New method steers protein design towards desired properties.
problem Challenges in designing proteins with specific structures and properties.
method Feynman-Kac framework applied to RFdiffusion models with guiding potentials.
result Significant improvement in predicted interface energetics and binder designability.
A new diffusion model generates novel protein backbones without relying on pretrained networks.
problem Generating novel protein backbones without relying on pretrained networks.
method Developed a SE(3) invariant diffusion model on multiple frames, called FrameDiff.
result Generated designable protein monomers up to 500 amino acids without pretrained networks.
DeepProteomics uses neural networks to classify protein families efficiently.
problem Lack of functional annotation for many protein sequences in databases.
method Used RNN, LSTM, GRU, and deep neural network models on a dataset of 40,433 proteins.
result Achieved maximum 78% accuracy in classifying protein families.
ProtTrans models predict protein features without evolutionary info.
problem Predicting protein features from amino acid sequences.
method Self-supervised deep learning on large protein datasets.
result ProtT5 embeddings outperform state-of-the-art for per-residue predictions.
WideDTA predicts drug-target binding affinity using text-based information.
problem Predicting drug-target binding affinity is a major challenge in drug discovery.
method WideDTA uses chemical and biological textual sequence information, including protein sequence, ligand SMILES, protein domains and motifs, and maximum common substructure words.
result WideDTA outperformed DeepDTA on the KIBA dataset, indicating the word-based sequence representation is a promising alternative.