Deep neural networks predict protein functions from sequences.
problem Accurately predicting protein functions from amino acid sequences.
method Artificial recurrent neural networks (RNN) with LSTM units trained on annotated datasets.
result RNN models achieved high performance for in-class and out-of-class protein function predictions.
New multitask algorithm separates rare from frequent protein functions.
problem Challenging automated protein function prediction with unbalanced data.
method Uses dissimilarity information to separate rare class labels, unlike similarity-based approaches.
result Multitask label propagation algorithm performs best with dissimilarity matrix.
Tail-GNNs improve protein function prediction using relational reinforcement.
problem Predicting hierarchical protein functions from sequence data.
method Combining Tail-GNNs with dilated convolutional networks for multi-task learning.
result Significant improvement in F_1 score for protein function prediction.
EBM predicts protein conformations at atomic scale using crystallized data.
problem Predicting the conformation of a side chain from its context within a protein structure.
method Energy-based model trained on crystallized protein data, evaluating performance on rotamer recovery task.
result EBM achieves performance close to state-of-the-art methods, including Rosetta energy function.
CNN scoring function predicts protein-ligand interactions.
problem Scoring protein-ligand interactions for drug discovery.
method Convolutional Neural Networks (CNN) for 3D protein-ligand interactions.
result CNN scoring function outperforms AutoDock Vina in ranking poses.
A new method predicts protein functions using variable-length sequences.
problem Computational methods for protein function prediction are slow and inaccurate for long sequences.
method Two feature sets: single fixed-sized segments and multi-sized segments, using bi-directional LSTM. Combined with MLDA features.
result Significant improvement in accuracy for long protein sequences.
Predicts cellular functions in human tissues using multi-layer networks.
problem Challenges in predicting tissue-specific cellular function.
method Hierarchy-aware unsupervised node feature learning for multi-layer networks.
result Improves prediction accuracy of cellular functions in 48 tissues.
A new framework uses text descriptions to improve protein design.
problem Lack of effective methods to incorporate textual descriptions in protein design.
method ProteinDT framework that combines text and protein structural information.
result ProteinDT significantly improves protein design accuracy and performance.
Combines active learning and imbalance-aware classification for protein function prediction.
problem Scarce positive labels and lack of explicit negative labels in supervised learning.
method Active learning for selecting negative examples and imbalance-aware classification for mitigating label imbalance.
result The combined techniques outperform state-of-the-art methods on protein function prediction benchmarks.
Improved protein structure classification using weighted graphlets and deep neural networks.
problem Protein structure classification for function prediction.
method Developed a weighted network and graphlet-based measure, combined with a deep neural network.
result Significantly improved performance on 36 real datasets compared to existing methods.
ChemBoost predicts protein-ligand binding affinity using SMILES syntax.
problem Predicting high affinity drug-target interactions from sequence similarity alone.
method ChemBoost uses SMILES syntax to represent ligands as documents and proteins as sequences or ligand-centric features. It learns chemical word embeddings and predicts affinities using eXtreme Gradient Boosting.
result ChemBoost outperforms state-of-the-art systems in predicting protein-ligand affinities.
This thesis improves protein contact prediction using unsupervised and supervised methods.
problem Improving accuracy of protein contact prediction.
method Unsupervised and supervised deep learning methods.
result A scoring system called diversity score for measuring contact novelty.
Deep learning predicts protein structures accurately.
problem Predicting the 3D structure of proteins from amino acid sequences.
method Embeddings and deep learning models for backbone atom distance matrices and torsion angles.
result Competitive results in CASP13 and CASP12, surpassing previous winners.
Article compares different machine learning techniques for protein classification.
problem Predicting enzyme class from unknown proteins is challenging.
method Implemented seven classification techniques on 4368 protein data.
result C5.0 classification technique gives highest accuracy and balanced performance.
Protein Thoughts interprets protein interactions with clear reasoning, improving prediction accuracy.
problem Lack of mechanistic justification in protein-protein interaction predictions.
method Interpretable search problem reformulation, hypothesis-guided entropy-regularized Tree-of-Thoughts search, embedding-space flow matching.
result Improves mean best-binder rank from 47.7 to 11.2 on SHS148k benchmark.
Sequence-based model predicts protein-protein interactions with high accuracy.
problem Predicting protein-protein interactions for alternative treatment options.
method Sequence clustering, discrete cosine transform, supervised machine learning, SVM with RBF.
result Mesh model achieved an average AUC of 0.84.
New method classifies protein structures using network features.
problem Efficiently predicting protein function from structural data.
method Modelled protein structures as PSNs, used graphlets and deep learning for features.
result Proposed methods outperform existing PSC approaches in accuracy.
Protein function prediction is the important problem in modern biology. In this paper, the un-normalized, symmetric normalized, and random walk graph Laplacian based semi-supervised learning methods will be applied to the integrated network combined from multiple networks to predict the functions of all yeast proteins …
Deep neural network improves amino acid side chain prediction accuracy.
problem Predicting amino acid side chain conformation for protein modeling and design.
method Deep neural network architecture without physics-based assumptions.
result Improved accuracy by more than 25% for aromatic residues.
Most network-based protein (or gene) function prediction methods are based on the assumption that the labels of two adjacent proteins in the network are likely to be the same. However, assuming the pairwise relationship between proteins or genes is not complete, the information a group of genes that show very similar p…
Novel ligand-based method improves protein representation performance.
problem Improving protein representation for bioinformatics tasks.
method Proposes SMILESVec method to represent ligands and compute protein similarity.
result Ligand-based protein representation performs as well as sequence-based methods.
Machine learning predicts signaling peptides from protein star graphs.
problem Predicting signaling activity of proteins from molecular structure.
method Protein star graphs, S2SNet topological indices, Machine Learning (SVM-RFE, Laplacian kernel).
result Best model predicts 98.0% signaling pathways with AUROC 0.961.
Visualizes 3D CNNs for protein-ligand scoring.
problem Interpreting complex neural network decisions for protein-ligand scoring.
method Three visualization methods for 3D CNNs, including filters and weights.
result Visualizations aid in tuning and designing neural networks.
PS8-Net improves eight-state protein secondary structure prediction accuracy.
problem Precise prediction of eight-state protein secondary structure (PSS) is crucial in bioinformatics.
method PS8-Net is a new deep convolutional neural network (DCNN) that uses a PS8 module with skip connections to enhance accuracy.
result PS8-Net achieves 76.89% Q8 accuracy on benchmark datasets.
mGPfusion predicts protein stability changes using a novel Gaussian process method.
problem Limited experimental data for predicting protein stability changes.
method Bayesian data fusion model combining experimental and molecular simulation data.
result mGPfusion outperforms state-of-the-art methods in predicting protein stability.
New method maps protein sequences to embeddings encoding structural information.
problem Inferring structural properties from amino acid sequences when structures are unknown.
method Representation learning using bidirectional LSTM models with structural similarity and residue contact maps.
result Trained embeddings improve structural similarity prediction and transfer to other tasks.
We use a semisupervised learning algorithm based on a topological data analysis approach to assign functional categories to yeast proteins using similarity graphs. This new approach to analyzing biological networks yields results that are as good as or better than state of the art existing approaches.
Deep learning predicts protein contacts with high accuracy.
problem Low quality contact predictions for proteins without homologs.
method Integrates evolutionary coupling and sequence conservation through an ultra-deep neural network.
result Significantly outperforms existing methods in contact prediction and ab initio folding.
New model predicts protein-ligand binding affinity from atomic coordinates.
problem Predicting protein-ligand binding affinity using empirical scoring functions.
method Developed atomic convolutional neural network to learn chemical interactions directly from atomic coordinates.
result Atomic convolutional networks outperform or compete with cheminformatics methods in predicting binding free energy.
Deep neural networks achieve near perfect protein classification.
problem Classifying protein sequences into families and Gene Ontology classes.
method Developed two new ANN models for multi-label protein classification.
result Achieved AUC scores of 99.99% for 698 UniProt families and 99.45% for 983 Gene Ontology classes.
PANDA predicts protein binding affinity changes from sequences, outperforming existing methods.
problem Accurately predicting changes in protein binding affinity due to mutations.
method Sequence-based machine learning approach using protein sequence information.
result PANDA achieves higher Pearson correlation coefficients than existing methods.
Paper improves Tm prediction of protein fragments using sparsity and probabilistic models.
problem Improving accuracy of melting temperature prediction for protein fragments.
method Promoting sparsity in pre-trained transformer models and adopting probabilistic frameworks.
result Mean absolute error of 0.23C for predicting melting temperature.
Motivation. Protein contact map describes the pairwise spatial and functional relationship of residues in a protein and contains key information for protein 3D structure prediction. Although studied extensively, it remains very challenging to predict contact map using only sequence information. Most existing methods pr…
WideDTA predicts drug-target binding affinity using text-based information.
problem Predicting drug-target binding affinity is a major challenge in drug discovery.
method WideDTA uses chemical and biological textual sequence information, including protein sequence, ligand SMILES, protein domains and motifs, and maximum common substructure words.
result WideDTA outperformed DeepDTA on the KIBA dataset, indicating the word-based sequence representation is a promising alternative.
EnzyNet classifies enzymes using 3D CNN on spatial structure.
problem Predicting enzyme function from amino acid sequence is unreliable.
method 3D convolutional neural networks on voxel-based spatial structure.
result Achieved 78.4% accuracy on 63,558 enzymes.
New method optimizes protein design by sampling from realistic inputs.
problem Optimizing properties of interest in design problems, especially with black box predictive models.
method Conditioning by Adaptive Sampling, using model-based adaptive sampling to estimate conditional input distributions.
result Achieves state-of-the-art results on protein fluorescence problem.
Bayesian Active Learning improves protein docking accuracy and uncertainty quantification.
problem Uncertainty quantification in protein docking optimization.
method Bayesian Active Learning (BAL) for optimization and uncertainty quantification of protein docking.
result BAL significantly improves docking accuracy and provides tight confidence intervals.
Flexible Kernels for Protein Property Prediction
problem Predicting protein properties from sparse experimental data
method Sequence kernels using evolutionary substitution matrices and local linearity
result Data-efficient models of protein property landscapes
New method predicts protein structural quality using random forests.
problem Predicting protein structural quality from models.
method Combines multi/single model quality assessment with chemical, physical, geometric features and random forests.
result Random forest technique improves local quality assessment accuracy.
Protein contacts contain important information for protein structure and functional study, but contact prediction from sequence remains very challenging. Both evolutionary coupling (EC) analysis and supervised machine learning methods are developed to predict contacts, making use of different types of information, resp…
Deep learning predicts molecular functions from 3D fields.
problem Predicting molecular functions from 3D fields.
method Deep learning models trained on approximated electron density and electrostatic potential fields.
result Deep learning achieves comparable performance to state-of-the-art methods.
New method steers protein design towards desired properties.
problem Challenges in designing proteins with specific structures and properties.
method Feynman-Kac framework applied to RFdiffusion models with guiding potentials.
result Significant improvement in predicted interface energetics and binder designability.
ProtTrans models predict protein features without evolutionary info.
problem Predicting protein features from amino acid sequences.
method Self-supervised deep learning on large protein datasets.
result ProtT5 embeddings outperform state-of-the-art for per-residue predictions.
Deep model learns protein interfaces from high-order interactions.
problem Predicting protein interfaces from amino acid pairs.
method Graph neural networks and convolutional neural networks for 2D dense predictions.
result Our method consistently improves interface prediction performance.
New neural network predicts accurate protein complex structures.
problem Predicting accurate protein complex structures from atomic coordinates.
method Rotation-equivariant neural network combining point-based representation, equivariance, local convolutions, and hierarchical subsampling.
result Significant improvement in identifying accurate structural models.
A new method predicts compounds for orphan proteins.
problem Predicting binding affinities for orphan proteins.
method Corresponding projections for transfer learning.
result The method outperforms state-of-the-art in orphan screening.
Two deep learning models predict protein-protein interactions with high accuracy.
problem Overfitting and information leak in deep learning models for PPI prediction.
method Carefully designed deep learning models, strict conditions for training and testing, and methodology to avoid information leak.
result Best model predicts more than 78% of human PPI with strong confidence.
Machine learning predicts protein structures and simulates dynamics.
problem Understanding and predicting protein folding and dynamics.
method Machine learning techniques for structure prediction and simulation.
result Machine learning enhances protein simulation and structure prediction.