New method classifies protein structures using network features.
problem Efficiently predicting protein function from structural data.
method Modelled protein structures as PSNs, used graphlets and deep learning for features.
result Proposed methods outperform existing PSC approaches in accuracy.
A new framework uses text descriptions to improve protein design.
problem Lack of effective methods to incorporate textual descriptions in protein design.
method ProteinDT framework that combines text and protein structural information.
result ProteinDT significantly improves protein design accuracy and performance.
A new method predicts protein functions using variable-length sequences.
problem Computational methods for protein function prediction are slow and inaccurate for long sequences.
method Two feature sets: single fixed-sized segments and multi-sized segments, using bi-directional LSTM. Combined with MLDA features.
result Significant improvement in accuracy for long protein sequences.
Automated protein function prediction is a challenging problem with distinctive features, such as the hierarchical organization of protein functions and the scarcity of annotated proteins for most biological functions. We propose a multitask learning algorithm addressing both issues. Unlike standard multitask algorithm…
Novel ligand-based method improves protein representation performance.
problem Improving protein representation for bioinformatics tasks.
method Proposes SMILESVec method to represent ligands and compute protein similarity.
result Ligand-based protein representation performs as well as sequence-based methods.
Paper proposes MLPCD for protein community detection in large PPI networks.
problem Identifying reliable protein communities from large-scale PPI networks.
method Integrates Gene Expression Data and uses Multi-source Learning with cloud computing.
result Demonstrates superior performance compared to existing methods.
Improved protein structure classification using weighted graphlets and deep neural networks.
problem Protein structure classification for function prediction.
method Developed a weighted network and graphlet-based measure, combined with a deep neural network.
result Significantly improved performance on 36 real datasets compared to existing methods.
EBM predicts protein conformations at atomic scale using crystallized data.
problem Predicting the conformation of a side chain from its context within a protein structure.
method Energy-based model trained on crystallized protein data, evaluating performance on rotamer recovery task.
result EBM achieves performance close to state-of-the-art methods, including Rosetta energy function.
Tail-GNNs improve protein function prediction using relational reinforcement.
problem Predicting hierarchical protein functions from sequence data.
method Combining Tail-GNNs with dilated convolutional networks for multi-task learning.
result Significant improvement in F_1 score for protein function prediction.
Mathematician summarizes protein geometry and mutation effects.
problem Understanding how proteins mutate and their structure-function relationship.
method Mathematical analysis of protein structures and functions, focusing on hydrogen bonds and secondary structure.
result Protein secondary structure regulates mutation by stabilizing or destabilizing regions.
As high-throughput biological sequencing becomes faster and cheaper, the need to extract useful information from sequencing becomes ever more paramount, often limited by low-throughput experimental characterizations. For proteins, accurate prediction of their functions directly from their primary amino-acid sequences h…
DeepProteomics uses neural networks to classify protein families efficiently.
problem Lack of functional annotation for many protein sequences in databases.
method Used RNN, LSTM, GRU, and deep neural network models on a dataset of 40,433 proteins.
result Achieved maximum 78% accuracy in classifying protein families.
PGEL learns embeddings to diversify protein motifs while maintaining biological function.
problem Generating diverse protein structures while preserving biological function.
method Embedding learning framework that enhances motif diversity in a diffusion model's frozen denoiser.
result PGEL achieves greater structural diversity, better designability, and improved self-consistency compared to partial diffusion.
A new model explains protein interactions via electron delocalization.
problem Understanding how protein interactions affect each other.
method Quantized discrete differential geometry of n-simplices.
result Allosteric regulation follows from the model of interactions.
Computational approaches to drug discovery can reduce the time and cost associated with experimental assays and enable the screening of novel chemotypes. Structure-based drug design methods rely on scoring functions to rank and predict binding affinities and poses. The ever-expanding amount of protein-ligand binding an…
EGR refines and assesses protein complex structures.
problem Improving the accuracy of protein complex 3D structures for drug discovery.
method E(3)-equivariant graph neural network (GNN) for multi-task refinement and assessment.
result EGR achieves state-of-the-art performance in refining and assessing protein complexes.
Knot theory applied to proteins, distinguishing folded linear chains.
problem Classifying proteins as unknots when intra-chain interactions are ignored.
method Developing knot theory for folded linear molecular chains, considering self-bonding, and using Gauss codes and quandles.
result Extended knot theory to distinguish topologies of proteins with intra-chain bonds.
ProGen models protein sequences for synthetic biology.
problem Generating proteins without structural annotations.
method Trained a 1.2B-parameter language model on 280M protein sequences.
result ProGen generates proteins with fine-grained control and accuracy.
Predicts cellular functions in human tissues using multi-layer networks.
problem Challenges in predicting tissue-specific cellular function.
method Hierarchy-aware unsupervised node feature learning for multi-layer networks.
result Improves prediction accuracy of cellular functions in 48 tissues.
New method detects and compares folding pathways of knotted proteins.
problem Understanding the function of knots in protein folding.
method Topological analysis of protein knotoid distributions and entanglement.
result Reveals unique folding pathway for shallow knotted Carbonic Anhydrases.
Combines active learning and imbalance-aware classification for protein function prediction.
problem Scarce positive labels and lack of explicit negative labels in supervised learning.
method Active learning for selecting negative examples and imbalance-aware classification for mitigating label imbalance.
result The combined techniques outperform state-of-the-art methods on protein function prediction benchmarks.
Visualizes 3D CNNs for protein-ligand scoring.
problem Interpreting complex neural network decisions for protein-ligand scoring.
method Three visualization methods for 3D CNNs, including filters and weights.
result Visualizations aid in tuning and designing neural networks.
New 3D protein analysis methods improve accuracy.
problem Lack of suitable learning algorithms for protein data.
method Intrinsic-Extrinsic Convolution and Pooling for 3D protein structures.
result Outperforms state-of-the-art methods on protein analysis tasks.
ChemBoost predicts protein-ligand binding affinity using SMILES syntax.
problem Predicting high affinity drug-target interactions from sequence similarity alone.
method ChemBoost uses SMILES syntax to represent ligands as documents and proteins as sequences or ligand-centric features. It learns chemical word embeddings and predicts affinities using eXtreme Gradient Boosting.
result ChemBoost outperforms state-of-the-art systems in predicting protein-ligand affinities.
This paper proposes a new method to generate protein structures using deep learning.
problem Weak correlation between current scoring functions and protein molecular activity.
method Graph-generative models to sample novel tertiary protein structures.
result Generative models can reveal latent space and highlight structural factors.
Protein Thoughts interprets protein interactions with clear reasoning, improving prediction accuracy.
problem Lack of mechanistic justification in protein-protein interaction predictions.
method Interpretable search problem reformulation, hypothesis-guided entropy-regularized Tree-of-Thoughts search, embedding-space flow matching.
result Improves mean best-binder rank from 47.7 to 11.2 on SHS148k benchmark.
Sequence-based model predicts protein-protein interactions with high accuracy.
problem Predicting protein-protein interactions for alternative treatment options.
method Sequence clustering, discrete cosine transform, supervised machine learning, SVM with RBF.
result Mesh model achieved an average AUC of 0.84.
New method steers protein design towards desired properties.
problem Challenges in designing proteins with specific structures and properties.
method Feynman-Kac framework applied to RFdiffusion models with guiding potentials.
result Significant improvement in predicted interface energetics and binder designability.
Protein function prediction is the important problem in modern biology. In this paper, the un-normalized, symmetric normalized, and random walk graph Laplacian based semi-supervised learning methods will be applied to the integrated network combined from multiple networks to predict the functions of all yeast proteins …
TUNet improves protein classification in cell images.
problem Classifying specific proteins in human cells using microscopy images.
method TUNet model incorporating segmentation maps for improved classification.
result TUNet achieves competitive performance in protein classification.
A new diffusion model generates novel protein backbones without relying on pretrained networks.
problem Generating novel protein backbones without relying on pretrained networks.
method Developed a SE(3) invariant diffusion model on multiple frames, called FrameDiff.
result Generated designable protein monomers up to 500 amino acids without pretrained networks.
Most network-based protein (or gene) function prediction methods are based on the assumption that the labels of two adjacent proteins in the network are likely to be the same. However, assuming the pairwise relationship between proteins or genes is not complete, the information a group of genes that show very similar p…
Improved scaffold generation for protein motifs using SE(3) flow matching.
problem Lack of structural diversity in generated scaffolds for protein motifs.
method Extended FrameFlow for motif-scaffolding with motif amortization and motif guidance.
result 2.5 times more designable and unique motif-scaffolds compared to state-of-the-art.
Article compares different machine learning techniques for protein classification.
problem Predicting enzyme class from unknown proteins is challenging.
method Implemented seven classification techniques on 4368 protein data.
result C5.0 classification technique gives highest accuracy and balanced performance.
Motivation: Proteins are known to undergo conformational changes in the course of their functions. The changes in conformation are often attributable to a small fraction of residues within the protein. Therefore identification of these variable regions is important for an understanding of protein function. Results: We …
New method maps protein sequences to embeddings encoding structural information.
problem Inferring structural properties from amino acid sequences when structures are unknown.
method Representation learning using bidirectional LSTM models with structural similarity and residue contact maps.
result Trained embeddings improve structural similarity prediction and transfer to other tasks.
We use a semisupervised learning algorithm based on a topological data analysis approach to assign functional categories to yeast proteins using similarity graphs. This new approach to analyzing biological networks yields results that are as good as or better than state of the art existing approaches.
EnzyNet classifies enzymes using 3D CNN on spatial structure.
problem Predicting enzyme function from amino acid sequence is unreliable.
method 3D convolutional neural networks on voxel-based spatial structure.
result Achieved 78.4% accuracy on 63,558 enzymes.
Deep learning predicts protein structures accurately.
problem Predicting the 3D structure of proteins from amino acid sequences.
method Embeddings and deep learning models for backbone atom distance matrices and torsion angles.
result Competitive results in CASP13 and CASP12, surpassing previous winners.
This thesis improves protein contact prediction using unsupervised and supervised methods.
problem Improving accuracy of protein contact prediction.
method Unsupervised and supervised deep learning methods.
result A scoring system called diversity score for measuring contact novelty.
A novel method clusters protein conformations from MD simulations.
problem Clustering long MD protein dynamics for identifying states and behavior.
method Adversarial Autoencoder (AAE) for conformation clustering.
result Identifies many salient features of the folding process.
Bayesian Active Learning improves protein docking accuracy and uncertainty quantification.
problem Uncertainty quantification in protein docking optimization.
method Bayesian Active Learning (BAL) for optimization and uncertainty quantification of protein docking.
result BAL significantly improves docking accuracy and provides tight confidence intervals.
Artificial neural networks (ANNs) have gained a well-deserved popularity among machine learning tools upon their recent successful applications in image- and sound processing and classification problems. ANNs have also been applied for predicting the family or function of a protein, knowing its residue sequence. Here w…
Machine learning predicts signaling peptides from protein star graphs.
problem Predicting signaling activity of proteins from molecular structure.
method Protein star graphs, S2SNet topological indices, Machine Learning (SVM-RFE, Laplacian kernel).
result Best model predicts 98.0% signaling pathways with AUROC 0.961.
New method optimizes protein design by sampling from realistic inputs.
problem Optimizing properties of interest in design problems, especially with black box predictive models.
method Conditioning by Adaptive Sampling, using model-based adaptive sampling to estimate conditional input distributions.
result Achieves state-of-the-art results on protein fluorescence problem.
Deep neural network improves amino acid side chain prediction accuracy.
problem Predicting amino acid side chain conformation for protein modeling and design.
method Deep neural network architecture without physics-based assumptions.
result Improved accuracy by more than 25% for aromatic residues.
pyLEMMINGS improves bioinformatics protein function prediction.
problem Lack of accurate instance-level protein annotations.
method Stochastic sub-gradient optimization for large-margin multiple instance classification and ranking.
result pyLEMMINGS achieves state-of-the-art performance in bioinformatics tasks.
Deep learning models optimize protein sequences.
problem Optimizing protein properties through sequence design.
method Deep generative models guided by machine learning.
result Improved protein sequence generation from prior knowledge.