Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

164329493657 · Jun 202019922001200920182026
48 results for protein function

A new method predicts protein functions using variable-length sequences.

problem Computational methods for protein function prediction are slow and inaccurate for long sequences.
method Two feature sets: single fixed-sized segments and multi-sized segments, using bi-directional LSTM. Combined with MLDA features.
result Significant improvement in accuracy for long protein sequences.

Automated protein function prediction is a challenging problem with distinctive features, such as the hierarchical organization of protein functions and the scarcity of annotated proteins for most biological functions. We propose a multitask learning algorithm addressing both issues. Unlike standard multitask algorithm…

2016-11-03abs ↗pdf ↗

Paper proposes MLPCD for protein community detection in large PPI networks.

problem Identifying reliable protein communities from large-scale PPI networks.
method Integrates Gene Expression Data and uses Multi-source Learning with cloud computing.
result Demonstrates superior performance compared to existing methods.

Improved protein structure classification using weighted graphlets and deep neural networks.

problem Protein structure classification for function prediction.
method Developed a weighted network and graphlet-based measure, combined with a deep neural network.
result Significantly improved performance on 36 real datasets compared to existing methods.

EBM predicts protein conformations at atomic scale using crystallized data.

problem Predicting the conformation of a side chain from its context within a protein structure.
method Energy-based model trained on crystallized protein data, evaluating performance on rotamer recovery task.
result EBM achieves performance close to state-of-the-art methods, including Rosetta energy function.

Mathematician summarizes protein geometry and mutation effects.

problem Understanding how proteins mutate and their structure-function relationship.
method Mathematical analysis of protein structures and functions, focusing on hydrogen bonds and secondary structure.
result Protein secondary structure regulates mutation by stabilizing or destabilizing regions.

DeepProteomics uses neural networks to classify protein families efficiently.

problem Lack of functional annotation for many protein sequences in databases.
method Used RNN, LSTM, GRU, and deep neural network models on a dataset of 40,433 proteins.
result Achieved maximum 78% accuracy in classifying protein families.

PGEL learns embeddings to diversify protein motifs while maintaining biological function.

problem Generating diverse protein structures while preserving biological function.
method Embedding learning framework that enhances motif diversity in a diffusion model's frozen denoiser.
result PGEL achieves greater structural diversity, better designability, and improved self-consistency compared to partial diffusion.

A new model explains protein interactions via electron delocalization.

problem Understanding how protein interactions affect each other.
method Quantized discrete differential geometry of n-simplices.
result Allosteric regulation follows from the model of interactions.

Computational approaches to drug discovery can reduce the time and cost associated with experimental assays and enable the screening of novel chemotypes. Structure-based drug design methods rely on scoring functions to rank and predict binding affinities and poses. The ever-expanding amount of protein-ligand binding an…

2016-12-08abs ↗pdf ↗

EGR refines and assesses protein complex structures.

problem Improving the accuracy of protein complex 3D structures for drug discovery.
method E(3)-equivariant graph neural network (GNN) for multi-task refinement and assessment.
result EGR achieves state-of-the-art performance in refining and assessing protein complexes.

Knot theory applied to proteins, distinguishing folded linear chains.

problem Classifying proteins as unknots when intra-chain interactions are ignored.
method Developing knot theory for folded linear molecular chains, considering self-bonding, and using Gauss codes and quandles.
result Extended knot theory to distinguish topologies of proteins with intra-chain bonds.

Combines active learning and imbalance-aware classification for protein function prediction.

problem Scarce positive labels and lack of explicit negative labels in supervised learning.
method Active learning for selecting negative examples and imbalance-aware classification for mitigating label imbalance.
result The combined techniques outperform state-of-the-art methods on protein function prediction benchmarks.

ChemBoost predicts protein-ligand binding affinity using SMILES syntax.

problem Predicting high affinity drug-target interactions from sequence similarity alone.
method ChemBoost uses SMILES syntax to represent ligands as documents and proteins as sequences or ligand-centric features. It learns chemical word embeddings and predicts affinities using eXtreme Gradient Boosting.
result ChemBoost outperforms state-of-the-art systems in predicting protein-ligand affinities.

This paper proposes a new method to generate protein structures using deep learning.

problem Weak correlation between current scoring functions and protein molecular activity.
method Graph-generative models to sample novel tertiary protein structures.
result Generative models can reveal latent space and highlight structural factors.

Protein Thoughts interprets protein interactions with clear reasoning, improving prediction accuracy.

problem Lack of mechanistic justification in protein-protein interaction predictions.
method Interpretable search problem reformulation, hypothesis-guided entropy-regularized Tree-of-Thoughts search, embedding-space flow matching.
result Improves mean best-binder rank from 47.7 to 11.2 on SHS148k benchmark.

Sequence-based model predicts protein-protein interactions with high accuracy.

problem Predicting protein-protein interactions for alternative treatment options.
method Sequence clustering, discrete cosine transform, supervised machine learning, SVM with RBF.
result Mesh model achieved an average AUC of 0.84.

A new diffusion model generates novel protein backbones without relying on pretrained networks.

problem Generating novel protein backbones without relying on pretrained networks.
method Developed a SE(3) invariant diffusion model on multiple frames, called FrameDiff.
result Generated designable protein monomers up to 500 amino acids without pretrained networks.

Most network-based protein (or gene) function prediction methods are based on the assumption that the labels of two adjacent proteins in the network are likely to be the same. However, assuming the pairwise relationship between proteins or genes is not complete, the information a group of genes that show very similar p…

2012-12-03abs ↗pdf ↗

Improved scaffold generation for protein motifs using SE(3) flow matching.

problem Lack of structural diversity in generated scaffolds for protein motifs.
method Extended FrameFlow for motif-scaffolding with motif amortization and motif guidance.
result 2.5 times more designable and unique motif-scaffolds compared to state-of-the-art.

Article compares different machine learning techniques for protein classification.

problem Predicting enzyme class from unknown proteins is challenging.
method Implemented seven classification techniques on 4368 protein data.
result C5.0 classification technique gives highest accuracy and balanced performance.

New method maps protein sequences to embeddings encoding structural information.

problem Inferring structural properties from amino acid sequences when structures are unknown.
method Representation learning using bidirectional LSTM models with structural similarity and residue contact maps.
result Trained embeddings improve structural similarity prediction and transfer to other tasks.

Deep learning predicts protein structures accurately.

problem Predicting the 3D structure of proteins from amino acid sequences.
method Embeddings and deep learning models for backbone atom distance matrices and torsion angles.
result Competitive results in CASP13 and CASP12, surpassing previous winners.

Bayesian Active Learning improves protein docking accuracy and uncertainty quantification.

problem Uncertainty quantification in protein docking optimization.
method Bayesian Active Learning (BAL) for optimization and uncertainty quantification of protein docking.
result BAL significantly improves docking accuracy and provides tight confidence intervals.

Machine learning predicts signaling peptides from protein star graphs.

problem Predicting signaling activity of proteins from molecular structure.
method Protein star graphs, S2SNet topological indices, Machine Learning (SVM-RFE, Laplacian kernel).
result Best model predicts 98.0% signaling pathways with AUROC 0.961.

New method optimizes protein design by sampling from realistic inputs.

problem Optimizing properties of interest in design problems, especially with black box predictive models.
method Conditioning by Adaptive Sampling, using model-based adaptive sampling to estimate conditional input distributions.
result Achieves state-of-the-art results on protein fluorescence problem.

pyLEMMINGS improves bioinformatics protein function prediction.

problem Lack of accurate instance-level protein annotations.
method Stochastic sub-gradient optimization for large-margin multiple instance classification and ranking.
result pyLEMMINGS achieves state-of-the-art performance in bioinformatics tasks.