Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

2825638451,126 · Jun 202019922001200920182026
48 results for Protein Data Bank

DiAMoNDBack models protein backmapping from coarse-grained Cα traces.

problem Restoring all-atom details from coarse-grained protein representations.
method Autoregressive denoising diffusion model for residue-by-residue backmapping.
result Achieves state-of-the-art reconstruction performance in diverse applications.

Long, flexible physical filaments are naturally tangled and knotted, from macroscopic string down to long-chain molecules. The existence of knotting in a filament naturally affects its configuration and properties, and may be very stable or disappear rapidly under manipulation and interaction. Knotting has been previou…

2016-11-18abs ↗pdf ↗

When analyzing the genome, researchers have discovered that proteins bind to DNA based on certain patterns of the DNA sequence known as "motifs". However, it is difficult to manually construct motifs due to their complexity. Recently, externally learned memory models have proven to be effective methods for reasoning ov…

2017-02-22abs ↗pdf ↗

mGPfusion predicts protein stability changes using a novel Gaussian process method.

problem Limited experimental data for predicting protein stability changes.
method Bayesian data fusion model combining experimental and molecular simulation data.
result mGPfusion outperforms state-of-the-art methods in predicting protein stability.

Mathematical pipeline identifies structural homology of knotted proteins.

problem Quantification and classification of protein structures, especially knotted proteins, require noise-free and complete data.
method Developed a geometric framework using persistent homology to analyze protein structures.
result Persistent homology accurately represents structural homology of knotted proteins and identifies geometric features of protein entanglement.

EBM predicts protein conformations at atomic scale using crystallized data.

problem Predicting the conformation of a side chain from its context within a protein structure.
method Energy-based model trained on crystallized protein data, evaluating performance on rotamer recovery task.
result EBM achieves performance close to state-of-the-art methods, including Rosetta energy function.

Paper proposes MLPCD for protein community detection in large PPI networks.

problem Identifying reliable protein communities from large-scale PPI networks.
method Integrates Gene Expression Data and uses Multi-source Learning with cloud computing.
result Demonstrates superior performance compared to existing methods.

Two deep learning models predict protein-protein interactions with high accuracy.

problem Overfitting and information leak in deep learning models for PPI prediction.
method Carefully designed deep learning models, strict conditions for training and testing, and methodology to avoid information leak.
result Best model predicts more than 78% of human PPI with strong confidence.

ProtTrans models predict protein features without evolutionary info.

problem Predicting protein features from amino acid sequences.
method Self-supervised deep learning on large protein datasets.
result ProtT5 embeddings outperform state-of-the-art for per-residue predictions.

Article compares different machine learning techniques for protein classification.

problem Predicting enzyme class from unknown proteins is challenging.
method Implemented seven classification techniques on 4368 protein data.
result C5.0 classification technique gives highest accuracy and balanced performance.

Develops a fast BMF approach for binary matrices.

problem Finding patterns in binary matrices for various applications.
method MEBF (Median Expansion for Boolean Factorization) using geometric segmentation and heuristic submatrix identification.
result Superior performance in reconstruction error and computational efficiency compared to existing methods.

End-to-end model predicts protein interfaces from atomic coordinates.

problem Improving protein interface prediction using large datasets.
method Developed SASNet, an end-to-end learning model using only atomic coordinates.
result SASNet outperforms state-of-the-art methods trained on gold-standard data.

Riemannian geometry improves protein dynamics analysis.

problem Efficient analysis of protein dynamics data in non-linear spaces.
method Developed a local approximation technique for geodesics and a smooth manifold of protein conformations.
result Geodesics approximate molecular dynamics trajectories and provide realistic summary statistics.

DeepAffinity predicts compound-protein affinity from sequences, outperforming existing methods.

problem Lack of methods to predict compound-protein affinity from sequences alone.
method Unified RNN/GCNN-CNN model that unifies recurrent and convolutional neural networks.
result Model outperforms conventional options in predicting affinities with high accuracy.

Deep learning predicts protein structures accurately.

problem Predicting the 3D structure of proteins from amino acid sequences.
method Embeddings and deep learning models for backbone atom distance matrices and torsion angles.
result Competitive results in CASP13 and CASP12, surpassing previous winners.

Pipeline learns topological features for protein stability prediction.

problem Predicting protein stability using topological features.
method Data-driven method to learn topological features, comparing with expert features.
result Topological features achieve 92%-99% of SME-based models' performance.

Researchers use shape analysis to recover protein structures from Cryo-EM data.

problem Recovering the three-dimensional backbone structure of single polypeptide proteins from noisy tomographic projections.
method Shape analysis and matrix Lie group actions to deform point clouds to match 2D tomography data.
result Optimal deformations are computed to recover the three-dimensional backbone structure of proteins.

DeepProteomics uses neural networks to classify protein families efficiently.

problem Lack of functional annotation for many protein sequences in databases.
method Used RNN, LSTM, GRU, and deep neural network models on a dataset of 40,433 proteins.
result Achieved maximum 78% accuracy in classifying protein families.

Study analyzes profitability and efficiency of Chinese banks, finding state-owned banks superior.

problem Analyzing efficiency and profitability of Chinese banks over time.
method Used Data envelopment analysis (Super-SBM-UND-VRS based DEA) model considering non-performing loans as undesired output.
result State-owned banks and Rural/City Commercial Banks have better profitability super-efficiency than Joint-stock Banks.

This study uses high-frequency data to identify early warning signals for bank crises.

problem Identifying early warning signals for impending bank crises.
method Constructing multiple recurrence networks (MRNs) based on high-frequency stock returns to monitor nonlinear dynamics.
result Key indicators of MRNs, particularly average mutual information, provide valuable insights into periods of extreme volatility.

Researchers infer gene activity in dividing cells, accounting for protein inheritance and division history.

problem Inferring protein production kinetics in dividing cells due to protein inheritance and division history.
method Adapted conditional normalizing flows to approximate intractable likelihoods from simulated data.
result Glc3 gene is mostly inactive under stress, with brief and transient expression.

The paper tests for association between latent community memberships in multi-view network data.

problem Evaluating the independence of latent community memberships in multi-view network data.
method Extended stochastic block model for two-view network data, developed a new hypothesis test.
result Evidence of weak association between latent community memberships in binary interaction and co-complex association data.

Automated protein structure prediction from cryo-EM data.

problem Challenging to build atomic models from cryo-EM densities without prior structure.
method Uses GCN and LSTM to automate model building from amino acid identities and candidate locations.
result Automated approach reduces time and eliminates human intervention for protein structure determination.

Study examines factors influencing lending to SMEs by Kenyan banks.

problem Lack of creditworthiness makes SMEs difficult to finance by banks.
method Descriptive research design, census of 43 banks, secondary data analysis.
result Bank size and liquidity significantly influence lending to SMEs, while credit risk and interest rates do not.

PLUS pre-trains protein sequences with structural info, improving performance.

problem Lack of labeled protein sequences for training models.
method PLUS combines masked language modeling with same-family prediction for pre-training.
result PLUS-RNN outperforms other models in protein biology tasks.

Deep Learning identifies 20 critical proteins linked to FLT3-ITD mutation in leukemia.

problem Identifying critical proteins associated with FLT3-ITD mutation in leukemia.
method Hierarchical Deep Learning network using autoencoders for feature extraction.
result Deep Learning accurately correlates 20 critical proteins with FLT3-ITD mutation (97% accuracy).

TAPE benchmarks protein learning tasks, finds self-supervised pretraining boosts performance.

problem Fragmented datasets and lack of standardized evaluation in protein modeling.
method TAPE introduces five semi-supervised learning tasks, curates splits, benchmarks models.
result Self-supervised pretraining more than doubles performance in some cases.

ChemBoost predicts protein-ligand binding affinity using SMILES syntax.

problem Predicting high affinity drug-target interactions from sequence similarity alone.
method ChemBoost uses SMILES syntax to represent ligands as documents and proteins as sequences or ligand-centric features. It learns chemical word embeddings and predicts affinities using eXtreme Gradient Boosting.
result ChemBoost outperforms state-of-the-art systems in predicting protein-ligand affinities.

Novel parallel GNN predicts protein-ligand interactions with high accuracy.

problem Accurate prediction of protein-ligand interactions for drug design.
method Parallel Graph Neural Networks (GNN) integrating 3D structural data.
result GNN achieves high accuracy in predicting binary interactions and activity.

Unified model learns from proteins and ligands for drug design.

problem Disjoint data sources and modeling assumptions limit joint use of structure- and ligand-based drug design.
method Contrastive Geometric Learning for Unified Computational Drug Design (ConGLUDe)
result Unified model achieves competitive zero-shot virtual screening performance and state-of-the-art ligand-conditioned pocket selection.