Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

25.0%50.0%75.0%100.0% · Apr 199319922001200920182026
48 results for protein topology

Mathematical pipeline identifies structural homology of knotted proteins.

problem Quantification and classification of protein structures, especially knotted proteins, require noise-free and complete data.
method Developed a geometric framework using persistent homology to analyze protein structures.
result Persistent homology accurately represents structural homology of knotted proteins and identifies geometric features of protein entanglement.

Knoto-ID studies the entanglement of open protein chains without closing them.

problem Analyzing the entanglement of open protein chains without altering their geometry.
method Using knotoids, a generalization of knot theory for open curves, to evaluate entanglement without closing the curve.
result Knoto-ID can analyze both global and local topologies of protein chains, identifying non-trivial folds.

Knot theory applied to proteins, distinguishing folded linear chains.

problem Classifying proteins as unknots when intra-chain interactions are ignored.
method Developing knot theory for folded linear molecular chains, considering self-bonding, and using Gauss codes and quandles.
result Extended knot theory to distinguish topologies of proteins with intra-chain bonds.

Pipeline learns topological features for protein stability prediction.

problem Predicting protein stability using topological features.
method Data-driven method to learn topological features, comparing with expert features.
result Topological features achieve 92%-99% of SME-based models' performance.

Study uses knotoids to analyze open protein chains, revealing new topological regions.

problem Characterizing the topology of open protein chains.
method Introduced knotoids as a generalization of knots for open curves, analyzing protein chains without closure.
result Identified new topological regions in protein chains, including pre-knotted regions.

Method optimizes knotting pathways in constrained polymers.

problem Understanding how geometric constraints affect knot formation in polymers.
method Topological steering using knotoid spectrum and mean unravelling number.
result Geometric constraints increase the frequency of twist knots in polymers.

Persistent homology provides a new, efficient molecular descriptor for protein dynamics.

problem Designing effective molecular descriptors for high-dimensional MD trajectories.
method Introduced masked Flood complex, a protein-tailored modification of simplicial complexes, for persistent homology.
result Persistent homology-based descriptors are competitive across protein dynamics tasks, including frame-level observable regression and MSM estimation.

Machine learning predicts signaling peptides from protein star graphs.

problem Predicting signaling activity of proteins from molecular structure.
method Protein star graphs, S2SNet topological indices, Machine Learning (SVM-RFE, Laplacian kernel).
result Best model predicts 98.0% signaling pathways with AUROC 0.961.

Classifies uncolored bonded knots with up to 7 singularity points.

problem Classifying uncolored bonded knots with up to 7 singularity points.
method Generation of planar graphs, conversion into bonded knot diagrams, use of Yamada polynomial, and brute-force Reidemeister moves.
result Systematic classification of uncolored bonded knots with singularity number at most seven.
Graphoidsmath.CO

Graphoids are topological invariants of virtual graph diagrams.

problem Understanding knotted graphs with open ends in proteins and simplifying virtual spatial graphs.
method Topological interpretations of graphoids using graph Reidemeister moves.
result Virtual graphoids are useful for studying knotted graphs and simplifying spatial graphs.

Deep learning model predicts protein-ligand binding modes from docking data.

problem Improving protein-ligand binding mode prediction accuracy.
method Dual-graph architecture with separate sub-networks for ligand topology and protein-ligand interactions.
result Deep learning model outperforms docking programs in binding mode prediction.

Polynomial invariants classify molecular chains based on their contact arrangements.

problem No established invariants for molecular chains with both hard and soft contacts.
method Developed polynomial invariants for circuit topology of molecular chains.
result Polynomial invariants efficiently classify chains with various contact types.

Review of mathematical representations for biomolecular data.

problem Complexity and high dimensionality of biomolecular datasets hinder ML applications.
method Developed low-dimensional and scalable mathematical representations using algebraic topology, differential geometry, and graph theory.
result Mathematical representations improve protein-ligand binding predictions and other biomolecular applications.

SL2MF predicts synthetic lethality using logistic matrix factorization.

problem Predicting synthetic lethality in human cancers from limited experimental data.
method Logistic matrix factorization incorporating biological knowledge.
result SL2MF effectively predicts known and unknown SL interactions.

Link prediction is one of the fundamental problems in network analysis. In many applications, notably in genetics, a partially observed network may not contain any negative examples of absent edges, which creates a difficulty for many existing supervised learning approaches. We develop a new method which treats the obs…

2013-01-29abs ↗pdf ↗

mGPfusion predicts protein stability changes using a novel Gaussian process method.

problem Limited experimental data for predicting protein stability changes.
method Bayesian data fusion model combining experimental and molecular simulation data.
result mGPfusion outperforms state-of-the-art methods in predicting protein stability.

Ensemble method ranks homologous proteins robustly across various similarity metrics.

problem Ranking homologous proteins in a candidate set with high accuracy and robustness.
method Ensemble of models and assessment metrics, phalanxes, and aggregation of diverse metrics.
result Ensemble of phalanxes identifies strong and diverse subsets of feature variables for robust ranking.

Paper proposes MLPCD for protein community detection in large PPI networks.

problem Identifying reliable protein communities from large-scale PPI networks.
method Integrates Gene Expression Data and uses Multi-source Learning with cloud computing.
result Demonstrates superior performance compared to existing methods.

Improved protein structure classification using weighted graphlets and deep neural networks.

problem Protein structure classification for function prediction.
method Developed a weighted network and graphlet-based measure, combined with a deep neural network.
result Significantly improved performance on 36 real datasets compared to existing methods.

PANDA predicts protein binding affinity changes from sequences, outperforming existing methods.

problem Accurately predicting changes in protein binding affinity due to mutations.
method Sequence-based machine learning approach using protein sequence information.
result PANDA achieves higher Pearson correlation coefficients than existing methods.

Sequence-based model predicts protein-protein interactions with high accuracy.

problem Predicting protein-protein interactions for alternative treatment options.
method Sequence clustering, discrete cosine transform, supervised machine learning, SVM with RBF.
result Mesh model achieved an average AUC of 0.84.

A new model explains protein interactions via electron delocalization.

problem Understanding how protein interactions affect each other.
method Quantized discrete differential geometry of n-simplices.
result Allosteric regulation follows from the model of interactions.

EBM predicts protein conformations at atomic scale using crystallized data.

problem Predicting the conformation of a side chain from its context within a protein structure.
method Energy-based model trained on crystallized protein data, evaluating performance on rotamer recovery task.
result EBM achieves performance close to state-of-the-art methods, including Rosetta energy function.

Mathematician summarizes protein geometry and mutation effects.

problem Understanding how proteins mutate and their structure-function relationship.
method Mathematical analysis of protein structures and functions, focusing on hydrogen bonds and secondary structure.
result Protein secondary structure regulates mutation by stabilizing or destabilizing regions.

New multitask algorithm separates rare from frequent protein functions.

problem Challenging automated protein function prediction with unbalanced data.
method Uses dissimilarity information to separate rare class labels, unlike similarity-based approaches.
result Multitask label propagation algorithm performs best with dissimilarity matrix.

Improved 3D generative models for drug design reduce bias and enhance data efficiency.

problem Data sparsity and bias in 3D molecular design models.
method Multi-level contrastive learning protocol for bias control and data efficiency.
result Hierarchical generative models that are topologically unbiased and explainable.