Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

136271407542 · Jun 202019922001200920172026
48 results for protein sampling

Few-step protein backbone generators reduce sampling time by over 20x.

problem Computational bottleneck in diffusion-based protein generation models.
method Score distillation adapted for protein backbone generation, combined with inference time noise modulation.
result Significant reduction in sampling time (20+ fold) while maintaining comparable performance.

New method optimizes protein design by sampling from realistic inputs.

problem Optimizing properties of interest in design problems, especially with black box predictive models.
method Conditioning by Adaptive Sampling, using model-based adaptive sampling to estimate conditional input distributions.
result Achieves state-of-the-art results on protein fluorescence problem.

Study infers evolutionary interactions from protein sequences using regularization methods.

problem Inferring evolutionary interactions from protein sequences.
method Regularization methods, including L2L_2 for fields and group L1L_1 for couplings, with parameter tuning.
result Effective regularization parameters for sparse couplings improve accuracy.

Equivariant flows sample symmetric multi-body systems like proteins.

problem Sampling symmetric multi-body systems like proteins with strong interactions.
method Developed equivariant flows that respect the symmetries of the energy function.
result Equivariant flows can sample new configurations not possible with non-equivariant flows.

A new diffusion model generates novel protein backbones without relying on pretrained networks.

problem Generating novel protein backbones without relying on pretrained networks.
method Developed a SE(3) invariant diffusion model on multiple frames, called FrameDiff.
result Generated designable protein monomers up to 500 amino acids without pretrained networks.

Generative model designs highly designable proteins using geometric algebra.

problem Creating proteins with diverse and statistically accurate secondary structures.
method Introduced a geometric algebra flow matching model (FrameFlow) with Clifford Frame Attention (CFA) for protein backbone design.
result Achieved high designability, diversity, and novelty in protein backbone sampling.

Boltzmann Generators use deep learning to efficiently sample complex systems.

problem Sampling equilibrium states in many-body systems like proteins is computationally challenging.
method Combining deep learning and statistical mechanics, Boltzmann Generators learn a coordinate transformation to generate unbiased samples.
result Boltzmann Generators can generate one-shot equilibrium samples of complex systems and proteins.

A faster method for optimizing DNA and protein sequences using machine learning.

problem Designing DNA and protein sequences with improved function.
method Activation maximization with a straight-through approximation and adaptive entropy variable.
result Fast SeqProp achieves up to 100-fold faster convergence and improved fitness optima.

This paper uses bandit theory and Thompson Sampling to optimize protein sequences.

problem Optimizing protein sequences using machine learning and directed evolution.
method Proposes a Thompson Sampling-guided Directed Evolution (TS-DE) framework.
result TS-DE achieves a nearly optimal Bayesian regret of order ildeO(d2MT) ilde O(d^{2}\sqrt{MT}).

LMI approximates mutual information in high dimensions using learned low-dimensional representations.

problem Estimating mutual information between high-dimensional variables is challenging due to sample size limitations.
method Developed a method called latent MI (LMI) approximation that applies a nonparametric MI estimator to low-dimensional representations learned by a simple model architecture.
result LMI can approximate MI well for variables with >10^3 dimensions if their dependence structure has low intrinsic dimensionality.

New method uses diffusion models to generate proteins with specific motifs.

problem Generating proteins with specific functional substructures (motifs) using diffusion models.
method Adapting SMC-aided diffusion posterior samplers to zero-shot scaffolding tasks.
result Proposed potentials and samplers improve performance in generating proteins with desired motifs.

Variational auto-encoder frameworks have demonstrated success in reducing complex nonlinear dynamics in molecular simulation to a single non-linear embedding. In this work, we illustrate how this non-linear latent embedding can be used as a collective variable for enhanced sampling, and present a simple modification th…

2018-01-02abs ↗pdf ↗

Bayesian Active Learning improves protein docking accuracy and uncertainty quantification.

problem Uncertainty quantification in protein docking optimization.
method Bayesian Active Learning (BAL) for optimization and uncertainty quantification of protein docking.
result BAL significantly improves docking accuracy and provides tight confidence intervals.

This paper proposes a new method to generate protein structures using deep learning.

problem Weak correlation between current scoring functions and protein molecular activity.
method Graph-generative models to sample novel tertiary protein structures.
result Generative models can reveal latent space and highlight structural factors.

New model uses pretrained biochemical language models to generate drug compounds.

problem Developing novel compounds targeting specific proteins.
method Exploits pretrained language models to initialize and fine-tune targeted molecule generation models.
result Warm-started models outperform baseline models, with one-stage strategy showing better generalization.

Mathematical pipeline identifies structural homology of knotted proteins.

problem Quantification and classification of protein structures, especially knotted proteins, require noise-free and complete data.
method Developed a geometric framework using persistent homology to analyze protein structures.
result Persistent homology accurately represents structural homology of knotted proteins and identifies geometric features of protein entanglement.

Link prediction is one of the fundamental problems in network analysis. In many applications, notably in genetics, a partially observed network may not contain any negative examples of absent edges, which creates a difficulty for many existing supervised learning approaches. We develop a new method which treats the obs…

2013-01-29abs ↗pdf ↗

Improved protein identification in mass spectrometry data.

problem Expanding peptide scoring capabilities in tandem mass spectrometry.
method Deriving concave emission distributions for dynamic Bayesian networks.
result Efficiently learned scoring function outperforms state-of-the-art.

Improved protein structure classification using weighted graphlets and deep neural networks.

problem Protein structure classification for function prediction.
method Developed a weighted network and graphlet-based measure, combined with a deep neural network.
result Significantly improved performance on 36 real datasets compared to existing methods.

PANDA predicts protein binding affinity changes from sequences, outperforming existing methods.

problem Accurately predicting changes in protein binding affinity due to mutations.
method Sequence-based machine learning approach using protein sequence information.
result PANDA achieves higher Pearson correlation coefficients than existing methods.

A new model explains protein interactions via electron delocalization.

problem Understanding how protein interactions affect each other.
method Quantized discrete differential geometry of n-simplices.
result Allosteric regulation follows from the model of interactions.

EBM predicts protein conformations at atomic scale using crystallized data.

problem Predicting the conformation of a side chain from its context within a protein structure.
method Energy-based model trained on crystallized protein data, evaluating performance on rotamer recovery task.
result EBM achieves performance close to state-of-the-art methods, including Rosetta energy function.

Deep generative model discovers inhibitors for unknown targets.

problem Discovering novel inhibitor molecules for unknown drug targets.
method Deep generative framework trained on protein sequences, small molecules, and interactions.
result Micromolar-level inhibition observed for two out of four synthesized candidates, including activity against SARS-CoV-2 variants.

Knot theory applied to proteins, distinguishing folded linear chains.

problem Classifying proteins as unknots when intra-chain interactions are ignored.
method Developing knot theory for folded linear molecular chains, considering self-bonding, and using Gauss codes and quandles.
result Extended knot theory to distinguish topologies of proteins with intra-chain bonds.

Experimental determination of protein function is resource-consuming. As an alternative, computational prediction of protein function has received attention. In this context, protein structural classification (PSC) can help, by allowing for determining structural classes of currently unclassified proteins based on thei…

2018-04-12abs ↗pdf ↗

Mathematician summarizes protein geometry and mutation effects.

problem Understanding how proteins mutate and their structure-function relationship.
method Mathematical analysis of protein structures and functions, focusing on hydrogen bonds and secondary structure.
result Protein secondary structure regulates mutation by stabilizing or destabilizing regions.

EGR refines and assesses protein complex structures.

problem Improving the accuracy of protein complex 3D structures for drug discovery.
method E(3)-equivariant graph neural network (GNN) for multi-task refinement and assessment.
result EGR achieves state-of-the-art performance in refining and assessing protein complexes.