Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

3.6%7.1%10.7%14.3% · Oct 199219922001200920182026
48 results for peptide identification

OLCS-Ranker improves peptide identification accuracy and speed on hard datasets.

problem Efficiently identifying peptides from MS/MS data, especially on hard datasets with many false positives.
method Cost-sensitive online learning model and iterative online learning algorithm.
result OLCS-Ranker outperforms existing methods in accuracy and speed on large datasets.

Major histocompatibility complex class two (MHC-II) molecules are trans-membrane proteins and key components of the cellular immune system. Upon recognition of foreign peptides expressed on the MHC-II binding groove, helper T cells mount an immune response against invading pathogens. Therefore, mechanistic identificati…

2017-12-01abs ↗pdf ↗

PepCVAE designs novel antimicrobial peptides using a semi-supervised VAE.

problem Designing novel antimicrobial peptides for next-generation antimicrobial resistance solutions.
method Semi-supervised variational autoencoder (VAE) model that learns latent and antimicrobial attribute spaces from unlabeled and labeled data.
result PepCVAE generates novel AMPs with higher long-range diversity and closer to biological peptide distribution.

Study improves peptide design efficiency using active and meta-learning.

problem Designing novel functional peptides is inefficient due to low throughput and lack of data.
method Investigated active learning and meta-learning for optimizing peptide design experiments.
result Meta-learning improved average accuracy, but neither method outperformed random choice.

DeepNovoV2 improves de novo peptide sequencing from mass spectrometry data.

problem De novo peptide sequencing from mass spectrometry data for personalized cancer vaccines.
method DeepNovoV2 combines T-Net and recurrent neural networks for end-to-end training and prediction.
result DeepNovoV2 achieves 13.01-23.95\% higher accuracy than previous methods.

Improved protein identification in mass spectrometry data.

problem Expanding peptide scoring capabilities in tandem mass spectrometry.
method Deriving concave emission distributions for dynamic Bayesian networks.
result Efficiently learned scoring function outperforms state-of-the-art.

Deep neural networks improve free energy calculations for peptide conformations.

problem Challenges in developing suitable mappings for free energy perturbation.
method Adapted machine learning approach to train deep neural networks for mapping between Boltzmann distributions.
result Accurate free energy differences calculated between thermodynamic states with spring centers separated by 1 Å and sometimes 2 Å.

Machine learning predicts signaling peptides from protein star graphs.

problem Predicting signaling activity of proteins from molecular structure.
method Protein star graphs, S2SNet topological indices, Machine Learning (SVM-RFE, Laplacian kernel).
result Best model predicts 98.0% signaling pathways with AUROC 0.961.

New method designs antimicrobial peptides with high potency and low toxicity.

problem Designing potent antimicrobial drugs with low toxicity.
method CLaSS method using deep generative autoencoder and atomistic simulations.
result Design and synthesis of two novel AMPs with high potency and low toxicity.

GMVAE improves clustering in molecular simulations data.

problem Clustering metastable states in multi-basin free-energy landscapes.
method Gaussian mixture variational autoencoder (GMVAE) for dimensionality reduction and clustering.
result Enhanced clustering of metastable states compared to standard VAEs.

We attempt to set a mathematical foundation of immunology and amino acid chains. To measure the similarities of these chains, a kernel on strings is defined using only the sequence of the chains and a good amino acid substitution matrix (e.g. BLOSUM62). The kernel is used in learning machines to predict binding affinit…

2012-05-28abs ↗pdf ↗

Boosted GFlowNets improve exploration by sequentially training GFlowNets with residual rewards.

problem GFlowNets struggle to evenly explore reward landscapes, leading to poor coverage of high-reward areas.
method Sequential training of an ensemble of GFlowNets, each optimizing a residual reward.
result Boosted GFlowNets achieve better exploration and sample diversity on multimodal benchmarks and peptide design tasks.

Generative models accelerate molecular dynamics by four orders of magnitude.

problem Femtosecond time steps limit access to slow molecular processes.
method Deep generative modeling framework that accelerates sampling.
result Quantitative characterization of equilibrium ensembles and dynamical relaxation processes.

Many proteoforms - arising from alternative splicing, post-translational modifications (PTMs), or paralogous genes - have distinct biological functions, such as histone PTM proteoforms. However, their quantification by existing bottom-up mass-spectrometry (MS) methods is undermined by peptide-specific biases. To avoid …

2017-08-05abs ↗pdf ↗

We introduce a machine learning approach for extracting fine-grained representations of protein evolution from molecular dynamics datasets. Metastable switching linear dynamical systems extend standard switching models with a physically-inspired stability constraint. This constraint enables the learning of nuanced repr…

2016-10-05abs ↗pdf ↗

The paper introduces tests for high-dimensional independence using maximum and average distance correlations.

problem Testing independence in high-dimensional data.
method Characterizes consistency properties, compares test statistics, examines null distributions, and presents a fast chi-square-based procedure.
result The proposed tests are non-parametric and applicable to various metrics.

Paper explores using EEG for better speaker identification, even in noisy environments.

problem Speaker identification performance degrades in background noise.
method Uses EEG signals to enhance speaker identification systems, comparing with acoustic features.
result Speaker identification system using only EEG features outperforms one using only acoustic features in high background noise.

Bayesian models discover CVs for complex systems, enhancing sampling methods.

problem Limitations in modeling complex systems in biochemistry and materials science.
method Formulated CV discovery as a Bayesian inference problem, using deep learning and variational inference.
result Discovered CVs improve predictive ability for alanine dipeptide and ALA-15 peptides.

In the present paper we study interval identification systems of order three. We prove that the Rauzy induction preserves symmetry: for any symmetric interval identification system of order three after finitely many iterations of the Rauzy induction we always obtain a symmetric system. We also provide an example of sym…

2010-10-09abs ↗pdf ↗

Simplified identification methods for causal inference with arbitrary interventional distributions.

problem Estimating cause-effect relationships from data with experimental interventions.
method Using Single World Intervention Graphs and nested model factorization, we provide algorithms for identifying causal parameters from mixed observational and interventional distributions.
result Our algorithms are complete for certain types of interventional marginal distributions.

Study on identifying and inferring nonlinear dynamics on unknown networks.

problem Identifying network structure in nonlinear dynamic systems with unknown interactions.
method Showed network structure is not generically identified, requiring sufficient spectral heterogeneity. Developed necessary and sufficient conditions for identification and proposed a semiparametric estimator.
result Necessary and sufficient conditions for identification of network structure in nonlinear dynamic systems.

New findings on complexity limits in fixed budget bandit identification.

problem Determining the best possible error rate for fixed budget bandit identification.
method Analyzing the best non-adaptive sampling procedures and showing the existence of complexities.
result No fixed complexity for certain bandit identification tasks.

BoGA combines evolutionary search with Bayesian optimization for efficient protein design.

problem Designing novel proteins with specific characteristics is challenging due to sequence space complexity.
method BoGA integrates a genetic algorithm with Bayesian optimization to efficiently explore sequence space.
result BoGA accelerates discovery of high-confidence binders for diverse protein design objectives.

Review of automatic de-identification systems for EHR, highlighting challenges beyond accuracy.

problem Challenges in surrogate generation and patient privacy in de-identification of EHR.
method Comprehensive review of 18 recently published systems, focusing on accuracy and challenges.
result Despite accuracy improvements, challenges remain in surrogate generation and patient privacy.

Improves writer identification with unlabeled data and weighted label smoothing.

problem Offline writer identification requires labeled data, which is costly and time-consuming.
method Proposed a semi-supervised feature learning pipeline with weighted label smoothing regularization.
result Significantly improved baseline performance on writer identification datasets.

Bayesian methods reduce variance in subspace identification for small data sets.

problem High variance in traditional subspace identification methods for large models or small sample sizes.
method Investigation of Bayesian estimation solutions (regularized and shrinkage estimators) for subspace identification.
result Bayesian estimators reduce estimation risk by up to 40% compared to traditional methods.

A new algorithm identifies one of several nearly optimal arms in linear bandits.

problem Identifying one arm that is close to the best arm in linear bandits.
method Developed a procedure to adapt best-arm identification algorithms for ε\varepsilon-best-answer identification in transductive linear stochastic bandits.
result Proposed an asymptotically optimal algorithm for ε\varepsilon-best-answer identification.

A distributed system identification method for LTI systems using reverse experience replay.

problem Online system identification of LTI systems over multi-agent networks.
method DSGD-RER, a distributed variant of SGD-RER with backward updates.
result The estimation error decreases as the network size grows.

dynoGP uses deep Gaussian processes for dynamic system identification.

problem System identification for complex dynamical systems.
method Interconnecting linear dynamic GPs and static GPs to model dynamic and static nonlinearities.
result Demonstrates effectiveness of the approach using both simulated and real-world data.

ADSGD method speeds up model identification in sparse optimization.

problem Implicit model identification in sparse optimization problems.
method Accelerated Doubly Stochastic Gradient Method (ADSGD) for faster explicit model identification.
result ADSGD achieves faster explicit model identification and improved algorithm efficiency.

Proposes SPCA to incorporate structural constraints in model identification.

problem Model identification with partial structural knowledge.
method Structural Principal Component Analysis (SPCA) that leverages structural information.
result Demonstrates improved model estimates using synthetic and industrial data.