OLCS-Ranker improves peptide identification accuracy and speed on hard datasets.
problem Efficiently identifying peptides from MS/MS data, especially on hard datasets with many false positives.
method Cost-sensitive online learning model and iterative online learning algorithm.
result OLCS-Ranker outperforms existing methods in accuracy and speed on large datasets.
Improved peptide identification from mass spectrometry data.
problem Lack of large ground truth datasets for protein identification.
method Deep neural networks trained on imperfect hand-coded models.
result 43% improvement over standard matching methods.
Liquid chromatography coupled with tandem mass spectrometry, also known as shotgun proteomics, is a widely-used high-throughput technology for identifying proteins in complex biological samples. Analysis of the tens of thousands of fragmentation spectra produced by a typical shotgun proteomics experiment begins by assi…
Generative model gradients enhance MS/MS peptide identification.
problem Improving peptide identification from MS/MS spectra.
method Leverage log-likelihood gradients of generative models in a kernel-based classifier.
result Fisher kernel outperforms other methods on MS/MS datasets.
Major histocompatibility complex class two (MHC-II) molecules are trans-membrane proteins and key components of the cellular immune system. Upon recognition of foreign peptides expressed on the MHC-II binding groove, helper T cells mount an immune response against invading pathogens. Therefore, mechanistic identificati…
PepCVAE designs novel antimicrobial peptides using a semi-supervised VAE.
problem Designing novel antimicrobial peptides for next-generation antimicrobial resistance solutions.
method Semi-supervised variational autoencoder (VAE) model that learns latent and antimicrobial attribute spaces from unlabeled and labeled data.
result PepCVAE generates novel AMPs with higher long-range diversity and closer to biological peptide distribution.
AMP0 predicts antimicrobial peptides targeting specific microbes.
problem Low-throughput screening of antimicrobial peptides.
method Zero-shot and few-shot machine learning.
result AMP0 can predict antimicrobial activity against specific microbes.
Study improves peptide design efficiency using active and meta-learning.
problem Designing novel functional peptides is inefficient due to low throughput and lack of data.
method Investigated active learning and meta-learning for optimizing peptide design experiments.
result Meta-learning improved average accuracy, but neither method outperformed random choice.
DeepNovoV2 improves de novo peptide sequencing from mass spectrometry data.
problem De novo peptide sequencing from mass spectrometry data for personalized cancer vaccines.
method DeepNovoV2 combines T-Net and recurrent neural networks for end-to-end training and prediction.
result DeepNovoV2 achieves 13.01-23.95\% higher accuracy than previous methods.
Bayesian network models are finding success in characterizing enzyme-catalyzed reactions, slow conformational changes, predicting enzyme inhibition, and genomics. In this work, we apply them to statistical modeling of peptides by simultaneously identifying amino acid sequence motifs and using a motif-based model to cla…
Improved protein identification in mass spectrometry data.
problem Expanding peptide scoring capabilities in tandem mass spectrometry.
method Deriving concave emission distributions for dynamic Bayesian networks.
result Efficiently learned scoring function outperforms state-of-the-art.
We propose a specialized string kernel for small bio-molecules, peptides and pseudo-sequences of binding interfaces. The kernel incorporates physico-chemical properties of amino acids and elegantly generalize eight kernels, such as the Oligo, the Weighted Degree, the Blended Spectrum, and the Radial Basis Function. We …
This paper presents regression models obtained from a process of blind prediction of peptide binding affinity from provided descriptors for several distinct datasets as part of the 2006 Comparative Evaluation of Prediction Algorithms (COEPRA) contest. This paper finds that kernel partial least squares, a nonlinear part…
Deep neural networks improve free energy calculations for peptide conformations.
problem Challenges in developing suitable mappings for free energy perturbation.
method Adapted machine learning approach to train deep neural networks for mapping between Boltzmann distributions.
result Accurate free energy differences calculated between thermodynamic states with spring centers separated by 1 Å and sometimes 2 Å.
Machine learning predicts signaling peptides from protein star graphs.
problem Predicting signaling activity of proteins from molecular structure.
method Protein star graphs, S2SNet topological indices, Machine Learning (SVM-RFE, Laplacian kernel).
result Best model predicts 98.0% signaling pathways with AUROC 0.961.
New method designs antimicrobial peptides with high potency and low toxicity.
problem Designing potent antimicrobial drugs with low toxicity.
method CLaSS method using deep generative autoencoder and atomistic simulations.
result Design and synthesis of two novel AMPs with high potency and low toxicity.
GMVAE improves clustering in molecular simulations data.
problem Clustering metastable states in multi-basin free-energy landscapes.
method Gaussian mixture variational autoencoder (GMVAE) for dimensionality reduction and clustering.
result Enhanced clustering of metastable states compared to standard VAEs.
We attempt to set a mathematical foundation of immunology and amino acid chains. To measure the similarities of these chains, a kernel on strings is defined using only the sequence of the chains and a good amino acid substitution matrix (e.g. BLOSUM62). The kernel is used in learning machines to predict binding affinit…
Boosted GFlowNets improve exploration by sequentially training GFlowNets with residual rewards.
problem GFlowNets struggle to evenly explore reward landscapes, leading to poor coverage of high-reward areas.
method Sequential training of an ensemble of GFlowNets, each optimizing a residual reward.
result Boosted GFlowNets achieve better exploration and sample diversity on multimodal benchmarks and peptide design tasks.
Generative models accelerate molecular dynamics by four orders of magnitude.
problem Femtosecond time steps limit access to slow molecular processes.
method Deep generative modeling framework that accelerates sampling.
result Quantitative characterization of equilibrium ensembles and dynamical relaxation processes.
A novel method clusters protein conformations from MD simulations.
problem Clustering long MD protein dynamics for identifying states and behavior.
method Adversarial Autoencoder (AAE) for conformation clustering.
result Identifies many salient features of the folding process.
Many proteoforms - arising from alternative splicing, post-translational modifications (PTMs), or paralogous genes - have distinct biological functions, such as histone PTM proteoforms. However, their quantification by existing bottom-up mass-spectrometry (MS) methods is undermined by peptide-specific biases. To avoid …
We introduce a machine learning approach for extracting fine-grained representations of protein evolution from molecular dynamics datasets. Metastable switching linear dynamical systems extend standard switching models with a physically-inspired stability constraint. This constraint enables the learning of nuanced repr…
The paper introduces tests for high-dimensional independence using maximum and average distance correlations.
problem Testing independence in high-dimensional data.
method Characterizes consistency properties, compares test statistics, examines null distributions, and presents a fast chi-square-based procedure.
result The proposed tests are non-parametric and applicable to various metrics.
Paper explores using EEG for better speaker identification, even in noisy environments.
problem Speaker identification performance degrades in background noise.
method Uses EEG signals to enhance speaker identification systems, comparing with acoustic features.
result Speaker identification system using only EEG features outperforms one using only acoustic features in high background noise.
Bayesian models discover CVs for complex systems, enhancing sampling methods.
problem Limitations in modeling complex systems in biochemistry and materials science.
method Formulated CV discovery as a Bayesian inference problem, using deep learning and variational inference.
result Discovered CVs improve predictive ability for alanine dipeptide and ALA-15 peptides.
In the present paper we study interval identification systems of order three. We prove that the Rauzy induction preserves symmetry: for any symmetric interval identification system of order three after finitely many iterations of the Rauzy induction we always obtain a symmetric system. We also provide an example of sym…
Improved driver identification accuracy using steering wheel data.
problem Accurately identifying drivers based on naturalistic driving behavior.
method Novel approach for window length parameter design, leveraging GRUs neural network.
result Increased driver identification accuracy from under 15% to over 65%.
Simplified identification methods for causal inference with arbitrary interventional distributions.
problem Estimating cause-effect relationships from data with experimental interventions.
method Using Single World Intervention Graphs and nested model factorization, we provide algorithms for identifying causal parameters from mixed observational and interventional distributions.
result Our algorithms are complete for certain types of interventional marginal distributions.
Cyclic coordinate descent identifies models in finite time and converges linearly.
problem Model identification in composite nonsmooth optimization problems.
method Cyclic coordinate descent for a wide class of functions.
result Explicit local linear convergence rates for coordinate descent.
WiPIN uses Wi-Fi signals to identify people without requiring them to walk.
problem Identification requires walking and is unreliable with many users.
method Extracts body information from Wi-Fi signals without user movement.
result Achieves 92% accuracy with 30 users, robust to various settings.
Study on identifying and inferring nonlinear dynamics on unknown networks.
problem Identifying network structure in nonlinear dynamic systems with unknown interactions.
method Showed network structure is not generically identified, requiring sufficient spectral heterogeneity. Developed necessary and sufficient conditions for identification and proposed a semiparametric estimator.
result Necessary and sufficient conditions for identification of network structure in nonlinear dynamic systems.
New findings on complexity limits in fixed budget bandit identification.
problem Determining the best possible error rate for fixed budget bandit identification.
method Analyzing the best non-adaptive sampling procedures and showing the existence of complexities.
result No fixed complexity for certain bandit identification tasks.
BoGA combines evolutionary search with Bayesian optimization for efficient protein design.
problem Designing novel proteins with specific characteristics is challenging due to sequence space complexity.
method BoGA integrates a genetic algorithm with Bayesian optimization to efficiently explore sequence space.
result BoGA accelerates discovery of high-confidence binders for diverse protein design objectives.
Smooth flows for physical systems with smooth energies and forces.
problem Smooth energies for physical simulations and force computation.
method Smooth mixture transformations on compact intervals and hypertori, using root-finding and the inverse function theorem.
result Smooth flows allow training by force matching and use as molecular dynamics potentials.
Deep learning's convolutional networks improve system identification.
problem Nonlinear system identification problems.
method Exploration of relationships between TCN and Volterra series/block-oriented models.
result TCN outperforms traditional models in sequence modeling tasks.
Review of automatic de-identification systems for EHR, highlighting challenges beyond accuracy.
problem Challenges in surrogate generation and patient privacy in de-identification of EHR.
method Comprehensive review of 18 recently published systems, focusing on accuracy and challenges.
result Despite accuracy improvements, challenges remain in surrogate generation and patient privacy.
Improves writer identification with unlabeled data and weighted label smoothing.
problem Offline writer identification requires labeled data, which is costly and time-consuming.
method Proposed a semi-supervised feature learning pipeline with weighted label smoothing regularization.
result Significantly improved baseline performance on writer identification datasets.
Bayesian methods reduce variance in subspace identification for small data sets.
problem High variance in traditional subspace identification methods for large models or small sample sizes.
method Investigation of Bayesian estimation solutions (regularized and shrinkage estimators) for subspace identification.
result Bayesian estimators reduce estimation risk by up to 40% compared to traditional methods.
Equation discovery method reconstructs model structure and parameters from data.
problem Nonlinear system identification challenges.
method Two interlaced parts: model structure identification and parameter estimation.
result Equation discovery method successfully reconstructs model structure and parameters from data.
A tutorial on non-asymptotic system identification methods.
problem Identifying system parameters in linear models.
method Covering technique, Hanson-Wright Inequality, method of self-normalized martingales.
result Streamlined proofs of least-squares based estimator performance.
A new algorithm identifies one of several nearly optimal arms in linear bandits.
problem Identifying one arm that is close to the best arm in linear bandits.
method Developed a procedure to adapt best-arm identification algorithms for ε-best-answer identification in transductive linear stochastic bandits. result Proposed an asymptotically optimal algorithm for ε-best-answer identification. Online algorithm identifies PDEs from noisy data snapshots.
problem Identifying PDEs from sequential solution snapshots.
method Combines weak-form discretization with online proximal gradient descent.
result Efficiently identifies and tracks systems with time-varying coefficients.
A distributed system identification method for LTI systems using reverse experience replay.
problem Online system identification of LTI systems over multi-agent networks.
method DSGD-RER, a distributed variant of SGD-RER with backward updates.
result The estimation error decreases as the network size grows.
Tutorial on using concentration inequalities for linear system identification.
problem Learning state-space parameters of linear systems.
method Large-deviations and self-normalized martingales.
result Data-dependent and independent bounds on learning rate.
dynoGP uses deep Gaussian processes for dynamic system identification.
problem System identification for complex dynamical systems.
method Interconnecting linear dynamic GPs and static GPs to model dynamic and static nonlinearities.
result Demonstrates effectiveness of the approach using both simulated and real-world data.
ADSGD method speeds up model identification in sparse optimization.
problem Implicit model identification in sparse optimization problems.
method Accelerated Doubly Stochastic Gradient Method (ADSGD) for faster explicit model identification.
result ADSGD achieves faster explicit model identification and improved algorithm efficiency.
Proposes SPCA to incorporate structural constraints in model identification.
problem Model identification with partial structural knowledge.
method Structural Principal Component Analysis (SPCA) that leverages structural information.
result Demonstrates improved model estimates using synthetic and industrial data.