Physics-informed machine learning models improve biomolecular system simulations.
problem Modeling unresolved interactions beyond classical force fields.
method Physics-informed neural networks and operator learning.
result Accurate, mechanistic, generalizable models for long-timescale kinetics.
Tabular in-context learners perform well on biomolecular tasks, but performance depends on the representation used.
problem Predicting biomolecular properties from limited labeled data.
method Evaluating tabular in-context learners on protein fitness regression and small-molecule classification tasks.
result Tabular in-context learners are competitive for protein fitness regression but not for small-molecule classification.
Review of mathematical representations for biomolecular data.
problem Complexity and high dimensionality of biomolecular datasets hinder ML applications.
method Developed low-dimensional and scalable mathematical representations using algebraic topology, differential geometry, and graph theory.
result Mathematical representations improve protein-ligand binding predictions and other biomolecular applications.
Design of experiments improves validation of biomolecular networks.
problem Efficiently validate non-machine learning designed biomolecular networks.
method Use Gaussian processes and Bayesian optimization to select experimental points.
result Developed a stopping criterion based on discrepancy metric and uncertainty.
Optimizes biomolecular simulations by ranking adaptive sampling policies.
problem Efficiently sampling biomolecular systems to capture complex dynamical behaviors.
method Metric-driven ranking of adaptive sampling policies to identify the optimal policy for each round.
result Different adaptive sampling policies lead to faster convergence and improved sampling performance.
As deep Variational Auto-Encoder (VAE) frameworks become more widely used for modeling biomolecular simulation data, we emphasize the capability of the VAE architecture to concurrently maximize the timescale of the latent space while inferring a reduced coordinate, which assists in finding slow processes as according t…
Pipeline learns topological features for protein stability prediction.
problem Predicting protein stability using topological features.
method Data-driven method to learn topological features, comparing with expert features.
result Topological features achieve 92%-99% of SME-based models' performance.
Equivariant networks improve geometric prediction without scalar approximations.
problem Efficiently predicting geometric tensors in real-world scenarios.
method Equivariant networks for geometric prediction.
result Equivariant networks can generalize to unseen systems for geometric prediction.
A deep learning model organizes RNA graphs to reveal folding patterns and properties.
problem Organizing and understanding the complex folding patterns of RNA secondary structures.
method Geometric scattering autoencoder (GSAE) network for learning graph embeddings.
result GSAE accurately reflects bistable RNA structures and can sample new folding trajectories.
Bayesian model clusters diverse 'omics data for disease subtyping.
problem Clustering diverse 'omics datasets conflates multiple structures.
method Multi-view Bayesian mixture model with semi-supervised learning.
result Identifies distinct clusters of patients for stratified medicine.
Method scales up ML science by measuring multiple molecules at once.
problem Scaling up ML-driven science with wet lab experiments.
method Neural extension of compressed sensing for function space.
result Proves orders-of-magnitude gains in information density.
With the advent of deep generative models in computational chemistry, in silico anticancer drug design has undergone an unprecedented transformation. While state-of-the-art deep learning approaches have shown potential in generating compounds with desired chemical properties, they disregard the genetic profile and prop…
This review explores the use of machine learning in discovering collective variables for biomolecular dynamics.
problem Understanding the conformational dynamics and molecular recognition in biomolecules.
method Statistical analysis of high-dimensional spatiotemporal data generated from molecular dynamics simulations.
result Machine learning algorithms can be used to discover abstract collective variables that describe biomolecular dynamics.
Enhanced diffusion sampling improves rare event sampling in biomolecular simulations.
problem Efficiently sampling rare transition events in biomolecular systems.
method Quantitative steering protocols to generate biased ensembles and exact reweighting.
result Fast, accurate, and scalable estimation of equilibrium properties.
Enhanced diffusion sampling tackles rare event sampling in biomolecular simulations.
problem Efficiently sampling rare transition events in biomolecular simulations.
method Quantitative steering protocols to generate biased ensembles, followed by exact reweighting.
result Fast, accurate, and scalable estimation of equilibrium properties for folding free energies.
Cryo-electron microscopy (cryo-EM) is an emerging experimental method to characterize the structure of large biomolecular assemblies. Single particle cryo-EM records 2D images (so-called micrographs) of projections of the three-dimensional particle, which need to be processed to obtain the three-dimensional reconstruct…
Develops an MS-inspired algorithm for regression mode finding and space partitioning.
problem Finding local modes of regression functions and partitioning input space.
method Mean-shift-inspired algorithm for iterative gradient ascent.
result Proves convergence and rates of convergence for estimated local modes.
There is an increasing demand for computing the relevant structures, equilibria and long-timescale kinetics of biomolecular processes, such as protein-drug binding, from high-throughput molecular dynamics simulations. Current methods employ transformation of simulated coordinates into structural features, dimension red…
Develops methods for spectral estimation and rare-event prediction in complex systems.
problem Challenges in understanding dynamics in complex systems with many degrees of freedom.
method Inexact iterative numerical linear algebra methods for spectral estimation and rare-event prediction.
result Demonstrates methods on low-dimensional and high-dimensional models, showing their effectiveness.
Diffusion maps are a nonlinear manifold learning technique based on harmonic analysis of a diffusion process over the data. Out-of-sample extensions with computational complexity O(N), where N is the number of points comprising the manifold, frustrate applications to online learning applications requiring…
We address the problem of analyzing sets of noisy time-varying signals that all report on the same process but confound straightforward analyses due to complex inter-signal heterogeneities and measurement artifacts. In particular we consider single-molecule experiments which indirectly measure the distinct steps in a b…
Cryo-em images are found to be low-dimensional.
problem Understanding the geometric structure of cryo-em data.
method Applied manifold learning techniques to CryoSBI representations.
result Cryo-em data inherently populate low-dimensional manifolds.
Macromolecular and biomolecular folding landscapes typically contain high free energy barriers that impede efficient sampling of configurational space by standard molecular dynamics simulation. Biased sampling can artificially drive the simulation along pre-specified collective variables (CVs), but success depends crit…
Paper uses transfer learning and Bayesian optimization to reduce DNA sequence design experiments.
problem Designing many similar DNA sequences for specific applications is expensive and time-consuming.
method Combines transfer learning with Bayesian optimization to reduce experiment count.
result Total number of experiments can be significantly reduced by sharing information between tasks.
Adapts safe policies for exploration in high-risk settings.
problem Balancing safety and exploration in high-risk environments.
method Uses conformal calibration on a safe reference policy to determine aggressive action limits.
result Safe exploration improves performance without requiring model class identification or hyperparameter tuning.
We introduce a novel class of localized atomic environment representations, based upon the Coulomb matrix. By combining these functions with the Gaussian approximation potential approach, we present LC-GAP, a new system for generating atomic potentials through machine learning (ML). Tests on the QM7, QM7b and GDB9 biom…
Extend CPS to non-exchangeable settings with observation-specific permutation weights
problem Calibrated predictive bands under distributional shifts
method Encoding distributional shifts through observation-specific permutation weights
result Shift-aware predictive systems remain valid
This manuscript contributes a general and practical framework for casting a Markov process model of a system at equilibrium as a structural causal model, and carrying out counterfactual inference. Markov processes mathematically describe the mechanisms in the system, and predict the system's equilibrium behavior upon i…
Enhanced coloring invariant distinguishes folded molecular chain topologies.
problem Apparent indistinguishability of folded chain topologies using current coloring invariants.
method Introduced Boltzmann weights to improve the resolving power of quandle colorings.
result Improved resolution in distinguishing folded chain topologies.
K-Models clusters functional data with ordinal constraints for better interpretability.
problem Challenges in extracting meaningful insights from functional data due to lack of interpretability.
method Integrates ordinal constraints into clustering to improve interpretability and structure identification.
result Enhances interpretability of clustering results while maintaining performance.
New method for causal discovery in high dimensions with confounder blanket assumption.
problem Inferring causal relationships from observational data in high dimensions.
method Relaxes parametric restrictions and sparsity constraints, focusing on confounder blanket.
result Provable sound and complete structure learning algorithm with finite sample error control.
GDML learns effective CG models from all-atom data.
problem Learning effective coarse-grained force fields efficiently.
method Ensemble learning with stratified sampling and GDML.
result GDML yields smaller free energy error than neural networks.
This abstract reviews recent methods for predicting protein-ligand binding affinity.
problem Predicting protein-ligand binding affinity for various applications in life sciences.
method Traditional and deep learning models for binding affinity prediction.
result Improved predictive performance of AI-driven models.
Study limits of circadian synchronization under different light signals.
problem Disruption of circadian rhythms due to misalignment with external light signals.
method Matrix-free approach for locating periodic steady states, numerical continuation, bifurcation diagrams, unsupervised learning.
result Limits of circadian synchronization to external light signals of different frequency and duty cycle.
Energy-based diffusion models improve molecular sampling and simulation.
problem Inconsistency between diffusion model scores and equilibrium distributions.
method Fokker-Planck regularization to enforce consistency.
result Improved consistency and efficient sampling of biomolecular systems.
The modeling of atomistic biomolecular simulations using kinetic models such as Markov state models (MSMs) has had many notable algorithmic advances in recent years. The variational principle has opened the door for a nearly fully automated toolkit for selecting models that predict the long-time kinetics from molecular…
Continuous-depth Evoformer reduces protein folding prediction time and resource usage.
problem Efficient protein structure prediction with reduced computational costs.
method Continuous-depth formulation of Evoformer using Neural Ordinary Differential Equations (Neural ODEs).
result The continuous-time Evoformer achieves constant memory cost and improved efficiency.
In this thesis we present the novel semi-supervised network-based algorithm P-Net, which is able to rank and classify patients with respect to a specific phenotype or clinical outcome under study. The peculiar and innovative characteristic of this method is that it builds a network of samples/patients, where the nodes …
Unified framework for sampling and approximating high-dimensional energy landscapes.
problem Sampling and approximating complex energy landscapes in physical systems with constraints and energy barriers.
method Formulates a minimax optimization problem that jointly adapts surrogate approximation and adaptive sampling.
result Demonstrates effectiveness in biomolecular systems with up to 30 collective variables.
Introduce a thermodynamically informed, temperature-transferable MLCG framework for proteins.
problem Temperature transferability of MLCG models for proteins.
method Explicit decomposition of CG potential into energetic and entropic components.
result Reproduces temperature-dependent quantities like heat capacity.
Machine learning generates coarse-grained force fields for molecular dynamics.
problem Creating thermodynamically consistent coarse-grained models for larger systems.
method Hybrid architecture using graph neural networks to learn molecular features.
result Framework reproduces thermodynamics for small biomolecular systems.
Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…
A new method for handling imbalanced big data using ensembles and smart data.
problem Imbalanced data distribution in big data scenarios.
method Smart Data driven Decision Trees Ensemble (SD_DeTE) methodology.
result SD_DeTE outperforms Random Forest in handling imbalanced binary classification problems in big data.
Prevents sensitive data generation in diffusion models using labeled and unlabeled data.
problem Generating sensitive data in diffusion models using unlabeled data.
method Positive-Unlabeled Diffusion Models, approximating ELBO with labeled and unlabeled data.
result Prevents the generation of sensitive data without compromising image quality.
Study reveals Data Shapley's inconsistent performance in data selection tasks.
problem Inconsistency of Data Shapley's performance in data selection across different settings.
method Hypothesis testing framework and identification of utility functions.
result Data Shapley's performance is no better than random selection without specific constraints.
PRRO generates synthetic tabular data that improves SL performance and class distribution.
problem Low SL utility of synthetic data due to class imbalance and overlooked data relationships.
method Data pruning and column reordering to optimize SL utility.
result Synthetic data generated with PRRO enhances predictive performance and class distribution.
Defines data science as a natural ecosystem with challenges and missions.
problem Challenges and missions in data science due to 5D complexities and data life cycle phases.
method Systemic and data-centric view of data science as a fusion of data universe and its challenges, formalizing a general-purpose architecture.
result Essential data science as a natural ecosystem integrating specific disciplines and high-impact applications.
Synthetic data enhances analytics but requires careful volume management.
problem Accuracy of statistical methods on synthetic data vs. raw data.
method Synthetic Data Generation for Analytics framework using tabular diffusion models.
result Error rate decreases with more synthetic data but may stabilize or increase.