Gaussian process regression loses locality in high dimensions, affecting molecular energy surface fitting.
problem Loss of locality in high-dimensional Gaussian process regression.
method Analysis of Matern family kernels and multi-zeta basis functions.
result The property of locality disappears in high dimensions, impacting regression quality.
We introduce a machine learning model to predict atomization energies of a diverse set of organic molecules, based on nuclear charges and atomic positions only. The problem of solving the molecular Schrödinger equation is mapped onto a non-linear statistical regression problem of reduced complexity. Regression models a…
Multitask Gaussian process regression reduces data generation costs for molecular property prediction.
problem Data bottleneck in training surrogate models for molecular properties.
method Multitask Gaussian process regression over heterogeneous data sources (CC and DFT).
result Predicts at CC-level accuracy with over an order of magnitude reduction in data generation cost.
Improving in-context learning for latent space Bayesian optimization by adapting pretraining on molecular latent space.
problem Improving in-context learning for latent space Bayesian optimization.
method Adapting pretraining on molecular latent space.
result Achieving strong performance on held-out molecular optimization benchmarks.
We introduce multiscale invariant dictionaries to estimate quantum chemical energies of organic molecules, from training databases. Molecular energies are invariant to isometric atomic displacements, and are Lipschitz continuous to molecular deformations. Similarly to density functional theory (DFT), the molecule is re…
Deep neural network generates molecular structures close to equilibrium.
problem Discovery of atomistic systems with desirable properties.
method Autoregressive, convolutional deep neural network architecture.
result Model generates molecules close to equilibrium for C7O2H10 isomers.
Framework learns physics-informed continuum models from molecular data.
problem Discovering accurate and robust data-driven continuum models from molecular simulation data.
method Operator regression framework using neural networks in modal space with physical inductive biases.
result Learned operators generalize to unseen system characteristics.
Paper proposes efficient training for normalizing flows in Boltzmann generators.
problem Training normalizing flows for Boltzmann generators is computationally challenging and unstable.
method Regression Training of Normalizing Flows (RegFlow) using ℓ2-regression. result RegFlow enables efficient and stable training of normalizing flows for Boltzmann generators.
Solves the initial CV problem for molecular simulations using machine learning.
problem Selecting appropriate collective variables for enhancing sampling in molecular simulations.
method Data-driven approach inspired by supervised machine learning (SML).
result Various SML algorithms can be used as initial collective variables (SML_cv) for accelerated sampling.
A set of molecular descriptors whose length is independent of molecular size is developed for machine learning models that target thermodynamic and electronic properties of molecules. These features are evaluated by monitoring performance of kernel ridge regression models on well-studied data sets of small organic mole…
A new method uses IVA to fuse diverse molecular features for better machine learning predictions.
problem Challenges in selecting features for accurate molecular property prediction.
method Independent Vector Analysis (IVA) for fusing multiple molecular feature vectors into a single, compact set.
result Improved prediction performance of regression models for molecular properties.
Persistent homology provides a new, efficient molecular descriptor for protein dynamics.
problem Designing effective molecular descriptors for high-dimensional MD trajectories.
method Introduced masked Flood complex, a protein-tailored modification of simplicial complexes, for persistent homology.
result Persistent homology-based descriptors are competitive across protein dynamics tasks, including frame-level observable regression and MSM estimation.
All SMILES VAE learns molecule latent representations from SMILES strings.
problem Non-unique SMILES strings and high computational cost of graph convolutions hinder VAEs for molecular property optimization.
method Stacked recurrent neural networks encode multiple SMILES strings, pooling hidden representations, and attentional pooling builds a final latent representation.
result All SMILES VAE significantly surpasses state-of-the-art in molecular property optimization tasks.
Spherical CNNs tackle 3D data analysis, especially spherical images.
problem Learning problems involving spherical images, like omnidirectional vision and molecular regression.
method Defined spherical cross-correlation, developed a generalized FFT for efficient computation.
result Demonstrated spherical CNNs' effectiveness in 3D model recognition and atomization energy regression.
A new QSAR model selects relevant molecular descriptors for bioactivity prediction.
problem Redundant, noisy, and irrelevant descriptors in QSAR models.
method SPL-Logsum method using regularization and self-paced learning.
result SPL-Logsum method outperforms other methods in classification performance and model interpretability.
Study evaluates uncertainty quantification methods for molecular property prediction.
problem Uncertainty in neural models for molecular property prediction.
method Systematically evaluated several UQ methods on five benchmark datasets.
result No single method is unequivocally superior, and none provides reliable error ranking across datasets.
Simple linear models outperform complex BO methods in high dimensions.
problem Overcoming the curse of dimensionality in Bayesian optimization.
method Bayesian linear regression with linear kernels, applied to high-dimensional search spaces.
result Simple linear models match or outperform state-of-the-art BO methods in high-dimensional tasks.
New method pools graphs with edge features for molecular data.
problem Pooling graphs with edge features for molecular data.
method Proposes two types of pooling layers compatible with edge-feature graph-convolutional architecture.
result Significantly outperforms previous benchmarks on three out of four MoleculeNet datasets.
Paper develops a method to predict cancer patient survival using molecular profiles.
problem Accurately predicting cancer patient survival with complex survival-molecular profile relationships.
method Kernel Cox partially linear regression with a novel regularized garrotized kernel machine (RegGKM) method.
result The proposed method outperforms other methods in predicting survival accuracy.
GDML learns effective CG models from all-atom data.
problem Learning effective coarse-grained force fields efficiently.
method Ensemble learning with stratified sampling and GDML.
result GDML yields smaller free energy error than neural networks.
GNN-FiLM uses feature-wise linear modulation to improve graph neural networks.
problem Improving graph neural networks for better performance.
method Feature-wise linear modulation applied to target node representations in GNNs.
result GNN-FiLM outperforms baseline methods on a regression task for molecular graphs.
Sketching reduces data size for accurate spectral estimation.
problem Estimating spectral density from large simulation datasets.
method Sketching for dimensionality reduction and data compression.
result Sketching provides 90% accurate spectral density estimate with 10% data.
Unified theory linking atom-centered and message-passing models for molecular properties.
problem Combining atom-centered and message-passing models for accurate molecular property prediction.
method Generalizing ACDC framework to include multi-centered information, providing a complete linear basis for regression.
result Unified understanding of atom-centered and message-passing models, providing a coherent foundation.
LNK improves uncertainty estimation for molecular dynamics, reducing errors by up to 2.5 times.
problem Uncertainty estimation for molecular force fields to improve model reliability.
method LNK: Gaussian Process-based extension to GNNs addressing six desiderata.
result LNK reduces out-of-equilibrium detection errors by up to 2.5 times compared to existing methods.
MuML models predict molecular dipole moments using atomic partial charges and dipoles.
problem Predicting molecular dipole moments accurately and efficiently.
method Combining atomic partial charges and atomic dipoles within a physically inspired ML model.
result MuML models achieve excellent transferability and accuracy, approaching DFT results at a fraction of the computational cost.
A new method learns graph-level features for drug properties predication.
problem Predicting drug efficacy and toxicity from molecular graphs.
method Introducing a dummy super node connected to all nodes and modifying graph operations to learn graph-level features.
result The method improves molecular properties predication performance on MoleculeNet.
Deep generative models have been wildly successful at learning coherent latent representations for continuous data such as video and audio. However, generative modeling of discrete data such as arithmetic expressions and molecular structures still poses significant challenges. Crucially, state-of-the-art methods often …
A new model designs molecular latent vectors for drug discovery.
problem Designing effective molecular descriptors from molecular structures.
method Proposes a denoising diffusion probabilistic model (DDPM) for variational autoencoding molecular graphs.
result Demonstrates superior prediction performance and robustness compared to existing approaches.
DOCKSTRING simplifies docking simulations for better drug design benchmarks.
problem Lack of meaningful benchmarks for ligand design.
method Open-source Python package for docking scores, extensive dataset, and pharmaceutically-relevant tasks.
result Docking scores are more appropriate benchmarks than simple physicochemical properties.
MOSES benchmarks molecular generation models using a standardized dataset and metrics.
problem Unclear comparison and ranking of molecular generation models.
method Developed MOSES platform with training and testing datasets, metrics.
result Suggested MOSES results as reference for advancements in generative chemistry.
MoFlow generates chemically valid molecular graphs from latent representations.
problem Generating chemically valid molecular graphs from latent representations is challenging.
method MoFlow uses a flow-based approach with Glow for bond generation and a novel graph conditional flow for atom generation, ensuring chemical validity and efficiency.
result MoFlow achieves state-of-the-art performance in molecular graph generation and optimization.
Generative model creates drug-like molecules with multiple properties.
problem Designing molecules with multiple desired properties.
method Conditional Variational Autoencoder (VAE) in latent space control.
result Can generate drug-like molecules with five target properties.
Study compares GNNs and classical molecular featurisations for molecular property and cliff prediction.
problem Comparing GNNs and classical featurisations for molecular property and cliff prediction.
method Systematic exploration and comparison of PDVs, ECFPs, and GNNs; introduction of substructure pooling.
result Sort & Slice outperforms hash-based folding in ECFP vectorization.
Gemini uses inexpensive measurements to correct biases in expensive property evaluations.
problem Accurate estimation of materials properties using expensive measurements is hindered in scientific discovery campaigns.
method Gemini is a data-driven model that corrects systematic biases between property evaluation methods using inexpensive measurements.
result Gemini reduces the number of expensive evaluations needed for Bayesian optimization in materials discovery.
Novel RL approach for molecular design using quantum mechanics.
problem Existing RL methods for molecular design are limited in scope and reward function.
method Formulation in Cartesian coordinates, direct use of quantum mechanics for reward function, translation and rotation invariant state-action space.
result Agent efficiently learns to solve molecular design tasks from scratch.
Automates molecule design with a novel variational autoencoder.
problem Designing molecules based on specific chemical properties.
method Junction tree variational autoencoder generating tree-structured scaffolds and combining them into molecules.
result Significantly outperforms previous models on molecular generation and optimization tasks.
IGNN improves GNNs by maximizing edge-state transform mutual information.
problem Optimizing GNNs for better relational information.
method Variational information maximization to learn optimal transform parameters.
result IGNN achieves state-of-the-art performance on molecular graph tasks.
LSS learns molecular trajectories from MD data.
problem Limited integration time steps in MD simulations.
method Three deep learning networks for slow collective variables, dynamics, and configuration reconstruction.
result Generates ultra-long synthetic folding trajectories.
RL-VAE uses RL to decode molecular graphs from latent embeddings.
problem Efficiently decoding molecular graphs from latent embeddings.
method Repurposed simple graph generator for efficient decoding.
result Decoding molecular graphs from latent embeddings is possible with a simple graph generator.
Bayesian neural networks quantify uncertainties in molecular property predictions.
problem Poor predictions in molecular property predictions due to unreliable training data.
method Bayesian neural networks to estimate model-driven and data-driven uncertainties.
result Uncertainty quantification is necessary for reliable molecular applications.
Graph neural networks improve molecular property prediction.
problem Efficiently predicting molecular properties with high accuracy and scalability.
method Gated Graph Recursive Neural Networks (GGNN) with skip connections.
result GGNN achieves state-of-the-art performance on molecular property prediction benchmarks.
Machine learning improves molecular dynamics simulations by reducing costs and enhancing accuracy.
problem Inaccurate and costly molecular dynamics simulations hinder chemical system description.
method Adaptive sampling of reference data points and machine learning models for predicting molecular properties.
result Machine learning models can predict molecular dipole moments and infrared spectra accurately.
Generative model learns molecular geometry from graph representations.
problem Generating equilibrium states for molecular systems is computationally expensive.
method Probabilistic model based on Euclidean distance geometry.
result Generative model achieves state-of-the-art accuracy in molecular conformation generation.
Improved molecular property prediction using attention and gate mechanisms.
problem Predicting molecular properties from chemical data.
method Attention and gate mechanisms in graph convolutional networks.
result Improved prediction of molecular properties, including photovoltaic efficiency.
AniDS improves molecular force field modeling by learning anisotropic noise.
problem Molecular force field modeling suffers from oversimplified assumptions about atomic motions.
method AniDS introduces anisotropic noise generation for better modeling of directional and structural variability.
result AniDS outperforms existing methods on benchmarks, achieving significant improvements in force prediction accuracy.
Machine learning models simulate molecular spectra and reactions in solvents.
problem Accurate simulation of molecular spectra and reactions in solvent environments.
method Introduced FieldSchNet, a deep neural network for modeling molecular interactions with external fields.
result Demonstrated significant lowering of Claisen rearrangement reaction activation barrier using FieldSchNet.
BBRT improves molecular properties through iterative translation.
problem Optimizing molecular structures for improved biochemical properties.
method Iterative translation of molecules using a black box approach.
result Improvement in molecular properties with each iteration of the translator.
LaPool improves molecular graph representation learning by capturing interaction importance.
problem Lack of efficient intermediate pooling steps in GNNs leads to poor molecular substructure representation.
method LaPool is a novel, data-driven, and interpretable hierarchical graph pooling method that considers node features and graph structure.
result LaPool outperforms recent GNNs on molecular graph prediction and understanding tasks.