Paper introduces a graph-based approach for retrosynthesis prediction.
problem Predicting precursor molecules for target molecules in organic synthesis.
method Graph-based approach that predicts graph edits and synthons expansion.
result Achieves top-1 accuracy of 53.7%.
All SMILES VAE learns molecule latent representations from SMILES strings.
problem Non-unique SMILES strings and high computational cost of graph convolutions hinder VAEs for molecular property optimization.
method Stacked recurrent neural networks encode multiple SMILES strings, pooling hidden representations, and attentional pooling builds a final latent representation.
result All SMILES VAE significantly surpasses state-of-the-art in molecular property optimization tasks.
A new deep model generates molecules by fragments, improving validity and uniqueness.
problem Generating valid and unique molecules using deep learning.
method Develops a language model for molecular fragments, using frequency-based masking.
result Significantly outperforms other language model-based competitors in molecule generation.
ConfFlow uses transformer networks to generate molecular conformations efficiently.
problem Efficient generation of valid conformations for large molecules.
method Flow-based model using transformer networks that directly samples in coordinate space.
result ConfFlow improves accuracy by up to 40% for large molecule conformations.
Generative model for 3D molecules respects symmetry for targeted properties.
problem Infeasibility of exhaustive exploration in chemical space.
method Symmetry-adapted 3D point set generation neural network.
result Model generates molecules with desired properties like small HOMO-LUMO gap.
Benchmark proposes to assess molecule docking efficiency.
problem Lack of realistic benchmarks for measuring progress in drug design.
method Proposes a docking-based benchmark using SMINA software.
result Graph-based generative models fail to generate high-scoring molecules.
New neural network models predict molecular properties without 3D geometry, speeding up high-throughput screening.
problem Predicting molecular properties for large, complex molecules without computationally expensive 3D geometry.
method Message-passing neural networks trained with and without 3D structural information.
result Message-passing neural networks achieve similar accuracy to state-of-the-art methods without 3D geometry.
SMILES Transformer learns molecular fingerprints for drug discovery.
problem Poor performance of rule-based molecular fingerprints in shallow prediction models or small datasets.
method Unsupervised pre-training of a sequence-to-sequence language model on a corpus of SMILES.
result SMILES Transformer outperformed existing methods in small-data settings.
Adversarial autoencoder learns data on curved manifolds.
problem Representing data with non-Euclidean geometry.
method CCM adversarial autoencoder (CCM-AAE) trained on constant-curvature Riemannian manifolds.
result CCM-AAE outperforms other autoencoders on non-Euclidean data.
Improved molecular property prediction using updated neural message passing.
problem Predicting properties of molecules and materials accurately.
method Extended neural message passing model with edge update network.
result Superior prediction of formation energies and other properties on multiple datasets.
Generative model predicts synthesizable molecules from reactants.
problem Generating molecules with desirable properties doesn't guarantee synthesis feasibility.
method Proposes a realistic synthesis process model with reactant selection and reaction prediction.
result Model generates diverse, valid, and unique synthesizable molecules.
Although machine learning has been successfully used to propose novel molecules that satisfy desired properties, it is still challenging to explore a large chemical space efficiently. In this paper, we present a conditional molecular design method that facilitates generating new molecules with desired properties. The p…
Molecular "fingerprints" encoding structural information are the workhorse of cheminformatics and machine learning in drug discovery applications. However, fingerprint representations necessarily emphasize particular aspects of the molecular structure while ignoring others, rather than allowing the model to make data-d…
SILVR generates new molecules fitting protein binding sites.
problem Generating novel small molecule compounds for drug design.
method Selective Iterative Latent Variable Refinement (SILVR) for diffusion-based molecule generation.
result SILVR can generate new molecules similar in shape to original fragments without protein knowledge.
Paper proposes a method to design molecules with specific properties.
problem Designing molecules with desired chemical and biological properties.
method Energy-based model in latent space, SGDS algorithm for gradual distribution shifting.
result Method achieves strong performances on various molecule design tasks.
In de novo drug design, computational strategies are used to generate novel molecules with good affinity to the desired biological target. In this work, we show that recurrent neural networks can be trained as generative models for molecular structures, similar to statistical language models in natural language process…
GCDM generates valid large 3D molecules and optimizes existing molecules.
problem Lack of geometric properties in 3D molecule generation models.
method Introduces Geometry-Complete Diffusion Model (GCDM) using equivariant GNNs.
result Significantly outperforms existing models in 3D molecule generation and optimization.
Generative model creates new molecules retaining a scaffold with certainty.
problem Designing new molecules with a specific scaffold.
method Generative model that extends scaffold graph by adding vertices and edges.
result Model can generate novel molecules with high validity, uniqueness, and novelty.
Improved RL model for fragment-based molecule generation.
problem Generating molecules with high docking scores.
method Thorough reproduction, scrutiny, and improvement of the FREED model.
result The improved model produces molecules with superior docking scores.
Generative model learns to create molecules with multiple properties using interpretable substructures.
problem Creating molecules with multiple chemical properties is challenging.
method Compose molecules from substructures identified as responsible for each property, using graph generative models.
result Significant improvements in accuracy, diversity, and novelty of generated compounds over state-of-the-art baselines.
CORE optimizes molecules by copying or generating substructures, improving accuracy.
problem Inaccurate substructure prediction in molecule optimization.
method Copy & Refine (CORE) strategy combining scaffolding tree generation and adversarial training.
result Significant improvement in various molecule optimization metrics.
A new method improves molecule generation accuracy and efficiency.
problem Posterior collapse in VAEs for molecule sequence generation.
method Re-balancing reconstruction loss to avoid posterior collapse.
result Our method achieves state-of-the-art reconstruction accuracy and competitive validity.
A new model designs molecules with desired properties.
problem Finding molecules with optimal chemical or biological properties.
method A latent prompt Transformer model with three components: latent vector, molecule generation, and property prediction.
result The model achieves state-of-the-art performance on molecule design tasks.
Equivariant diffusion model generates 3D molecules efficiently.
problem Generating high-quality 3D molecules efficiently.
method Equivariant Diffusion Model (EDM) that operates on atom coordinates and types.
result Significantly outperforms previous methods in molecule quality and training efficiency.
Generative neural network designs novel 3D molecules with specified properties.
problem Designing molecules with desired properties in chemistry.
method Conditional generative neural network for 3D molecular structures.
result Demonstrated utility in generating novel molecules with specified motifs or composition.
ALMGIG uses adversarial learning to generate and infer novel molecules efficiently.
problem Efficiently generating and inferring novel molecules using graph representations.
method Adversarial learning framework that avoids explicit graph isomorphism, using cycle-consistency loss and multi-graph Graph Isomorphism Network.
result ALMGIG more accurately learns the distribution over the space of molecules and efficiently searches the molecular space.
Designing a new drug is a lengthy and expensive process. As the space of potential molecules is very large (10^23-10^60), a common technique during drug discovery is to start from a molecule which already has some of the desired properties. An interdisciplinary team of scientists generates hypothesis about the required…
Modof-pipe optimizes molecules by modifying a single site, outperforming state-of-the-art methods.
problem Improving drug candidates' properties through chemical modification.
method Deep generative model Modof over molecular graphs for molecule optimization.
result Modof-pipe achieves significant improvements in octanol-water partition coefficient and molecule similarity constraints.
New model generates larger molecules more effectively.
problem Previous graph generation techniques struggle with larger molecules.
method Hierarchical graph encoder-decoder using structural motifs.
result Model significantly outperforms previous baselines on molecule generation tasks.
Two-step process generates molecules from latent vectors.
problem Generating valid molecules from latent representations.
method Two-step decoding: first formula, then bonds.
result Highest reconstruction rate of 90.5%.
VecMol generates 3D molecules as continuous vector fields, overcoming modality and geometry constraints.
problem Challenges in generating 3D molecules, especially in drug discovery and materials science.
method VecMol reimagines molecular representation by modeling 3D molecules as continuous vector fields over Euclidean space, parameterized by a neural field and generated using a latent diffusion model.
result Vector-field-based representations show promise for 3D molecular generation, validated on benchmarks.
SELFIES solves molecular string representation weaknesses for material design.
problem Weaknesses in SMILES for representing valid molecules in material design.
method Introducing SELFIES, a 100% robust string-based molecular representation.
result SELFIES strings correspond to valid molecules, allowing arbitrary machine learning applications.
The new wave of successful generative models in machine learning has increased the interest in deep learning driven de novo drug design. However, assessing the performance of such generative models is notoriously difficult. Metrics that are typically used to assess the performance of such generative models are the perc…
MHG-VAE achieves 100% valid molecules with simpler architecture.
problem Generating valid molecules and evaluating properties efficiently.
method MHG-VAE uses molecular hypergraph grammar to guide a single VAE.
result 100% validity achieved with simpler architecture.
In this study, we intend to solve a mutual information problem in interacting molecules of any type, such as proteins, nucleic acids, and small molecules. Using machine learning techniques, we accurately predict pairwise interactions, which can be of medical and biological importance. Graphs are are useful in this prob…
CogMol designs novel drug-like molecules for SARS-CoV-2 targets.
problem Designing efficient drugs for novel viral proteins.
method End-to-end framework combining VAE, controlled sampling, and predictors.
result Highly selective and affinity molecules for SARS-CoV-2 targets.
Generative models encode and decode 3D crystal structures from a large dataset.
problem Challenges in encoding and decoding 3D crystal structures from large datasets.
method Training two neural networks on a dataset of over 120,000 crystal structures to encode and decode 3D atom positions.
result Ability to generate compressed, continuous latent space representations and decode molecules accurately.
Automates molecule design with simpler SMILES generation and reinforcement learning.
problem Designing molecules with specific chemical properties.
method Combines context-free grammar for SMILES strings and reinforcement learning with a Transformer model.
result Significantly reduces model steps per atom and beats previous baselines.
A new framework optimizes molecules using deep reinforcement learning.
problem Optimizing molecules while maintaining chemical validity and drug-likeness.
method Combining deep reinforcement learning with domain knowledge of chemistry, MolDQN directly modifies molecules.
result MolDQN achieves optimization of molecules without bias from pre-training datasets.
ChemBO optimizes small organic molecules for synthesis and desired properties.
problem Designing and optimizing new organic molecules for specific properties.
method Bayesian optimization framework that considers synthesizability constraints.
result ChemBO generates synthesizable candidates efficiently and effectively.
AI and HPC help screen millions of molecules for SARS-CoV-2 treatments.
problem Finding effective treatments for SARS-CoV-2.
method AI and HPC enable screening of large molecule datasets.
result Data release of 23 datasets with 4.2 billion molecules.
Mol-CycleGAN generates optimized molecules with similar structure.
problem Designing molecules with desired properties is challenging.
method CycleGAN-based model that generates optimized compounds with high structural similarity.
result Significantly outperforms previous results in optimizing penalized logP of drug-like molecules.
Predicting the biological function of molecules, be it proteins or drug-like compounds, from their atomic structure is an important and long-standing problem. Function is dictated by structure, since it is by spatial interactions that molecules interact with each other, both in terms of steric complementarity, as well …
MoleculeSTM learns from molecule structures and texts for better drug design.
problem Lack of integration between chemical structures and textual knowledge in AI drug discovery.
method Jointly learns chemical structures and texts via contrastive learning, using a large dataset.
result MoleculeSTM achieves state-of-the-art performance in zero-shot tasks like structure-text retrieval and molecule editing.
BBRT improves molecular properties through iterative translation.
problem Optimizing molecular structures for improved biochemical properties.
method Iterative translation of molecules using a black box approach.
result Improvement in molecular properties with each iteration of the translator.
Generative model creates drug-like molecules with multiple properties.
problem Designing molecules with multiple desired properties.
method Conditional Variational Autoencoder (VAE) in latent space control.
result Can generate drug-like molecules with five target properties.
Model predicts stable molecules with AI and physics constraints.
problem Designing stable molecules with limited data.
method Graph Scattering Variational Autoencoder with physical constraints.
result Model generates stable molecules with desired properties.
New RL formulation for maximizing maximum reward in molecule generation.
problem Traditional RL frameworks do not fit real-world applications like drug discovery.
method Formulated a new objective function to maximize maximum reward, derived Bellman equation, introduced operators, and proved convergence.
result Achieved state-of-the-art results in molecule generation.