New method explains complex fuel compound classifications.
problem Understanding complex quantitative structure-activity relationship models.
method Locally Interpretable Machine-Agnostic Explanations (LIME) applied to 2-D chemical structures.
result Replicates chemical intuition, allowing direct acceptance/rejection of decisions.
Optimizes molecular generation for chemist preferences.
problem Models lack inherent preferences for chemist-desired structures.
method Fine-tuning with Direct Preference Optimization.
result Approach is simple, efficient, and highly effective.
Two regularization techniques improve GCNN explainability and preference from chemists.
problem Difficulty in rationalizing molecular graph neural network predictions.
method Batch Representation Orthonormalization (BRO) and Gini regularization applied during GCNN training.
result Regularization improves GCNN attribution methods and preference from chemists.
Retro* uses neural networks to efficiently find high-quality synthetic routes in organic chemistry.
problem Finding efficient synthetic routes in organic chemistry is challenging due to the vast search space.
method Retro* is a neural-based A*-like algorithm that learns a neural search bias to guide efficient best-first search.
result Retro* outperforms existing methods in both success rate and solution quality while being more efficient.
AI models predict new opioid ligands from molecular dynamics.
problem Lack of crystal structures limits virtual screening of drug candidates.
method Molecular dynamics simulation and machine learning.
result Identified a novel μ opioid chemotype. NLP techniques improve drug discovery by analyzing chemical and protein text.
problem Improving drug discovery through better analysis of chemical and protein text.
method Natural language processing techniques applied to biochemical entities.
result Enhanced prediction of molecular properties and design of novel molecules.
Paper introduces a graph-based approach for retrosynthesis prediction.
problem Predicting precursor molecules for target molecules in organic synthesis.
method Graph-based approach that predicts graph edits and synthons expansion.
result Achieves top-1 accuracy of 53.7%.
Neural networks predict substructures from mass spectra to identify chemical threats.
problem Identifying unknown chemical threats from mass spectra and formulas.
method Data-driven approach using neural networks to rank and match substructures.
result Substructure classifiers achieve over 90% micro F1-score and correctly identify structures in 88-71% of cases.
Model predicts chemical reactions from patent data.
problem Predicting chemical reactions from patent data.
method Neural sequence-to-sequence model trained on 50,000 reactions.
result Model performs comparably with expert systems and overcomes their limitations.
Neural model predicts organic reactions with high accuracy.
problem Predicting outcomes of complex organic chemistry reactions.
method Sequence-to-sequence model trained end-to-end, data-driven.
result Achieved 80.1% top-1 accuracy on a smaller dataset.
New method uses Riemannian geometry to quantify molecular shapes.
problem Quantifying molecular similarity for drug discovery.
method Riemannian geometry and Kähler quantization (KQMolSA).
result KQMolSA method compares well to existing shape similarity methods.
Sparse molecular representations improve interpretability in graph neural networks.
problem Difficulty in understanding which molecular graph aspects drive deep learning predictions.
method Constrain weights in a graph convolutional neural network using the Gini index to maximize representation inequality.
result The Gini-constrained approach does not degrade evaluation metrics and allows for interpretable representation combination.
Generative model predicts synthesizable molecules from reactants.
problem Generating molecules with desirable properties doesn't guarantee synthesis feasibility.
method Proposes a realistic synthesis process model with reactant selection and reaction prediction.
result Model generates diverse, valid, and unique synthesizable molecules.
RNNs generate drug-like molecules with good affinity to targets.
problem Generating novel drug-like molecules with desired biological affinity.
method Recurrent Neural Networks trained as generative models for molecular structures.
result Fine-tuned RNNs can generate 14% of molecules active against Staphylococcus aureus and 28% against Plasmodium falciparum.
Deep RL finds efficient pathways for sugar to chemicals.
problem Finding efficient pathways from sugar to value-added chemicals.
method Markov decision process with deep reinforcement learning.
result Promising preliminary results in efficient biomass conversion.
Model predicts electron paths in chemical reactions.
problem Predicting electron movements in chemical reactions.
method Designing a model to learn electron paths from raw reaction data.
result Model achieves excellent performance on USPTO reaction dataset.
Deep reinforcement learning optimizes retrosynthetic planning for chemical synthesis.
problem Optimizing chemical synthesis plans from molecular targets to simpler starting materials.
method Deep reinforcement learning to estimate synthesis costs and values of molecules.
result Trained neural networks outperform heuristic approaches in synthesizing unfamiliar molecules.
ChemCrow enhances LLMs for chemistry tasks, automating complex chemical processes.
problem Limited access to computational chemistry tools for large-language models.
method Integrating 18 expert-designed chemistry tools into an LLM (ChemCrow).
result ChemCrow autonomously plans and executes chemical syntheses and discoveries.
Machine learning recreates the periodic table from element properties.
problem Recreating the periodic table using machine learning.
method Unsupervised machine learning with GTM for feature embedding.
result PTG autonomously generates various periodic table layouts.
Simple machine learning models achieve high accuracy in toxicity prediction.
problem Toxicity prediction of chemical compounds using complex models.
method Using shallow neural networks and decision trees with 2D features.
result Achieves similar or better performance than deep neural networks with less computing time.
EAGCN learns attention weights and node features for multi-relational graphs.
problem Learning molecular properties from complex graph structures.
method Edge attention-based multi-relational GCN (EAGCN) that learns attention weights and node features.
result EAGCN predicts compound properties from molecular graphs efficiently and interprets attention weights.
Synthetic augmentation helps but not always in imbalanced learning.
problem Imbalanced learning causes poor performance on rare classes.
method Developed a statistical framework for synthetic augmentation in imbalanced learning.
result Synthetic augmentation is not always beneficial and depends on the imbalance regime.
Synthetic data enhances analytics but requires careful volume management.
problem Accuracy of statistical methods on synthetic data vs. raw data.
method Synthetic Data Generation for Analytics framework using tabular diffusion models.
result Error rate decreases with more synthetic data but may stabilize or increase.
Synthetic dataset for deep learning with known Gaussian distribution.
problem Lack of datasets with known distribution for deep learning verification.
method Proposes a method to generate a synthetic dataset with Gaussian distribution.
result Synthetic dataset mimics MNIST and can be used with DNNs.
Paper creates fair synthetic data ensuring equal predictions across sensitive attributes.
problem Ensuring fair predictions across sensitive attributes in synthetic data.
method Equalizing target probability distributions across sensitive attributes in synthetic data generation.
result Synthetic data provides strong fair predictions, equal across all thresholds.
Paper establishes utility theory for synthetic data generation.
problem Lack of theoretical understanding in synthetic data utility.
method Statistical learning framework with two utility metrics: generalization and model ranking.
result Theoretical bounds for synthetic data utility metrics ensure comparable generalization and consistent model comparison.
SMOTE-DP enhances synthetic data privacy without sacrificing utility.
problem Balancing privacy and utility in synthetic data generation.
method Integrating SMOTE with differential privacy mechanisms.
result SMOTE-DP produces synthetic data that is both private and useful.
A new method combines synthetic data analysis and DP generation to produce accurate uncertainty estimates.
problem Invalid inferences from DP synthetic data analysis.
method Combining synthetic data analysis techniques from MI and NA Bayesian modeling with a novel noise-aware synthetic data generation algorithm.
result Accurate confidence intervals from DP synthetic data are produced, wider with tighter privacy.
Enhances statistical inference using synthetic data.
problem Limited labeled data for statistical inference.
method GESPI framework that combines synthetic and real data.
result Error rate remains below a user-specified bound and decreases with synthetic data quality.
AI techniques explain synthetic tabular data weaknesses.
problem Challenges in evaluating synthetic tabular data quality.
method Apply explainable AI to a binary detection classifier.
result Reveals inconsistencies, unrealistic dependencies, or missing patterns in synthetic data.
The study improves theoretical understanding of using multiple synthetic datasets for better model accuracy.
problem Lack of theoretical understanding of using multiple synthetic datasets for supervised learning.
method Derive bias-variance decompositions for multiple synthetic datasets settings.
result A simple rule of thumb to select the appropriate number of synthetic datasets.
Synthetic data training improves neural networks for captcha breaking.
problem Training neural networks with synthetic data for improved performance.
method Connecting synthetic data training to model-based Bayesian inference.
result Demonstrated state-of-the-art performance and posterior uncertainty in captcha breaking.
Paper proposes a new method to assess synthetic data generators.
problem Assessing the quality of implicit generative models.
method Kernelised Stein Statistic (KSD) test based on non-parametric Stein operator.
result Improved power performance compared to existing approaches.
Paper evaluates synthetic retail data for fidelity, utility, and privacy.
problem Ensuring accurate synthetic data in retail.
method Differentiates between continuous and discrete data, measures fidelity and utility, and uses Differential Privacy for privacy.
result Validated framework for reliable and scalable synthetic data evaluation.
Enhanced synthetic dataset improves asset allocation analysis.
problem Lack of realistic synthetic data for fixed income portfolio construction.
method Improved CorrGAN model for synthetic correlation matrices and Encoder-Decoder model for additional data conditioning.
result Synthetic dataset enhances portfolio construction and asset allocation analysis.
This survey introduces synthetic timelike Ricci curvature bounds in Lorentzian spaces.
problem Synthetic timelike Ricci curvature bounds in non-smooth Lorentzian spaces.
method Optimal transport and entropy tools.
result Synthetic version of Hawking's singularity theorem and synthetic characterisation of Einstein's vacuum equations.
Differentially private synthetic control estimates treatment effects while protecting privacy.
problem Estimating treatment effects on sensitive data without revealing individual information.
method Combines non-private synthetic control and differentially private empirical risk minimization.
result Private synthetic control produces accurate predictions with minimal privacy cost.
GAN-AE generates synthetic data for imbalanced sequence classification.
problem Imbalanced sequence classification in medical devices.
method Autoencoder and GAN architecture for generating synthetic data.
result GAN-AE outperforms other synthetic data generation methods.
Develops a private synthetic graph generator using Gromov-Wasserstein distance.
problem Creating private synthetic networks for complex data.
method Random connection model, fused Gromov-Wasserstein distance, differential privacy.
result Effective algorithm for generating private synthetic graphs with theoretical guarantees.
New method generates private synthetic data without domain size dependence.
problem Private synthetic data generation for unbounded query classes.
method Constructs a private synthetic data generator for privately PAC learnable query classes.
result Sample complexity independent of domain size for privately PAC learnable query classes.
Synthetic data can amplify privacy in linear regression models.
problem Understanding how synthetic data can enhance privacy in linear regression models.
method Investigated through the linear regression framework, analyzing synthetic data generated from random inputs and controlled inputs.
result Releasing a limited number of synthetic data points amplifies privacy beyond the model's inherent guarantees when inputs are random, but not when inputs are controlled by an adversary.
Synthetic data improves financial models without real data.
problem Lack of real financial data due to privacy and regulation.
method Application of synthetic data across various financial data types.
result Synthetic data enhances financial model accuracy and fairness.
This paper uses LLMs to generate synthetic data to improve classification accuracy in imbalanced datasets.
problem Imbalanced classification and spurious correlation in data science.
method Develops novel theoretical foundations and uses transformer models to generate synthetic data.
result Transformer models can generate high-quality synthetic data to improve classification accuracy.
This paper proposes a generalization bound for GAN-synthetic data.
problem Improving classification accuracy and privacy in supervised learning.
method Proposes a generalization bound to measure the gap between synthetic and real data.
result Guarantees the generalization capability of classifiers learning from GAN-synthetic data.
Paper reviews synthetic data from AI models for statistical inference.
problem When can synthetic data be used reliably in statistical inference?
method Survey of generative models, statistical analysis of pitfalls.
result Principled use of synthetic data requires careful model specification.
Paper proposes a sparse synthetic control method to select important predictors.
problem Choosing and weighting predictors affects synthetic control estimator performance.
method Sparse synthetic control procedure that penalizes predictors, derived in a linear factor model.
result Sparse synthetic control achieves lower bias and better post-treatment performance.
SPI uses synthetic data to improve predictive inference efficiency.
problem Inefficient predictive inference with scarce calibration data.
method Integrates synthetic data to align nonconformity scores and improve coverage guarantees.
result SPI yields substantially tighter and more informative prediction sets.
Bayesian approach for learning from synthetic data, improving model accuracy.
problem Lack of statistical properties and robust methods for learning from synthetic data.
method Bayesian paradigm to update model parameters considering synthetic data generating process and learning task.
result Novel approach outperforms standard methods in supervised learning and inference problems.