Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

4488131175 · Jun 202019922001200920182026
48 results for synthetic chemists

New method explains complex fuel compound classifications.

problem Understanding complex quantitative structure-activity relationship models.
method Locally Interpretable Machine-Agnostic Explanations (LIME) applied to 2-D chemical structures.
result Replicates chemical intuition, allowing direct acceptance/rejection of decisions.

Two regularization techniques improve GCNN explainability and preference from chemists.

problem Difficulty in rationalizing molecular graph neural network predictions.
method Batch Representation Orthonormalization (BRO) and Gini regularization applied during GCNN training.
result Regularization improves GCNN attribution methods and preference from chemists.

Retro* uses neural networks to efficiently find high-quality synthetic routes in organic chemistry.

problem Finding efficient synthetic routes in organic chemistry is challenging due to the vast search space.
method Retro* is a neural-based A*-like algorithm that learns a neural search bias to guide efficient best-first search.
result Retro* outperforms existing methods in both success rate and solution quality while being more efficient.

NLP techniques improve drug discovery by analyzing chemical and protein text.

problem Improving drug discovery through better analysis of chemical and protein text.
method Natural language processing techniques applied to biochemical entities.
result Enhanced prediction of molecular properties and design of novel molecules.

Neural networks predict substructures from mass spectra to identify chemical threats.

problem Identifying unknown chemical threats from mass spectra and formulas.
method Data-driven approach using neural networks to rank and match substructures.
result Substructure classifiers achieve over 90% micro F1-score and correctly identify structures in 88-71% of cases.

Sparse molecular representations improve interpretability in graph neural networks.

problem Difficulty in understanding which molecular graph aspects drive deep learning predictions.
method Constrain weights in a graph convolutional neural network using the Gini index to maximize representation inequality.
result The Gini-constrained approach does not degrade evaluation metrics and allows for interpretable representation combination.

RNNs generate drug-like molecules with good affinity to targets.

problem Generating novel drug-like molecules with desired biological affinity.
method Recurrent Neural Networks trained as generative models for molecular structures.
result Fine-tuned RNNs can generate 14% of molecules active against Staphylococcus aureus and 28% against Plasmodium falciparum.

Deep reinforcement learning optimizes retrosynthetic planning for chemical synthesis.

problem Optimizing chemical synthesis plans from molecular targets to simpler starting materials.
method Deep reinforcement learning to estimate synthesis costs and values of molecules.
result Trained neural networks outperform heuristic approaches in synthesizing unfamiliar molecules.

ChemCrow enhances LLMs for chemistry tasks, automating complex chemical processes.

problem Limited access to computational chemistry tools for large-language models.
method Integrating 18 expert-designed chemistry tools into an LLM (ChemCrow).
result ChemCrow autonomously plans and executes chemical syntheses and discoveries.

Simple machine learning models achieve high accuracy in toxicity prediction.

problem Toxicity prediction of chemical compounds using complex models.
method Using shallow neural networks and decision trees with 2D features.
result Achieves similar or better performance than deep neural networks with less computing time.

EAGCN learns attention weights and node features for multi-relational graphs.

problem Learning molecular properties from complex graph structures.
method Edge attention-based multi-relational GCN (EAGCN) that learns attention weights and node features.
result EAGCN predicts compound properties from molecular graphs efficiently and interprets attention weights.

Synthetic augmentation helps but not always in imbalanced learning.

problem Imbalanced learning causes poor performance on rare classes.
method Developed a statistical framework for synthetic augmentation in imbalanced learning.
result Synthetic augmentation is not always beneficial and depends on the imbalance regime.

Paper creates fair synthetic data ensuring equal predictions across sensitive attributes.

problem Ensuring fair predictions across sensitive attributes in synthetic data.
method Equalizing target probability distributions across sensitive attributes in synthetic data generation.
result Synthetic data provides strong fair predictions, equal across all thresholds.

Paper establishes utility theory for synthetic data generation.

problem Lack of theoretical understanding in synthetic data utility.
method Statistical learning framework with two utility metrics: generalization and model ranking.
result Theoretical bounds for synthetic data utility metrics ensure comparable generalization and consistent model comparison.

A new method combines synthetic data analysis and DP generation to produce accurate uncertainty estimates.

problem Invalid inferences from DP synthetic data analysis.
method Combining synthetic data analysis techniques from MI and NA Bayesian modeling with a novel noise-aware synthetic data generation algorithm.
result Accurate confidence intervals from DP synthetic data are produced, wider with tighter privacy.

AI techniques explain synthetic tabular data weaknesses.

problem Challenges in evaluating synthetic tabular data quality.
method Apply explainable AI to a binary detection classifier.
result Reveals inconsistencies, unrealistic dependencies, or missing patterns in synthetic data.

The study improves theoretical understanding of using multiple synthetic datasets for better model accuracy.

problem Lack of theoretical understanding of using multiple synthetic datasets for supervised learning.
method Derive bias-variance decompositions for multiple synthetic datasets settings.
result A simple rule of thumb to select the appropriate number of synthetic datasets.

Synthetic data training improves neural networks for captcha breaking.

problem Training neural networks with synthetic data for improved performance.
method Connecting synthetic data training to model-based Bayesian inference.
result Demonstrated state-of-the-art performance and posterior uncertainty in captcha breaking.

Enhanced synthetic dataset improves asset allocation analysis.

problem Lack of realistic synthetic data for fixed income portfolio construction.
method Improved CorrGAN model for synthetic correlation matrices and Encoder-Decoder model for additional data conditioning.
result Synthetic dataset enhances portfolio construction and asset allocation analysis.

This survey introduces synthetic timelike Ricci curvature bounds in Lorentzian spaces.

problem Synthetic timelike Ricci curvature bounds in non-smooth Lorentzian spaces.
method Optimal transport and entropy tools.
result Synthetic version of Hawking's singularity theorem and synthetic characterisation of Einstein's vacuum equations.

Differentially private synthetic control estimates treatment effects while protecting privacy.

problem Estimating treatment effects on sensitive data without revealing individual information.
method Combines non-private synthetic control and differentially private empirical risk minimization.
result Private synthetic control produces accurate predictions with minimal privacy cost.

Develops a private synthetic graph generator using Gromov-Wasserstein distance.

problem Creating private synthetic networks for complex data.
method Random connection model, fused Gromov-Wasserstein distance, differential privacy.
result Effective algorithm for generating private synthetic graphs with theoretical guarantees.

Synthetic data can amplify privacy in linear regression models.

problem Understanding how synthetic data can enhance privacy in linear regression models.
method Investigated through the linear regression framework, analyzing synthetic data generated from random inputs and controlled inputs.
result Releasing a limited number of synthetic data points amplifies privacy beyond the model's inherent guarantees when inputs are random, but not when inputs are controlled by an adversary.

This paper uses LLMs to generate synthetic data to improve classification accuracy in imbalanced datasets.

problem Imbalanced classification and spurious correlation in data science.
method Develops novel theoretical foundations and uses transformer models to generate synthetic data.
result Transformer models can generate high-quality synthetic data to improve classification accuracy.

This paper proposes a generalization bound for GAN-synthetic data.

problem Improving classification accuracy and privacy in supervised learning.
method Proposes a generalization bound to measure the gap between synthetic and real data.
result Guarantees the generalization capability of classifiers learning from GAN-synthetic data.

Paper proposes a sparse synthetic control method to select important predictors.

problem Choosing and weighting predictors affects synthetic control estimator performance.
method Sparse synthetic control procedure that penalizes predictors, derived in a linear factor model.
result Sparse synthetic control achieves lower bias and better post-treatment performance.

Bayesian approach for learning from synthetic data, improving model accuracy.

problem Lack of statistical properties and robust methods for learning from synthetic data.
method Bayesian paradigm to update model parameters considering synthetic data generating process and learning task.
result Novel approach outperforms standard methods in supervised learning and inference problems.