Differential privacy allows quantifying privacy loss resulting from accessing sensitive personal data. Repeated accesses to underlying data incur increasing loss. Releasing data as privacy-preserving synthetic data would avoid this limitation, but would leave open the problem of designing what kind of synthetic data. W…
Fixed-parameter tractability of private synthetic data generation
problem Generating synthetic data under differential privacy
method Linear programming and subsampled private multiplicative weights method
result Optimal error rates across all regimes
Synthetic data helps prevent forgetting when learning sequentially.
problem Catastrophic forgetting in neural networks.
method Generate synthetic data via two-step optimisation process using meta-gradients.
result Training on synthetic data prevents forgetting when learning sequentially.
Method preserves correlations in synthetic data.
problem Preserving dependence structure of original data.
method Orthogonal Procrustes problem for restoring Pearson correlation.
result Restores Pearson correlation structure while preserving feature distributions and downstream tasks performance.
This paper tackles imbalanced data in binary classification problems.
problem Imbalanced data leads to skewed results in classification problems.
method Synthetic Minority Oversampling Technique (SMOTE) and Adaptive Synthetic (ADASYN) Sampling Approach.
result Synthetic data points enhance understanding of oversampling techniques.
The paper corrects bias in synthetic data for imbalanced learning.
problem Challenges in balancing false positive and negative rates in imbalanced data.
method Proposes a bias correction procedure to generate synthetic data for minority groups.
result Enhances prediction accuracy while avoiding overfitting.
This paper investigates how synthetic data and overparameterization improve VAE generalization.
problem Improving generalization performance of Variational Autoencoders (VAEs).
method Investigates the effectiveness of synthetic data and overparameterization in VAEs.
result Training on synthetic data and using more parameters improves VAE generalization, inference, and robustness.
Enhances statistical inference using synthetic data.
problem Limited labeled data for statistical inference.
method GESPI framework that combines synthetic and real data.
result Error rate remains below a user-specified bound and decreases with synthetic data quality.
Paper develops an efficient algorithm for experiment design using synthetic controls.
problem Designing optimal experiments for treatment effect estimation.
method Solves a phase synchronization problem via a normalized generalized power method.
result First global optimality guarantee for experiment design with pre-treatment data.
Survey on curvature bounds and isoperimetric inequalities.
problem Interplay between curvature bounds and isoperimetric problems.
method Survey of classical and recent results.
result Recent developments on curvature bounds and isoperimetric problems.
A new method combines synthetic data analysis and DP generation to produce accurate uncertainty estimates.
problem Invalid inferences from DP synthetic data analysis.
method Combining synthetic data analysis techniques from MI and NA Bayesian modeling with a novel noise-aware synthetic data generation algorithm.
result Accurate confidence intervals from DP synthetic data are produced, wider with tighter privacy.
Bayesian approach for learning from synthetic data, improving model accuracy.
problem Lack of statistical properties and robust methods for learning from synthetic data.
method Bayesian paradigm to update model parameters considering synthetic data generating process and learning task.
result Novel approach outperforms standard methods in supervised learning and inference problems.
New synthetic data analysis reveals high type 1 error rates.
problem Analyzing synthetic data for inference raises significant methodological challenges.
method Developed statistical inference tools and conducted a simulation study.
result Type 1 error rates are unacceptably high in synthetic data analysis.
New approach generates better synthetic data for neural program synthesis.
problem Current approaches to neural program synthesis generalize poorly to real data.
method Adversarial approach to control synthetic data distributions.
result Proposed method outperforms current approaches.
Paper reviews synthetic data from AI models for statistical inference.
problem When can synthetic data be used reliably in statistical inference?
method Survey of generative models, statistical analysis of pitfalls.
result Principled use of synthetic data requires careful model specification.
Method for creating synthetic multi-fidelity data sets.
problem Lack of representative synthetic datasets for multifidelity optimisation benchmarks.
method Systematic generation of synthetic fidelities from preexisting datasets.
result Allows systematic investigation of lower fidelity proxies' influence.
This paper tackles imbalanced classification with weakly supervised oversampling.
problem Imbalanced classification in high-dimensional datasets.
method Weakly supervised SMOTE, cost-sensitive NCA, bootstrap ensemble.
result Improved classification performance on synthetic and real-world datasets.
Data augmentation is rapidly gaining attention in machine learning. Synthetic data can be generated by simple transformations or through the data distribution. In the latter case, the main challenge is to estimate the label associated to new synthetic patterns. This paper studies the effect of generating synthetic data…
Study shows sample noise impacts active learning performance.
problem Impact of sample noise on active learning performance.
method Proposed Incremental Weighted K-Means for noisy samples.
result Robust sampler improves synthetic tasks but only marginally in real-life.
Synthetic data improves credit scoring models' performance without compromising borrower privacy.
problem Scarcity of real data for credit scoring models due to privacy concerns.
method Privacy-preserving training with synthetic data.
result Credit scoring models trained with synthetic data show a reduction of 3% in AUC and 6% in KS compared to real data models.
SynthBH uses synthetic data to control FDR in multiple testing.
problem Controlling false discovery rate in multiple hypothesis testing.
method SynthBH, a synthetic-powered multiple testing procedure.
result SynthBH guarantees FDR control with synthetic data.
ProGen models protein sequences for synthetic biology.
problem Generating proteins without structural annotations.
method Trained a 1.2B-parameter language model on 280M protein sequences.
result ProGen generates proteins with fine-grained control and accuracy.
New KD-tree based method for private synthetic data generation.
problem Creating private synthetic data that accurately represents real data.
method KD-trees combined with noise perturbation for differentially private synthetic data generation.
result Our data-dependent approach improves utility over prior work and scales well.
Generative synthetic data can preserve predictive accuracy but distort causal inference.
problem Distortion of average treatment effect estimates in synthetic data.
method Hybrid synthetic-data framework that generates covariates while modeling treatment and outcome mechanisms separately.
result Hybrid synthesis improves causal fidelity compared to fully generative baselines.
Labels distilled from images improve model training efficiency and flexibility.
problem Creating synthetic labels for a small set of real images to train models effectively.
method Introduce a more robust and flexible meta-learning algorithm for distillation and an effective first-order strategy based on convex optimization layers.
result Label distillation leads to improved results and greater flexibility in neural architectures.
Solves multi-class imbalanced data problem with geometry-based sampling and synthetic data.
problem Handling imbalanced multi-class data in classification problems.
method Two novel methods: undersampling and oversampling.
result Efficacy demonstrated through comparison with state-of-the-art methods.
The paper analyzes SMOTE for imbalanced classification, providing theoretical bounds and guidelines.
problem The challenge of imbalanced classification problems, especially with minority classes.
method Theoretical analysis of SMOTE and related oversampling techniques for minority classes.
result Derives concentration and excess risk bounds for SMOTE and kernel-based classifiers.
Unified framework for generating synthetic financial time series that accurately capture both marginal distributions and temporal dynamics.
problem Generating synthetic financial time series that reproduce both marginal distributions and temporal dynamics.
method SBBTS: A unified Schrödinger-Bass framework for synthetic financial time series.
result SBBTS accurately recovers stochastic volatility and correlation parameters that prior methods fail to capture.
Study detects synthetic tabular data across different tables.
problem Detecting synthetic tabular data in varied tables.
method Four table-agnostic detectors combined with preprocessing schemes.
result Cross-table learning possible with naive preprocessing, but cross-table transfer challenging.
Synthetic Petri Dish predicts neural architecture performance faster.
problem Expensive NAS evaluation process with ground-truth data.
method Instantiates motifs in small networks, evaluates with few synthetic samples.
result Significantly higher accuracy in predicting motif performance.
FedSyn generates synthetic data from multiple organizations' datasets.
problem Generating diverse synthetic data from limited datasets.
method Federated learning and GAN for privacy-preserving synthetic data generation.
result Synthetic data can be generated from diverse datasets without accessing individual data.
A new RL framework optimizes drug-like molecules synthetically.
problem Optimizing drug-like molecules for specific criteria.
method Deep Reinforcement Learning framework for chemical space optimization.
result Outperforms existing methods in pharmacological optimization.
MixBoost generates synthetic instances to balance imbalanced datasets.
problem Training models on imbalanced datasets.
method Iterative data augmentation method that selects and combines instances from majority and minority classes.
result MixBoost outperforms existing approaches on 20 benchmark datasets.
The paper shows how training with synthetic data can lead to model improvement, not degradation, under certain conditions.
problem Model collapse in iterative training on contaminated sources.
method Statistical analysis of iterative training on a mixture of true and synthetic data.
result Training with synthetic data can lead to model improvement, not degradation, under specific conditions.
Proposes a new method for generating synthetic data using copula flows.
problem Challenges of current synthetic data generation methods, especially with mixed real and categorical variables.
method Uses normalizing flows to learn copula density and univariate marginals based on copula theory.
result Demonstrates improved synthetic data generation and density estimation.
FSPO optimizes synthetic preferences for LLM personalization.
problem Personalizing large language models for diverse users.
method FSPO reframes reward modeling as a meta-learning problem, using few labeled preferences and synthetic data.
result FSPO achieves high winrates in personalized responses, both synthetic and real.
Generative model creates synthetic unlabeled data for SSL.
problem Training SSL models without real unlabeled datasets.
method Meta-optimized synthetic samples generated from generative models.
result Synthetic samples improve SSL performance more efficiently than real unlabeled data.
Modified CTGAN-Plus-Features method optimizes asset allocation with CVaR constraint.
problem Optimizing portfolio weights in asset allocation problems.
method Combines synthetic data generation with CVaR-constraint optimization.
result Synthetic data captures key characteristics of original data and outperforms conventional strategies.
Optimizes experimental design using synthetic controls for better outcomes.
problem Estimating average treatment effects in studies with pre-treatment data.
method Mixed-integer programming for selecting treated and control units and weights.
result Improves mean squared error and statistical power compared to simple alternatives.
KT models improved slightly with synthetic student data.
problem Limited access to real student data and lack of diversity in public datasets.
method Simulated student data using three statistical strategies and tested on KT baselines.
result Synthetic data can lead to similar performance as real data.
This work generates synthetic EHRs with privacy guarantees for machine learning tasks.
problem Privacy concerns and heterogeneity in EHR data limit their use in machine learning.
method Generative Adversarial Networks (GANs) with differential privacy (DP) for synthetic data generation.
result Synthetic EHRs maintain performance close to real data, even with DP applied.
A new method for generating synthetic time series improves forecasting model accuracy.
problem Training forecasting models on imbalanced time series datasets.
method Data augmentation using oversampling strategies for imbalanced learning.
result The proposed method outperforms global and local models.
We explore several oversampling techniques for an imbalanced multi-label classification problem, a setting often encountered when developing models for Computer-Aided Diagnosis (CADx) systems. While most CADx systems aim to optimize classifiers for overall accuracy without considering the relative distribution of each …
Visual Domain Adaptation is a problem of immense importance in computer vision. Previous approaches showcase the inability of even deep neural networks to learn informative representations across domain shift. This problem is more severe for tasks where acquiring hand labeled data is extremely hard and tedious. In this…
Estimates treatment effects in panel data with general intervention patterns.
problem Estimating average treatment effects in panel data with heterogeneous treatment effects.
method Extends synthetic control framework to allow rate-optimal recovery of average treatment effects for general intervention patterns.
result First rate-optimal guarantees for general intervention patterns in estimating average treatment effects.
Optimal transport explored on a specific geometric space.
problem Optimal transport problem in sub-Lorentzian Heisenberg group.
method Synthetic metric spacetime structure analysis and sub-Lorentzian version of Brenier's theorem.
result Established sub-Lorentzian version of Brenier's theorem and derived Monge-Ampère equation.
PrAda-GAN improves synthetic data generation under differential privacy.
problem Generating synthetic data under differential privacy with marginal-based methods.
method Sequential generator architecture integrating GAN and marginal-based approaches, with adaptive regularization of Bayes network structure.
result PrAda-GAN outperforms existing methods in privacy-utility trade-off on synthetic and real-world datasets.
Bayesian algorithm discovers synthetic routes from target molecules.
problem Identifying synthetic routes from desired products.
method Bayesian inference and combinatorial optimization.
result Algorithm rediscovered 80.3% and 50.0% of known synthetic routes.