Framework generates precise synthetic populations for scalable modeling.
problem Generating accurate synthetic populations without personal data.
method Constraint-programming framework encoding aggregated statistics and structural relations.
result Exact control of demographic profiles without requiring microdata.
New method constructs synthetic treatment groups without mean exchangeability assumption.
problem Violations of mean exchangeability assumption in randomized controlled trials.
method Weighted mixture of treatment groups from source populations, minimizing conditional maximum mean discrepancy.
result Asymptotic normality of synthetic treatment group estimator established.
New method generates geolocated synthetic populations from real data.
problem Generating synthetic populations with explicit geographic coordinates.
method Mapping coordinates into a latent space using Normalizing Flows (NF), then combining with other features in a Variational Autoencoder (VAE).
result NF+VAE architecture outperforms existing methods in generating geolocated synthetic populations.
Copula-based method generates synthetic populations from marginal distributions.
problem Generating realistic synthetic populations from limited data.
method Copula-based framework for population synthesis.
result Copula framework enhances transferability and realism of synthetic populations.
New method synthesizes large populations using deep learning.
problem Generating realistic synthetic populations for complex models.
method Deep generative modeling with Variational Autoencoder (VAE).
result VAE outperforms traditional methods in scalability and detail.
A new model synthesizes population with fewer structural and sampling zeros.
problem Synthesizing a feasible and diverse synthetic population from limited data.
method A deep generative model with two regularizations to minimize structural zeros and preserve sampling zeros.
result The model significantly improves feasibility and diversity of synthetic populations.
SynC generates synthetic population from aggregated data using Gaussian copula.
problem Generating individual-level data from aggregated datasets is challenging.
method SynC removes outliers, fits with Gaussian copula, merges datasets, and scales them.
result SynC efficiently combines multiple datasets into synthetic individual-level data.
CTGAN synthesizes population data for travel behavior simulation.
problem Synthesizing population data for agent-based transportation modeling.
method Composite Travel Generative Adversarial Network (CTGAN).
result Consistent and accurate generation of synthetic populations with tabular and sequential mobility data.
PolicySynth improves synthetic data alignment with real data for better campaign decisions.
problem Synthetic data used in decision support systems often leads to incorrect decisions.
method PolicySynth framework that conditions synthetic data on churn scorer to align with real data decisions.
result PolicySynth achieves high strategy simulation fidelity (0.923-0.960) on churn and acquisition datasets.
Syntax designs adaptive trials for subpopulations with potential benefits.
problem Identifying subpopulations with positive treatment effects in diverse patient populations.
method Adaptive patient recruitment and synthetic control estimation.
result Syntax outperforms conventional trial designs in identifying beneficial subpopulations.
A new method combines synthetic data analysis and DP generation to produce accurate uncertainty estimates.
problem Invalid inferences from DP synthetic data analysis.
method Combining synthetic data analysis techniques from MI and NA Bayesian modeling with a novel noise-aware synthetic data generation algorithm.
result Accurate confidence intervals from DP synthetic data are produced, wider with tighter privacy.
New method for valid prediction intervals in counterfactual outcomes with runtime confounding.
problem Valid prediction intervals for counterfactual outcomes under runtime confounding.
method Debiased machine learning framework grounded in semiparametric efficiency theory.
result Prediction intervals achieve desired coverage rates with faster convergence compared to standard methods.
Synthetic data augmentation can improve imbalanced classification metrics.
problem Improving imbalanced classification metrics
method Developing a framework for analyzing the effects of synthetic data augmentation on score-based classification
result Augmentation can improve AUROC, AUPRC, balanced accuracy, and F1 score
New framework for adaptive clinical trials to address real-world challenges.
problem Real-world challenges in post-regulatory clinical trials.
method RFAN framework integrating regulatory constraints and treatment policy value.
result Empirical evaluation of RFAN's performance.
New method infers dynamical systems from population data.
problem Inferring dynamical systems from population data.
method Deducing and estimating Fokker-Planck equation, projecting to test functions, sparse inference.
result Induces driving forces of dynamical systems.
Method predicts computational reproducibility of large population studies data analysis pipelines.
problem Difficulty in evaluating reproducibility of large population studies due to computational and storage requirements.
method Formulated as collaborative filtering process with constraints on training set construction.
result One sampling method, 'Random File Numbers (Uniform)', predicts reproducibility with good accuracy.
Paper tackles estimating individual treatment effects from observational data.
problem Estimating the difference between outcomes with and without treatment from single observation.
method Formulated as inference from hidden variables, uses a model of four causal populations, proposes ECM algorithm.
result ECM algorithm provides better performance compared to baseline methods on synthetic and real-world data.
Synthetic social networks closely match real-world interactions.
problem Evaluating realism of synthetic social contact networks.
method Used multiple measures of graph complexity to compare synthetic networks with stylized models and empirical data.
result Synthetic networks are more realistic than stylized models.
New method corrects biased comparisons in two-group data.
problem Unreliable inferences from biased sampling.
method Developed an inference method resilient to sampling biases.
result Controls false positives under moderate bias levels.
Active learning method for neural population dynamics using optogenetics.
problem Efficiently selecting neurons to stimulate for identifying neural population dynamics.
method Developed active learning procedure for low-rank regression to determine informative photostimulation patterns.
result Demonstrated a two-fold reduction in data required for predictive power using low-rank linear dynamical systems model.
Semi-supervised GAN creates synthetic genetic data for disease prediction.
problem Expensive and time-consuming to build large labeled genetic databases.
method Semi-supervised Genetic Generative Adversarial Network (gGAN).
result Model achieved satisfactory results with real genetic data.
A new method identifies sub-populations in unlabelled heterogeneous data by accounting for co-features.
problem Estimating sub-populations in unlabelled heterogeneous data with co-features.
method Mixture of Conditional Gaussian Graphical Models (CGGM) with penalized EM algorithm.
result The method successfully identifies sub-populations disrupted by co-features.
Analyzes how bias evolves in SGD training across different data sub-populations.
problem Understanding bias formation during machine learning training.
method Analytical description of SGD dynamics in a teacher-student setup with Gaussian-mixture model.
result Different sub-populations influence bias at different timescales, revealing shifting classifier preferences.
Kernel measures similarity of nonlinear causal structures in heterogeneous populations.
problem Learning causal structure in populations with diverse underlying structures.
method Distance covariance-based kernel for measuring similarity of causal structures.
result Kernel enables clustering of homogeneous subpopulations for causal structure learning.
New method predicts model performance under selection bias in healthcare.
problem Selection bias limits model generalizability in healthcare.
method Proposes a novel upper bound method for estimating model performance.
result Validates and demonstrates the practical utility of the method.
We consider continuous time Markovian processes where populations of individual agents interact stochastically according to kinetic rules. Despite the increasing prominence of such models in fields ranging from biology to smart cities, Bayesian inference for such systems remains challenging, as these are continuous tim…
New method estimates treatment effects across different populations.
problem Estimating treatment effects across populations with changing distributions.
method SBRL-HAP framework combining balancing and independence regularizers with hierarchical attention.
result Significant improvement in HTE estimation across out-of-distribution populations.
We investigate the problems of identity and closeness testing over a discrete population from random samples. Our goal is to develop efficient testers while guaranteeing Differential Privacy to the individuals of the population. We describe an approach that yields sample-efficient differentially private testers for the…
The paper analyzes SMOTE for imbalanced classification, providing theoretical bounds and guidelines.
problem The challenge of imbalanced classification problems, especially with minority classes.
method Theoretical analysis of SMOTE and related oversampling techniques for minority classes.
result Derives concentration and excess risk bounds for SMOTE and kernel-based classifiers.
This work generates synthetic EHRs with privacy guarantees for machine learning tasks.
problem Privacy concerns and heterogeneity in EHR data limit their use in machine learning.
method Generative Adversarial Networks (GANs) with differential privacy (DP) for synthetic data generation.
result Synthetic EHRs maintain performance close to real data, even with DP applied.
Uncertainty sampling is explained as a gradient step on a smoothed loss, leading to better parameters.
problem Reducing the amount of data required to learn a classifier.
method Interprets uncertainty sampling as a preconditioned stochastic gradient step on a smoothed zero-one loss.
result Uncertainty sampling converges to stationary points of the smoothed population zero-one loss.
New method tests weighted networks without thresholding, improving accuracy.
problem Testing and anomaly detection on weighted network data.
method Hierarchical Bayesian hypothesis testing framework for weighted networks.
result Method shows lower Type I error and higher statistical power compared to alternatives.
Develops a model to detect shared communities in non-aligned graphs.
problem Clustering and community detection in non-aligned graphs with heterogeneous populations.
method Joint Stochastic Blockmodel (Joint SBM) and efficient spectral clustering.
result The joint model better estimates communities compared to separate SBMs on individual graphs.
Framework evaluates quality of synthetic data generated with differential privacy.
problem Ensuring synthetic data retains statistical quality after applying differential privacy.
method Developed a framework to evaluate synthetic data quality from a practical researcher's viewpoint.
result Synthetic data can be evaluated against training data or underlying populations, and for specific tasks like inference or prediction.
Approach for recovering shared structure from multiple networks with unknown noise.
problem Recovering shared structure from multiple networks with unknown edge distributions.
method Exploits shared mean structure to denoise edge-level measurements and estimate population-level parameters.
result Established a finite-sample concentration inequality for low-rank eigenvalue truncation of a random weighted adjacency matrix.
NPE trains neural networks to approximate posterior distributions in SIR models from final outcome data.
problem Computational challenges in Bayesian inference for SIR models with final outcome data.
method Neural posterior estimation (NPE) using a logNormal posterior approximated by a neural network.
result NPE accurately recovers reference posteriors across various population sizes and transmission regimes.
Bayesian method identifies causal sets across populations without graph knowledge.
problem Transporting causal information across populations without causal graph knowledge.
method Combines observational and experimental data to identify s-admissible backdoor sets.
result Proves asymptotic convergence and corrects transportability bias in simulations.
Estimates class prior for unlabeled data using kernel embedding.
problem Estimating class prior in PU learning scenario where only positive and full population samples are available.
method Direct estimator based on distribution matching and kernel embedding in Reproducing Kernel Hilbert Space.
result Asymptotic consistency and explicit deviation bound for the estimator.
Proposes efficient data acquisition for personalized treatment effects from observational data.
problem Efficiently acquiring outcomes for personalized treatment effects in observational studies.
method Introduces causal, Bayesian acquisition functions to select points with overlapping support.
result Demonstrates improved sample efficiency and accuracy in learning personalized treatment effects.
A new algorithm reconstructs population dynamics from coarse samples.
problem Reconstructing population dynamics from unlabeled samples at coarse time intervals.
method Deep Momentum Multi-Marginal Schrödinger Bridge (DMSB) framework.
result Significantly outperforms baselines in synthetic and real-world datasets.
Synthetic augmentation improves financial machine learning performance in variance-dominant regimes.
problem Data scarcity in financial machine learning.
method Formalized synthetic augmentation, introduced size-matched null augmentation, and developed a non-parametric block permutation test.
result Synthetic augmentation is beneficial only in variance-dominant regimes, such as persistent volatility forecasting.
We lay theoretical foundations for new database release mechanisms that allow third-parties to construct consistent estimators of population statistics, while ensuring that the privacy of each individual contributing to the database is protected. The proposed framework rests on two main ideas. First, releasing (an esti…
A new evolutionary algorithm improves k-means clustering by recombining the entire population.
problem Optimizing the k-means clustering problem, especially in non-convex cases.
method Recombinator-k-means uses stochastic recombination with a reweighting mechanism.
result Recombinator-k-means outperforms standard genetic algorithms in optimization objective.
Paper explores how unsupervised learning can be understood through linear algebra concepts.
problem Understanding unsupervised learning through linear algebra concepts.
method Introducing the concept of linearly independent populations and using them to solve for prevalence values.
result Unsupervised learning can be realized as a generalization of supervised learning.
Improving Bayesian filtering with strictly proper scoring rules
problem Bayesian filtering of partially and noisily observed dynamical systems
method Proper scoring ensemble filter (PSEF)
result Accurate approximation of challenging filtering distributions
Proposes a method to correct for covariate shift in meta-analysis of randomized trials.
problem Invalidation of standard IPD meta-analysis due to covariate shift across studies.
method Placebo-anchored transport framework that treats source-trial outcomes as proxy signals and target-trial placebo outcomes as gold labels.
result Yields target-identified effect estimates in connected targets and a principled screen--then--transport procedure in disconnected targets.
The paper studies stability of generative models trained on mixed data.
problem Training generative models on mixed datasets (real and synthetic data).
method Developed a framework to rigorously study stability under specific conditions.
result Proved the stability of iterative training under certain conditions.
Modeling hidden neurons in SNNs using mesoscopic approximations.
problem Underconstrained problem of modeling unobserved neurons in SNNs.
method Coarse-graining and mean-field approximations to derive neuLVM.
result neuLVM can efficiently model large SNNs and recover connectivity parameters.