Researchers formalize PD and PFI to relate them to data generating process.
problem Lack of theory linking PD and PFI to data generating process.
method Formalize PD and PFI as estimators of ground truth estimands, account for model variance with learner-PD and learner-PFI.
result PD and PFI estimates deviate from ground truth due to statistical biases, model variance, and Monte Carlo approximation errors.
This work analyzes VQ-VAEs using information theory, focusing on latent variables and their impact on generalization and data generation.
problem Lack of theoretical analysis for latent variables in unsupervised models like VQ-VAEs.
method Information-theoretic analysis, introducing a novel data-dependent prior.
result Derives a generalization error bound for VQ-VAEs that depends on LV complexity and encoder, not decoder.
Estimates copula density for complex data distributions.
problem Estimating copula density from observed data.
method Neural network-based copula density neural estimation (CODINE).
result Novel approach capable of modeling complex distributions.
Generative method avoids function estimation for data generation.
problem Challenges in function estimation for generative models.
method Deterministic point transport with gradient descent.
result Data generation possible without function estimation.
Proposes DP-MERF for privacy-preserving synthetic data generation.
problem Privacy-preserving data generation for synthetic datasets.
method Differentially private mean embeddings with random features.
result Achieves better privacy-utility trade-offs than existing methods.
New KD-tree based method for private synthetic data generation.
problem Creating private synthetic data that accurately represents real data.
method KD-trees combined with noise perturbation for differentially private synthetic data generation.
result Our data-dependent approach improves utility over prior work and scales well.
New algorithms for clustering and synthetic data generation of heterogeneous tabular datasets.
problem Clustering and generating synthetic data from heterogeneous tabular datasets with hidden cluster structure.
method Developed MMM and MMMsynth algorithms for clustering and synthetic data generation.
result MMMsynth algorithm outperforms other literature tabular-data generators and approaches real data performance.
It is generally difficult to make any statements about the expected prediction error in an univariate setting without further knowledge about how the data were generated. Recent work showed that knowledge about the real underlying causal structure of a data generation process has implications for various machine learni…
Unified framework for generating data by modeling causal and correlational dependencies.
problem Modeling both causal and correlational dependencies among latent factors.
method Causal-Correlation Variational Autoencoder (C2VAE) framework.
result Improves generation quality, disentanglement, and intervention fidelity.
ASADG improves data generation for accurate surrogate modeling of complex physical problems.
problem Training surrogate models on imbalanced data leads to inaccurate predictions.
method ASADG iteratively adds input data to improve the representation of the response manifold.
result ASADG generates more representative input data compared to LHS for better model accuracy.
BSTabDiff: Block-Subunit Diffusion Priors for HDLSS Tabular Data Generation
problem High-dimensional tabular data generation in HDLSS
method Block-subunit generative framework
result More realistic and stable synthetic data
GUIDE-VAE generates user-guided data with improved realism and performance.
problem Generating data points for multi-user datasets while considering user information.
method Conditional generative model that integrates user embeddings and a pattern dictionary-based covariance composition.
result GUIDE-VAE outperforms conventional VAEs in multi-user settings, especially under data imbalance.
Calibration of simplified vine copulas using noise contrastive estimation
problem Modeling complex multivariate dependence structures
method Noise contrastive estimation for calibration
result Improved model accuracy when simplifying assumption is violated
Survey on learning with graph-dependent data, deriving new generalization bounds.
problem Traditional i.i.d. data assumption fails in many real-life applications.
method Collect and analyze graph-dependent concentration bounds, derive generalization bounds.
result New generalization bounds for graph-dependent data.
Improves normalizing flows by incorporating data dependencies.
problem Current normalizing flow learning assumes independent data, leading to errors.
method Proposes a likelihood objective with dependencies and efficient learning algorithm.
result Improves density estimation and data generation on real-world data.
This study benchmarks tabular data generation models, optimizing hyperparameters and feature encodings.
problem Generating realistic tabular data is challenging due to heterogeneity, non-smooth distributions, and complex dependencies.
method Comprehensive evaluation of five model families on 16 datasets, considering hyperparameters, feature encodings, and architectures.
result Large-scale dataset-specific tuning significantly improves model performance, especially for diffusion-based models.
A vine copula model is a flexible high-dimensional dependence model which uses only bivariate building blocks. However, the number of possible configurations of a vine copula grows exponentially as the number of variables increases, making model selection a major challenge in development. In this work, we formulate a v…
TVineSynth generates synthetic data to balance privacy and utility.
problem Balancing privacy and utility in synthetic data generation.
method Uses vine copula with truncation to control privacy and utility trade-off.
result Achieves superior privacy-utility balance compared to competitors.
uTSGAN improves on TSGAN for generating time series data.
problem Challenges in generating time-dependent data.
method Unified training of independent networks in TSGAN.
result uTSGAN outperforms TSGAN in 80% of benchmark datasets.
Gibbs-ERM learning is a natural idealized model of learning with stochastic optimization algorithms (such as Stochastic Gradient Langevin Dynamics and ---to some extent--- Stochastic Gradient Descent), while it also arises in other contexts, including PAC-Bayesian theory, and sampling mechanisms. In this work we study …
Temporal coarse-graining of multi-sector default count data generates effective correlation matrices and rank copulas.
problem Explaining the difference in default dependence between monthly and annual aggregation.
method Dynamic low-rank state-space model with AR(1) latent credit-state factors.
result Effective correlation matrices and rank copulas are generated from monthly default count data.
We study the sample complexity of private synthetic data generation over an unbounded sized class of statistical queries, and show that any class that is privately proper PAC learnable admits a private synthetic data generator (perhaps non-efficient). Previous work on synthetic data generators focused on the case that …
What do auto-encoders learn about the underlying data generating distribution? Recent work suggests that some auto-encoder variants do a good job of capturing the local manifold structure of data. This paper clarifies some of these previous observations by showing that minimizing a particular form of regularized recons…
Several techniques for domain adaptation have been proposed to account for differences in the distribution of the data used for training and testing. The majority of this work focuses on a binary domain label. Similar problems occur in a scientific context where there may be a continuous family of plausible data genera…
Chordal graphs can be used to encode dependency models that are representable by both directed acyclic and undirected graphs. This paper discusses a very simple and efficient algorithm to learn the chordal structure of a probabilistic model from data. The algorithm is a greedy hill-climbing search algorithm that uses t…
This paper improves generative models by using data scaling and theoretical analysis.
problem Challenges in selecting noise distributions for stable learning in generative models.
method Introduces Scale-GAN, which uses data scaling and variance-based regularization.
result Data scaling controls the bias-variance trade-off and improves stability and accuracy.
Paper proposes a new method to assess synthetic data generators.
problem Assessing the quality of implicit generative models.
method Kernelised Stein Statistic (KSD) test based on non-parametric Stein operator.
result Improved power performance compared to existing approaches.
Study improves ERM for heavy-tailed data with dependent inputs.
problem Empirical Risk Minimization with dependent and heavy-tailed data.
method Extending risk bounds for ERM with heavy-tailed, dependent data.
result Established risk bounds for ERM with dependent and heavy-tailed data.
Study shows sample complexity for logistic regression with normal covariates.
problem Estimating parameters of logistic regression with normal design.
method Analyzes sample complexity in terms of dimension and inverse temperature.
result Shows two change-points in sample complexity curve based on inverse temperature.
In most papers establishing consistency for learning algorithms it is assumed that the observations used for training are realizations of an i.i.d. process. In this paper we go far beyond this classical framework by showing that support vector machines (SVMs) essentially only require that the data-generating process sa…
Generative model captures complex dependence in financial data.
problem Complex dependence structure in business and financial data.
method Multivariate generative model with heterogeneous and asymmetric tail dependence.
result Novel moment learning algorithm for scalable parameter estimation.
Multitask Gaussian process regression reduces data generation costs for molecular property prediction.
problem Data bottleneck in training surrogate models for molecular properties.
method Multitask Gaussian process regression over heterogeneous data sources (CC and DFT).
result Predicts at CC-level accuracy with over an order of magnitude reduction in data generation cost.
Diffusion Transformer captures spatial-temporal dependencies in sequential data.
problem Capturing rich spatial and temporal dependencies in sequential data.
method Established theoretical guarantees for diffusion transformers learning Gaussian process data.
result Spatial-temporal dependencies are captured within attention layers of diffusion transformers.
Early detection of Alzheimer's disease (AD) and identification of potential risk/beneficial factors are important for planning and administering timely interventions or preventive measures. In this paper, we learn a disease model for AD that combines genotypic and phenotypic profiles, and cognitive health metrics of pa…
New criteria distinguish cause from effect in data, overcoming statistical limitations.
problem Determining causal direction from statistical dependence alone.
method Intuitive criteria based on simplicity of prediction, tested on synthetic data.
result Criteria accurately distinguish cause from effect in various scenarios.
Paper adapts causal analysis for time-dependent systems, especially energy management.
problem Challenges in root-cause analysis for systems with lagged time-dependencies, particularly in energy management.
method Adapts causal root-cause analysis method to time-dependent systems, discusses two truncation approaches.
result Extension effectively localizes root-causes in feature and time domain with enough lags.
New empirical PAC-Bayes bound for Markov chains with finite state space.
problem Lack of empirical bounds for Markov chains with temporal dependence.
method Proved a new PAC-Bayes bound for Markov chains, providing an empirical pseudo-spectral gap.
result First fully empirical PAC-Bayes bound for Markov chains with finite state space.
AutoSimulate efficiently optimizes synthetic data generation.
problem Optimizing synthetic data generation for machine learning.
method Differentiable approximation of the objective function for efficient optimization.
result Significantly faster (up to 50x) and more efficient (up to 30x) synthetic data generation.
Recently, many regularized procedures have been proposed for variable selection in linear regression, but their performance depends on the tuning parameter selection. Here a criterion for the tuning parameter selection is proposed, which combines the strength of both stability selection and cross-validation and therefo…
The paper analyzes learning rates for non-irreducible Markov chains.
problem Real-world data often violates i.i.d. assumptions.
method Examines iterated random functions and contractive functions.
result Derives data-distribution dependent learning rates.
Paper addresses OPE for dependent bandit samples using MDS and batch updates.
problem Evaluating policies from non-i.i.d. historical data in contextual bandits.
method Constructs an MDS-based estimator for dependent samples, solves batch update and deficient support issues.
result Derives an asymptotically normal estimator for evaluation policy value.
The study addresses overlooked data-generating processes in time-series asset pricing.
problem The literature on time-series asset pricing overlooks the data-generating processes for factors expressed in return differences.
method The study proposes a new definition of returns and compound returns for factors, and uses OLS with net returns for single-index models.
result OLS with net returns for single-index models leads to inflated alphas, exaggerated t-values, and overestimated Sharpe ratios.
mGRN improves multivariate time series prediction by managing marginal and joint memories.
problem Extracting dependencies in multivariate sequential data with strong serial and cross-sectional dependencies.
method Developed a novel recurrent network architecture, Memory-Gated Recurrent Networks (mGRN), with gates for marginal and joint memories.
result mGRN consistently outperforms state-of-the-art architectures on various public datasets.
DecoupleNets use neural networks to assess and select dependence models.
problem Assessing and selecting dependence models for multivariate data.
method Neural networks (DecoupleNets) transform data to uniformity, then assess and select models.
result DecoupleNets provide a novel, efficient method for dependence model assessment and selection.
A neural network approach to compute stable metrics for numerical simulation data.
problem Computing stable and generalizing metrics for diverse numerical simulation data.
method A Siamese neural network architecture with a specialized loss function trained on a controlled data generation setup.
result LSiM outperforms existing metrics for vector spaces and image-based metrics.
Hermite polynomials improve private data generation by reducing feature count.
problem Infinite-dimensional features in kernel mean embedding are impractical for private data generation.
method Replace random features with Hermite polynomial features, leveraging their ordered nature.
result Hermite polynomial features yield a more accurate approximation of kernel mean embedding with fewer features.
Discrimination between non-stationarity and long-range dependency is a difficult and long-standing issue in modelling financial time series. This paper uses an adaptive spectral technique which jointly models the non-stationarity and dependency of financial time series in a non-parametric fashion assuming that the time…
Method preserves correlations in synthetic data.
problem Preserving dependence structure of original data.
method Orthogonal Procrustes problem for restoring Pearson correlation.
result Restores Pearson correlation structure while preserving feature distributions and downstream tasks performance.