Bayesian methods often misinterpret data and asymptotic concepts.
problem Misunderstandings in Bayesian predictive inference.
method Discussion of two specific misunderstandings.
result Consequences of misinterpretations illustrated through examples.
Framework captures missing data in sparse data sets.
problem Capturing missing data in extremely sparse data sets.
method Coupled compound Poisson factorization with stochastic variational inference.
result Explicitly modeling missing data improves results in clustering, prediction, and matrix factorization.
G-PATE generates private data with high utility using teacher-discriminator aggregation.
problem Privacy concerns in large-scale data sharing for machine learning.
method Generative adversarial nets combined with private gradient aggregation among discriminators.
result Significantly improves privacy budget efficiency and data utility.
Graphical interpretation of unfairness in causal Bayesian networks.
problem Unfairness in datasets and models.
method Causal Bayesian networks to interpret and measure unfairness.
result Causal Bayesian networks provide a tool to measure and design fair models.
VAEs improve representation learning by inverting the data-generating process through self-consistency.
problem VAEs struggle to invert the data-generating process, yet often succeed in representation learning.
method Studied VAEs in the limit of near-deterministic decoders, proving self-consistency and showing ELBO convergence to a regularized log-likelihood.
result VAEs can perform independent mechanism analysis (IMA), recovering true latent factors under specific conditions.
Robust Bayesian inference improves model performance on discrete data.
problem Misspecification of discrete-valued models leads to poor inference and prediction.
method Total Variation Distance (TVD) for discrepancy, efficient estimator and inference method.
result Our approach significantly improves predictive performance on various data.
A new framework for private Bayesian tests maintains interpretability and computational efficiency.
problem Lack of interpretability and inability to quantify evidence in confidential data.
method Differentially private Bayesian tests based on test statistics.
result Established results on Bayes factor consistency under the proposed framework.
It is generally difficult to make any statements about the expected prediction error in an univariate setting without further knowledge about how the data were generated. Recent work showed that knowledge about the real underlying causal structure of a data generation process has implications for various machine learni…
A group theory framework unifies causal discovery methods.
problem Variations in interpreting independence across causal discovery algorithms.
method Random generic group transformations to assess cause-mechanism relationship.
result Group theoretic framework provides a general tool for studying data generating mechanisms.
Algorithm recovers independent causal mechanisms from transformed data.
problem Learning independent causal mechanisms from data.
method Unsupervised algorithm based on competing experts.
result Learned mechanisms generalize to novel domains.
SMOTE-DP enhances synthetic data privacy without sacrificing utility.
problem Balancing privacy and utility in synthetic data generation.
method Integrating SMOTE with differential privacy mechanisms.
result SMOTE-DP produces synthetic data that is both private and useful.
New approach identifies latent properties from mechanisms, not just data.
problem Identifying latent properties from data generating processes.
method Equivariance perspective on identifiable representation learning.
result Identification of latent properties is possible up to shared equivariances in known mechanisms.
New method generates synthetic survival data by conditioning on event times and censoring indicators.
problem Generating accurate synthetic survival data with censored event times.
method Conditioning covariates on event times and censoring indicators using existing tabular data generation models.
result Our method consistently outperforms baselines and improves survival model performance.
The paper tackles extrapolation in generative models by enforcing independence of mechanisms.
problem How to make generative models extrapolate to new, unseen environments?
method Developed a theoretical framework for independence of mechanisms, demonstrated on toy examples and real-world data.
result Extrapolation capabilities of generative models can be improved by enforcing independence of mechanisms explicitly during training.
When the training data in a two-class classification problem is overwhelmed by one class, most classification techniques fail to correctly identify the data points belonging to the underrepresented class. We propose Similarity-based Imbalanced Classification (SBIC) that learns patterns in the training data based on an …
Transformers adapted to spherical geometry using space-filling curves.
problem Generalizing transformers to geometric domains like spheres.
method Attention heads following a space-filling curve.
result Introduction of the Spiroformer on a 2-sphere.
New method transfers causal mechanisms for few-shot domain adaptation.
problem Few labeled target domain data for regression problems.
method Mechanism transfer using structural equations in causal modeling.
result Method can adapt from apparently different distributions.
This paper tackles time series imputation by identifying and modeling different missing mechanisms.
problem Different types of missing mechanisms (MAR, MNAR) in time series data.
method Proposes a framework for time series imputation by analyzing data generation processes and modeling latent variables via variational inference and normalizing flow.
result Establishes identifiability results for latent variables under nonlinear independent component analysis, showing that latent variables are identifiable.
Differentially private statistical inference using β-divergence.
problem Achieving differential privacy without altering data generation.
method Sampling from a generalised posterior minimizing β-divergence. result More precise inference with broader applicability.
Algorithm determines if missing data makes probabilistic queries estimable.
problem Determining estimability of probabilistic queries from missing data.
method Algorithm based on Bayesian networks and missingness mechanisms.
result Advances existing work by systematically determining estimability.
A privacy-preserving synthetic data generation framework that distinguishes between true and phantom data disclosures.
problem Detecting and explaining data disclosures in synthetic datasets.
method Customizable empirical auditing framework with statistical hypothesis testing.
result Demonstrated tighter privacy leakage bounds than prior methods.
New principle for disentangling latent factors using sparse regularization.
problem Disentangling latent factors from complex data.
method Sparse regularization of latent mechanisms to induce disentanglement.
result Recovery of latent variables up to permutation under certain conditions.
Paper adapts causal analysis for time-dependent systems, especially energy management.
problem Challenges in root-cause analysis for systems with lagged time-dependencies, particularly in energy management.
method Adapts causal root-cause analysis method to time-dependent systems, discusses two truncation approaches.
result Extension effectively localizes root-causes in feature and time domain with enough lags.
Generative Intervention Models predict perturbation effects without knowing the underlying mechanisms.
problem Predicting perturbation effects when the mechanisms are unknown.
method Generative Intervention Models (GIM) that map perturbation features to distributions over atomic interventions in a causal model.
result GIMs achieve robust out-of-distribution predictions and infer underlying perturbation mechanisms.
Improves privacy guarantees by analyzing randomness in privacy-preserving mechanisms.
problem Balancing user privacy and business constraints in privacy-preserving mechanisms.
method Analyzes explicit and implicit randomness in privacy mechanisms and proposes a probabilistic calibration method.
result Proposes privacy at risk, providing stronger privacy guarantees with quantifiable risks.
DoWhy-GCM extends causal inference in graphical models for diverse queries.
problem Addressing diverse causal queries in graphical causal models.
method Specify cause-effect relations via a causal graph, fit causal mechanisms, pose causal queries.
result Identification of root causes, attribution of causal influences, diagnosis of causal structures.
A deep learning framework discovers causal relationships from incomplete data.
problem Discovering causal knowledge from incomplete observational data.
method Imputated Causal Learning (ICL) framework for iterative missing data imputation and causal structure discovery.
result ICL outperforms state-of-the-art methods in various missing data scenarios.
Estimates how changing features affects predictions.
problem Unclear causal relationships between predictors and predictions.
method Connects causal structure of data generation and prediction mechanism, identifies feature with greatest causal influence, and estimates necessary causal intervention.
result Identifies and estimates the impact of features on prediction.
SGNs use Hamiltonian mechanics for invertible deep generative modeling.
problem Efficient and exact likelihood evaluation for deep generative models.
method Symplectic structure in latent space, Hamiltonian dynamics for data generation.
result Exact likelihood evaluation without Jacobian calculations.
This paper provides a guide to feature importance methods for better scientific inference.
problem Limited understanding of data-generating process due to opaque ML model mechanisms.
method Comprehensive review and new proofs of global feature importance methods.
result Facilitates a thorough understanding and concrete recommendations for FI methods.
A new method combines conformal prediction with Super Learner for interval predictions.
problem Constructing reliable interval predictions for complex regression functions.
method Coupling conformal prediction with Super Learner framework.
result The conformalized SL achieves valid finite-sample coverage with competitive performance.
Unified framework for representation and causal structure learning using exchangeable data.
problem Identifying latent representations or causal structures in non-i.i.d. data.
method Identifiable Exchangeable Mechanisms (IEM) framework for representation and structure learning.
result New insights and identifiability results for causal structure and representation learning.
Paper proposes SRA algorithm for online learning robustness and adaptivity.
problem Quantifying and evaluating tradeoff between robustness and adaptivity in online learning.
method SRA algorithm using biased stochastic approximation scheme with adaptive threshold.
result SRA algorithm provides superior performance in synthetic and real datasets.
BSTabDiff: Block-Subunit Diffusion Priors for HDLSS Tabular Data Generation
problem High-dimensional tabular data generation in HDLSS
method Block-subunit generative framework
result More realistic and stable synthetic data
FinStressTS creates synthetic benchmarks for financial forecasting, revealing model weaknesses.
problem Limited failure attribution in real-world financial benchmarks.
method Synthetic benchmark with 30 diagnostic environments linked to six mechanism families.
result Model performance varies by mechanism type, with autoregressive models often outperforming Transformers.
CTSyn generates high-quality synthetic tabular data.
problem Challenges in generating high-quality synthetic tabular data.
method Diffusion-based generative foundation model with autoencoder and conditional latent diffusion.
result CTSyn outperforms existing table synthesizers on standard benchmarks.
Detect hidden confounding in observational data using multiple environments.
problem Detect hidden confounding in observational data.
method Theoretical framework and simulation studies to test for hidden confounding.
result The proposed procedure correctly predicts hidden confounding, especially when bias is large.
New method learns exogenous variable distributions for better causal optimization.
problem Maximizing target variables in structural causal models.
method Learn exogenous variable distributions to improve surrogate models' fidelity.
result Improves approximation of structural causal models and broader application scenarios.
Deep learning framework improves accuracy in solid mechanics.
problem Improving accuracy in solid mechanics simulations.
method Physics Informed Neural Networks (PINN) with multi-network model.
result PINN framework leads to more accurate predictions and improved robustness.
A new STAR framework models integer-valued data with flexible distributions.
problem Modeling integer-valued data with flexibility and accuracy.
method Simultaneously Transforming and Rounding (STAR) a continuous-valued process.
result STAR framework designs a new BART model for integer-valued data with impressive predictive accuracy.
Survey on combining causal models with deep generative models for improved explainability and fairness.
problem Deep generative models lack explainability, induce spurious correlations, and poor out-of-distribution extrapolation.
method Structural causal models (SCMs) combined with deep generative models to address shortcomings.
result Causal generative models offer robustness, fairness, and interpretability.
Study on Gibbs-ERM learning, focusing on excess risk bounds and effective dimension.
problem Understanding the interplay between data distribution and learning in large hypothesis spaces.
method Distribution-dependent analysis of Gibbs-ERM, focusing on excess risk and effective dimension.
result Distribution-dependent upper bounds on excess risk, showing effective dimension controls risk.
New algorithm learns invariant representations for robust neural networks.
problem Learning robust neural network representations that are invariant to certain factors.
method Causal perspective and distribution matching approach.
result Empirically, the algorithm achieves state-of-the-art performance on domain generalization.
SCARY dataset generates complex causal scenarios for causality research.
problem Lack of complexity in existing causal datasets.
method Synthetic dataset with 40 scenarios, three seeds, and two data generation mechanisms.
result Provides a valuable resource for realistic causal discovery.
New method generates private synthetic data without domain size dependence.
problem Private synthetic data generation for unbounded query classes.
method Constructs a private synthetic data generator for privately PAC learnable query classes.
result Sample complexity independent of domain size for privately PAC learnable query classes.
GAN-based data augmentation can perpetuate biases in synthetic data.
problem Biases in synthetic data generated by GANs.
method Used a dataset of engineering researchers' head-shots to demonstrate how GANs can reinforce and amplify biases.
result GAN-based data augmentation can amplify biases in synthetic data.
Paper proposes a new method to assess synthetic data generators.
problem Assessing the quality of implicit generative models.
method Kernelised Stein Statistic (KSD) test based on non-parametric Stein operator.
result Improved power performance compared to existing approaches.
Noise increases the Rashomon ratio, leading simpler models to perform similarly to complex ones.
problem Why simpler models perform similarly to complex models on noisy datasets.
method Analyzed the data generation process and model training choices, introduced pattern diversity.
result Noisier datasets lead to larger Rashomon ratios, explaining simpler models' performance.