Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

4997146194 · May 202619922001200920182026
48 results for data-generation mechanism

G-PATE generates private data with high utility using teacher-discriminator aggregation.

problem Privacy concerns in large-scale data sharing for machine learning.
method Generative adversarial nets combined with private gradient aggregation among discriminators.
result Significantly improves privacy budget efficiency and data utility.

VAEs improve representation learning by inverting the data-generating process through self-consistency.

problem VAEs struggle to invert the data-generating process, yet often succeed in representation learning.
method Studied VAEs in the limit of near-deterministic decoders, proving self-consistency and showing ELBO convergence to a regularized log-likelihood.
result VAEs can perform independent mechanism analysis (IMA), recovering true latent factors under specific conditions.

Robust Bayesian inference improves model performance on discrete data.

problem Misspecification of discrete-valued models leads to poor inference and prediction.
method Total Variation Distance (TVD) for discrepancy, efficient estimator and inference method.
result Our approach significantly improves predictive performance on various data.

It is generally difficult to make any statements about the expected prediction error in an univariate setting without further knowledge about how the data were generated. Recent work showed that knowledge about the real underlying causal structure of a data generation process has implications for various machine learni…

2016-10-11abs ↗pdf ↗

New approach identifies latent properties from mechanisms, not just data.

problem Identifying latent properties from data generating processes.
method Equivariance perspective on identifiable representation learning.
result Identification of latent properties is possible up to shared equivariances in known mechanisms.

New method generates synthetic survival data by conditioning on event times and censoring indicators.

problem Generating accurate synthetic survival data with censored event times.
method Conditioning covariates on event times and censoring indicators using existing tabular data generation models.
result Our method consistently outperforms baselines and improves survival model performance.

The paper tackles extrapolation in generative models by enforcing independence of mechanisms.

problem How to make generative models extrapolate to new, unseen environments?
method Developed a theoretical framework for independence of mechanisms, demonstrated on toy examples and real-world data.
result Extrapolation capabilities of generative models can be improved by enforcing independence of mechanisms explicitly during training.

This paper tackles time series imputation by identifying and modeling different missing mechanisms.

problem Different types of missing mechanisms (MAR, MNAR) in time series data.
method Proposes a framework for time series imputation by analyzing data generation processes and modeling latent variables via variational inference and normalizing flow.
result Establishes identifiability results for latent variables under nonlinear independent component analysis, showing that latent variables are identifiable.

Paper adapts causal analysis for time-dependent systems, especially energy management.

problem Challenges in root-cause analysis for systems with lagged time-dependencies, particularly in energy management.
method Adapts causal root-cause analysis method to time-dependent systems, discusses two truncation approaches.
result Extension effectively localizes root-causes in feature and time domain with enough lags.

Generative Intervention Models predict perturbation effects without knowing the underlying mechanisms.

problem Predicting perturbation effects when the mechanisms are unknown.
method Generative Intervention Models (GIM) that map perturbation features to distributions over atomic interventions in a causal model.
result GIMs achieve robust out-of-distribution predictions and infer underlying perturbation mechanisms.

Improves privacy guarantees by analyzing randomness in privacy-preserving mechanisms.

problem Balancing user privacy and business constraints in privacy-preserving mechanisms.
method Analyzes explicit and implicit randomness in privacy mechanisms and proposes a probabilistic calibration method.
result Proposes privacy at risk, providing stronger privacy guarantees with quantifiable risks.

DoWhy-GCM extends causal inference in graphical models for diverse queries.

problem Addressing diverse causal queries in graphical causal models.
method Specify cause-effect relations via a causal graph, fit causal mechanisms, pose causal queries.
result Identification of root causes, attribution of causal influences, diagnosis of causal structures.

A deep learning framework discovers causal relationships from incomplete data.

problem Discovering causal knowledge from incomplete observational data.
method Imputated Causal Learning (ICL) framework for iterative missing data imputation and causal structure discovery.
result ICL outperforms state-of-the-art methods in various missing data scenarios.

Estimates how changing features affects predictions.

problem Unclear causal relationships between predictors and predictions.
method Connects causal structure of data generation and prediction mechanism, identifies feature with greatest causal influence, and estimates necessary causal intervention.
result Identifies and estimates the impact of features on prediction.

SGNs use Hamiltonian mechanics for invertible deep generative modeling.

problem Efficient and exact likelihood evaluation for deep generative models.
method Symplectic structure in latent space, Hamiltonian dynamics for data generation.
result Exact likelihood evaluation without Jacobian calculations.

This paper provides a guide to feature importance methods for better scientific inference.

problem Limited understanding of data-generating process due to opaque ML model mechanisms.
method Comprehensive review and new proofs of global feature importance methods.
result Facilitates a thorough understanding and concrete recommendations for FI methods.

Unified framework for representation and causal structure learning using exchangeable data.

problem Identifying latent representations or causal structures in non-i.i.d. data.
method Identifiable Exchangeable Mechanisms (IEM) framework for representation and structure learning.
result New insights and identifiability results for causal structure and representation learning.

FinStressTS creates synthetic benchmarks for financial forecasting, revealing model weaknesses.

problem Limited failure attribution in real-world financial benchmarks.
method Synthetic benchmark with 30 diagnostic environments linked to six mechanism families.
result Model performance varies by mechanism type, with autoregressive models often outperforming Transformers.

Detect hidden confounding in observational data using multiple environments.

problem Detect hidden confounding in observational data.
method Theoretical framework and simulation studies to test for hidden confounding.
result The proposed procedure correctly predicts hidden confounding, especially when bias is large.

New method learns exogenous variable distributions for better causal optimization.

problem Maximizing target variables in structural causal models.
method Learn exogenous variable distributions to improve surrogate models' fidelity.
result Improves approximation of structural causal models and broader application scenarios.

A new STAR framework models integer-valued data with flexible distributions.

problem Modeling integer-valued data with flexibility and accuracy.
method Simultaneously Transforming and Rounding (STAR) a continuous-valued process.
result STAR framework designs a new BART model for integer-valued data with impressive predictive accuracy.

Survey on combining causal models with deep generative models for improved explainability and fairness.

problem Deep generative models lack explainability, induce spurious correlations, and poor out-of-distribution extrapolation.
method Structural causal models (SCMs) combined with deep generative models to address shortcomings.
result Causal generative models offer robustness, fairness, and interpretability.

Study on Gibbs-ERM learning, focusing on excess risk bounds and effective dimension.

problem Understanding the interplay between data distribution and learning in large hypothesis spaces.
method Distribution-dependent analysis of Gibbs-ERM, focusing on excess risk and effective dimension.
result Distribution-dependent upper bounds on excess risk, showing effective dimension controls risk.

GAN-based data augmentation can perpetuate biases in synthetic data.

problem Biases in synthetic data generated by GANs.
method Used a dataset of engineering researchers' head-shots to demonstrate how GANs can reinforce and amplify biases.
result GAN-based data augmentation can amplify biases in synthetic data.

Noise increases the Rashomon ratio, leading simpler models to perform similarly to complex ones.

problem Why simpler models perform similarly to complex models on noisy datasets.
method Analyzed the data generation process and model training choices, introduced pattern diversity.
result Noisier datasets lead to larger Rashomon ratios, explaining simpler models' performance.