Estimates joint causal effects using single-variable interventions on nonlinear models.
problem Estimating joint causal effects from single-variable interventions.
method Identifiability result and practical estimator for decomposing causal effects.
result Joint effects can be inferred without joint interventional data for nonlinear additive models.
New method improves clustering accuracy in noisy single-cell data.
problem Challenges in clustering single-cell RNA sequencing data due to noise and variability.
method Latent plug-and-play diffusion framework with input-space steering.
result Improved clustering accuracy on synthetic and real-world single-cell data.
We compare observed corporate cumulative default probabilities to those calculated using a stochastic model based on an extension of the work of Black and Cox and find that corporations default as if via diffusive dynamics. The model, based on a contingent-claims analysis of corporate capital structure, is easily calib…
New method learns cell trajectories and network interactions from single-cell data.
problem Network inference in systems biology from steady-state data.
method Min-entropy estimation for stochastic dynamics, leveraging both temporal and perturbational data.
result Jointly learns cellular trajectories and network interactions.
CNNs predict spatial fields from sparse data.
problem Predicting complete spatial fields from limited observations.
method Convolutional Neural Networks (CNNs) trained on a single partially observed field.
result CNNs can flexibly capture local spatial patterns without explicit covariance modeling.
New method for tensor completion from specific mode observations.
problem Recovering multiway data tensors from partial observations.
method Tensor train decomposition for fiber-wise observations.
result Deterministic recovery guarantees for specific observation patterns.
Data thinning splits observations into independent parts for convolution-closed distributions.
problem Validation of unsupervised learning results in settings with limited data.
method Data thinning, splitting observations into independent parts following the same distribution.
result Data thinning provides an attractive alternative to cross-validation in settings with limited sample splitting.
Consider the case that one observes a single time-series, where at each time t one observes a data record O(t) involving treatment nodes A(t), possible covariates L(t) and an outcome node Y(t). The data record at time t carries information for an (potentially causal) effect of the treatment A(t) on the outcome Y(t), in…
New model clusters cells and individuals, revealing genetic influences on cell types.
problem Clustering nested data with group-level and observation-level variables.
method Nested Atoms Model (NAM), Bayesian nonparametric approach.
result Identifies clusters of genetically similar individuals with homogeneous cell-type profiles.
KRCD detects unobserved confounders in nonlinear observational data.
problem Detecting unobserved confounders in nonlinear observational studies.
method Kernel Regression Confounder Detection (KRCD) using reproducing kernel Hilbert spaces.
result KRCD outperforms existing methods and achieves superior computational efficiency.
We investigate the variety of a portfolio of stocks in normal and extreme days of market activity. We show that the variety carries information about the market activity which is not present in the single-index model and we observe that the variety time evolution is not time reversal around the crash days. We obtain th…
Give deterministic necessary and sufficient conditions to guarantee that if a subspace fits certain partially observed data from a union of subspaces, it is because such data really lies in a subspace. Furthermore, Give deterministic necessary and sufficient conditions to guarantee that if a subspace fits certain parti…
Motivated by the task of 2-D classification in single particle reconstruction by cryo-electron microscopy (cryo-EM), we consider the problem of heterogeneous multireference alignment of images. In this problem, the goal is to estimate a (typically small) set of target images from a (typically large) collection of obser…
New method approximates M-estimator and predictions without solving fixed-point equations.
problem Characterize behavior of M-estimator and predictions in single index models.
method Develops data-driven observable adjustments to proximal operators.
result Empirical distributions of M-estimator and predictions are approximated without solving fixed-point equations.
The inference of the causal relationship between a pair of observed variables is a fundamental problem in science, and most existing approaches are based on one single causal model. In practice, however, observations are often collected from multiple sources with heterogeneous causal models due to certain uncontrollabl…
The input data features set for many data driven tasks is high-dimensional while the intrinsic dimension of the data is low. Data analysis methods aim to uncover the underlying low dimensional structure imposed by the low dimensional hidden parameters by utilizing distance metrics that consider the set of attributes as…
New model generates realistic single-cell gene expression data.
problem Generating realistic single-cell gene expression profiles is challenging.
method scLDM, a latent diffusion model using Diffusion Transformers and linear interpolants.
result Superior performance in generating realistic single-cell gene expression data.
In data science and machine learning, hierarchical parametric models, such as mixture models, are often used. They contain two kinds of variables: observable variables, which represent the parts of the data that can be directly measured, and latent variables, which represent the underlying processes that generate the d…
Estimates effects of multiple interventions with hidden confounders using single-variable interventions.
problem Estimating effects of multiple interventions in the presence of hidden confounders.
method Identifiability under nonlinear structural causal model with additive Gaussian noise; pooling and joint likelihood maximization.
result Proven identifiability and superior performance compared to baseline.
We show how any dataset of any modality (time-series, images, sound...) can be approximated by a well-behaved (continuous, differentiable...) scalar function with a single real-valued parameter. Building upon elementary concepts from chaos theory, we adopt a pedagogical approach demonstrating how to adjust this paramet…
Detect hidden confounding in observational data using multiple environments.
problem Detect hidden confounding in observational data.
method Theoretical framework and simulation studies to test for hidden confounding.
result The proposed procedure correctly predicts hidden confounding, especially when bias is large.
Proposes a method to estimate SDE noise from a single trajectory.
problem Estimating SDE noise from a single data trajectory without ergodicity or stationarity.
method Combining Taylor expansions, Girsanov transformations, and drift function's initial value for drift and noise estimation.
result First SSISDE algorithm capable of identifying SDE dynamics from a single trajectory.
Predict missing and future data points in light curves using scalable Gaussian Processes.
problem Gappy time-series data from commercial cameras confound light curve prediction.
method MuyGPs, a scalable framework for hyperparameter estimation of Gaussian Processes using nearest neighbors sparsification and local cross-validation.
result MuyGPs enable accurate prediction of missing and future data points in light curves.
A Python package solves source duplication in single channel LVMs using spectral regularisation.
problem Source duplication in LVMs hampers their practical use in single channel applications.
method Spectral regularisation term added to address source duplication issue.
result Spectral regularisation framework enables easier investigation and utilisation of LVMs.
Framework identifies causal direction from single data setting.
problem Identify causal direction from single observational data.
method VCEI framework based on ICM principle and artificial variation.
result VCEI is competitive to other frameworks in identifying causal direction.
Simplified identification methods for causal inference with arbitrary interventional distributions.
problem Estimating cause-effect relationships from data with experimental interventions.
method Using Single World Intervention Graphs and nested model factorization, we provide algorithms for identifying causal parameters from mixed observational and interventional distributions.
result Our algorithms are complete for certain types of interventional marginal distributions.
Study learns linear system dynamics from noisy bilinear data.
problem Learning linear dynamics from bilinear observations with process and measurement noise.
method Regression with Kronecker product design, data-dependent and independent error bounds.
result Upper bounds on statistical error rates and sample complexity for learning dynamics matrices.
New method identifies latent causal factors from observational data alone.
problem Identifying latent causal factors without interventions or graphical restrictions.
method Characterization of latent factors in nonlinear causal models with additive Gaussian noise and linear mixing, using a practical algorithm based on solving a quadratic program over observed data.
result Latent causal variables can be identified up to a layer-wise transformation, and further disentanglement is not possible.
IMPACC improves consensus clustering for bioinformatics data.
problem Consensus clustering's inefficiency and lack of interpretability for large-scale data.
method Ensemble minipatch co-occurrences, adaptive sampling of observations and features.
result Significantly improved accuracy and interpretability with substantial computational savings.
Paper tackles catastrophic overfitting in single-step adversarial training.
problem Catastrophic overfitting leads to sudden drop in robust accuracy.
method Proposes a method to prevent overfitting by using all adversarial examples.
result Demonstrates prevention of catastrophic overfitting and improves robustness.
FoundCause: Causal Discovery with Latent Confounders from Observational Data
problem Causal discovery from observational data
method FoundCause, an amortized causal discovery model trained on synthetic data
result FoundCause outperforms classical and amortized methods on real-world datasets
Method learns causal effects from multiple interventions in presence of unobserved confounders.
problem Disentangling causal effects from sets of interventions in the presence of unobserved confounders.
method Non-linear structural causal models with additive, multivariate Gaussian noise; algorithm that learns causal model parameters by pooling data from different regimes and maximizing combined likelihood.
result Identification proofs demonstrate that causal effects of single interventions can be learned from sets of interventions, even with unobserved confounders.
Unified framework for simulation-based inference learns a single model for multiple tasks.
problem Simulation-based inference for multiple tasks with limited model retraining.
method Unified flow-matching generative model with query-aware masking distribution.
result Competitive performance on various inference tasks and real-world problems.
Network models have been popular for modeling and representing complex relationships and dependencies between observed variables. When data comes from a dynamic stochastic process, a single static network model cannot adequately capture transient dependencies, such as, gene regulatory dependencies throughout a developm…
Method estimates causal effects from combined interventional and observational data.
problem Estimating causal effects from unobserved confounders.
method Causal reduction method replacing latent confounders with a single latent confounder.
result Improves estimation accuracy from combined data without observing all confounders.
Robust Trimmed k-means improves clustering with outliers and mixed data.
problem Real-world data often contains outliers and mixed membership clusters, complicating traditional clustering methods.
method Proposes Robust Trimmed k-means (RTKM) that robustifies k-means for both single- and multi-membership data.
result RTKM outperforms other methods on multi-membership data with outliers and single membership data with outliers.
Approach generates multiple correct predictions from single supervision.
problem Single correct prediction from multiple possible alternatives.
method Develops an approach to generate multiple high-quality predictions.
result Can generate high-quality outputs different from observed.
In this paper we present detailed simulation results on the wealth distribution model with quenched saving propensities. Unlike other wealth distribution models where the saving propensities are either zero or constant, this model is not found to be ergodic and self-averaging. The wealth distribution statistics with a …
We propose a probabilistic model for interpreting gene expression levels that are observed through single-cell RNA sequencing. In the model, each cell has a low-dimensional latent representation. Additional latent variables account for technical effects that may erroneously set some observations of gene expression leve…
Neuroscience is experiencing a data revolution in which many hundreds or thousands of neurons are recorded simultaneously. Currently, there is little consensus on how such data should be analyzed. Here we introduce LFADS (Latent Factor Analysis via Dynamical Systems), a method to infer latent dynamics from simultaneous…
In many estimation problems, e.g. linear and logistic regression, we wish to minimize an unknown objective given only unbiased samples of the objective function. Furthermore, we aim to achieve this using as few samples as possible. In the absence of computational constraints, the minimizer of a sample average of observ…
Develops a framework for causal structure learning using both interventional and observational data.
problem Lack of identifiability of causal structures with only observational data.
method Bilevel polynomial optimization (Bloom) framework for causal structure discovery from interventional and observational data.
result Bloom framework provides convergence and optimality guarantees, surpassing other learning algorithms in experiments.
Learning the parameters of a (potentially partially observable) random field model is intractable in general. Instead of focussing on a single optimal parameter value we propose to treat parameters as dynamical quantities. We introduce an algorithm to generate complex dynamics for parameters and (both visible and hidde…
Recommender systems often face heterogeneous datasets containing highly personalized historical data of users, where no single model could give the best recommendation for every user. We observe this ubiquitous phenomenon on both public and private datasets and address the model selection problem in pursuit of optimizi…
We apply a fast kernel method for mask-based single-channel speech enhancement. Specifically, our method solves a kernel regression problem associated to a non-smooth kernel function (exponential power kernel) with a highly efficient iterative method (EigenPro). Due to the simplicity of this method, its hyper-parameter…
The paper formalizes criteria for non-spurious and disentangled representations using causal methods.
problem Formalizing criteria for non-spurious and disentangled representations in representation learning.
method Causal perspective, counterfactual quantities, observable consequences of causal assertions.
result Computable metrics for assessing representation learning based on observed data.
Framework for causal discovery using multi-modal data.
problem Failure of representation learning in causal tasks.
method Statistical and computational framework combining representation learning and causal inference.
result Effective use of observational and perturbational data for causal discovery.
CausalPFN automates causal effect estimation from observational data.
problem Manual selection of causal effect estimators is time-consuming and requires domain expertise.
method CausalPFN is a transformer that learns to infer causal effects from raw observations without task-specific adjustments.
result CausalPFN achieves superior performance on various benchmarks and real-world tasks.