The paper develops robust tests for detecting independence in synchronous stochastic systems with finite sample guarantees.
problem Detecting independence in synchronous stochastic systems with finite sample guarantees.
method Combines confidence region estimates with permutation tests and dependence measures to detect nonlinear dependence.
result Consistent hypothesis tests for detecting independence under mild assumptions.
A new test detects non-linear independence in censored survival data.
problem Detecting non-linear independence between survival times and covariates.
method A kernel log-rank test using reproducing kernel Hilbert spaces.
result The test correctly rejects the null hypothesis under any alternative.
Discusses MultiFIT for multivariate dependence, comparing it to HSIC tests.
problem Comparing Multiscale Fisher's Independence Test (MultiFIT) to HSIC tests for multivariate dependence.
method Compares MultiFIT to HSIC tests, highlighting exact level control and performance limitations.
result Observes performance limitations of MultiFIT in terms of test power.
GaussDetect-LiNGAM eliminates Gaussianity tests for causal discovery.
problem Causal direction identification without Gaussianity assumptions.
method Leverages the equivalence between noise Gaussianity and residual independence in reverse regression.
result Gaussianity tests replaced with robust kernel-based independence tests.
Paper introduces a new test for conditional independence using weighted partial copulas.
problem Testing conditional independence between variables.
method The approach uses a weighted partial copula function and a bootstrap procedure to compute regions of rejection.
result The proposed test has competitive power compared to existing methods.
Entropy regularized OT test assesses independence between samples.
problem Testing independence between two samples.
method Entropy regularized optimal transport.
result Non-asymptotic bounds for test statistic established.
LCIT tests conditional independence using latent representations.
problem Detecting conditional independencies in statistical and machine learning tasks.
method Generative framework for learning latent representations of target variables X and Y, then testing for remaining dependencies.
result LCIT outperforms state-of-the-art baselines consistently under different metrics and settings.
Efficient tests for various statistical problems using incomplete U-statistics.
problem Nonparametric tests for two-sample, independence, and goodness-of-fit problems.
method Proposes MMDAggInc, HSICAggInc, and KSDAggInc tests aggregating over multiple kernel bandwidths.
result Aggregated tests provide a solution to the kernel selection problem and achieve optimal rates.
A new non parametric approach to the problem of testing the independence of two random process is developed. The test statistic is the Hilbert Schmidt Independence Criterion (HSIC), which was used previously in testing independence for i.i.d pairs of variables. The asymptotic behaviour of HSIC is established when compu…
This paper deals with the problem of nonparametric independence testing, a fundamental decision-theoretic problem that asks if two arbitrary (possibly multivariate) random variables X,Y are independent or not, a question that comes up in many fields like causality and neuroscience. While quantities like correlation o…
New method detects causal relationships from noisy measurements.
problem Discover causal relationships from noisy, imperfect measurements.
method Transformed Independent Noise (TIN) condition and ordered group decomposition.
result Identifies causal graph structure without over-complete ICA.
A new computationally efficient dependence measure, and an adaptive statistical test of independence, are proposed. The dependence measure is the difference between analytic embeddings of the joint distribution and the product of the marginals, evaluated at a finite set of locations (features). These features are chose…
New method tests independence with single nonstationary time series.
problem Testing independence in nonstationary nonlinear time series.
method Time-varying nonlinear regression, local long-run covariance estimation, strong Gaussian approximation.
result First framework for conditional independence testing with a single realization of a nonstationary nonlinear process.
The study finds a trade-off between model size, test loss, and training loss for linear predictors.
problem Finding the optimal balance between model size, test loss, and training loss for linear predictors.
method Established an algorithm and distribution-independent trade-off using non-asymptotic analysis.
result Models with low test loss are either classical (close to noise level training loss) or modern (large number of parameters).
New method tests DAGs without assuming linear or independent data.
problem Testing DAGs with nonlinear and time-dependent data.
method Structural, supervised and generative adversarial learning.
result Asymptotic guarantees for the test, allowing diverging data dimensions.
ACID neural network tests conditional independence efficiently.
problem Testing conditional independence in data.
method Amortized conditional independence testing using transformer-based neural networks.
result ACID achieves state-of-the-art performance and robust generalization.
Temporal data are increasingly prevalent in modern data science. A fundamental question is whether two time series are related or not. Existing approaches often have limitations, such as relying on parametric assumptions, detecting only linear associations, and requiring multiple tests and corrections. While many non-p…
Independent component analysis (ICA) decomposes multivariate data into mutually independent components (ICs). The ICA model is subject to a constraint that at most one of these components is Gaussian, which is required for model identifiability. Linear non-Gaussian component analysis (LNGCA) generalizes the ICA model t…
Linear independence testing is a fundamental information-theoretic and statistical problem that can be posed as follows: given n points {(Xi,Yi)}i=1n from a p+q dimensional multivariate distribution where Xi∈Rp and Yi∈Rq, determine whether aTX and bTY are uncorrela…
E-CIT framework reduces CITs' computational burden and improves causal discovery performance.
problem High computational cost of traditional CITs in causal discovery.
method E-CIT framework using divide-and-aggregate strategy with stable distribution p-value combination.
result Significant reduction in computational burden and competitive performance in causal discovery.
We propose a test of independence of two multivariate random vectors, given a sample from the underlying population. Our approach, which we call MINT, is based on the estimation of mutual information, whose decomposition into joint and marginal entropies facilitates the use of recently-developed efficient entropy estim…
New method for learning causal relationships in PNL models.
problem Learning causal relationships from empirical observations in PNL models.
method Rank-based methods to estimate non-linear functions, disentangling from independence tests.
result Consistent method for PNL causal discovery, validated in experiments.
Many model selection algorithms produce a path of fits specifying a sequence of increasingly complex models. Given such a sequence and the data used to produce them, we consider the problem of choosing the least complex model that is not falsified by the data. Extending the selected-model tests of Fithian et al. (2014)…
Proposes a method to calibrate data for more accurate linear correlation testing.
problem Inaccurate Pearson's correlation coefficient due to sample size and data non-normality.
method Predictive data calibration using machine learning to condition data on expected linear relationship.
result Calibrated Pearson's correlation coefficient yields a calibrated p-value and r estimate for posterior probability interpretation.
Proposes a modified Morgan-Pitman test for evaluating variances in machine learning models.
problem Limited ability to account for sampling variability in model selection.
method Enhances the classic Morgan-Pitman test for robustness in non-linear models with heavy-tailed distributions or outliers.
result Demonstrates the test's effectiveness and practical utility in model evaluation and selection.
We propose three measures of mutual dependence between multiple random vectors. All the measures are zero if and only if the random vectors are mutually independent. The first measure generalizes distance covariance from pairwise dependence to mutual dependence, while the other two measures are sums of squared distance…
We propose a method for learning Markov network structures for continuous data without invoking any assumptions about the distribution of the variables. The method makes use of previous work on a non-parametric estimator for mutual information which is used to create a non-parametric test for multivariate conditional i…
Identifying dependency in multivariate data is a common inference task that arises in numerous applications. However, existing nonparametric independence tests typically require computation that scales at least quadratically with the sample size, making it difficult to apply them to massive data. Moreover, resampling i…
A new method tests conditional independence by transforming it into an unconditional problem using transport maps.
problem Testing conditional independence between two random vectors given a third.
method Constructing transport maps to transform conditional independence into unconditional independence, estimating these maps from data using conditional continuous normalizing flow models.
result The proposed method is validated through simulations and real-data analysis, demonstrating practical effectiveness.
This work develops a non-parametric test for relational independence in non-i.i.d. data.
problem Testing independence in relational systems where data samples are not i.i.d.
method Kernel mean embedding for relational variables, consistent non-parametric scalable kernel test.
result Empirically validated effectiveness compared to state-of-the-art tests.
Conditional independence testing is an important problem, especially in Bayesian network learning and causal discovery. Due to the curse of dimensionality, testing for conditional independence of continuous variables is particularly challenging. We propose a Kernel-based Conditional Independence test (KCI-test), by con…
BERET improves binary expansion test for multivariate independence.
problem Testing independence of random vectors in arbitrary dimensions.
method Ensemble approach using sum of squared symmetry statistics and distance correlation.
result Improves power while preserving interpretability.
USP test improves on Pearson's chi-squared and G-test for independence.
problem Deficiencies in Pearson's chi-squared and G-test for independence. method USP test based on U-statistic estimator of population dependence measure. result USP test controls size, handles small cell counts, and detects minimal violations of independence.
New test for conditional independence using kernel embeddings.
problem Testing conditional independence in high-dimensional settings.
method Analytic kernel embeddings, asymptotic distribution.
result New test outperforms existing methods in high-dimensional settings.
A new test for conditional independence in discretized data.
problem Testing conditional independence when only discretized observations are available.
method Proposes a conditional independence test designed for discretized observations, using bridge equations to recover latent variables' information.
result Demonstrates the effectiveness of the proposed test through theoretical and empirical validation.
MULTIFIT tests independence between two random vectors using multiscale Fisher's test.
problem Detecting local dependence between two random vectors.
method MULTIFIT uses a resampling-free approach to test independence.
result MULTIFIT can easily handle large sample sizes and interpret dependency nature.
Study shows optimal rates for independence testing via U-statistic permutation tests.
problem Developing a valid test of independence for pairs with additional smoothness constraints.
method Defining a measure of dependence, using a permutation test based on a basis expansion and U-statistic estimator.
result Proves minimax optimality of the test in separation rates for certain cases.
DIET tests conditional independence using marginal dependence measures of residual information.
problem Computational intractability of conditional randomization tests (CRTs).
method DIET avoids fitting large models by leveraging marginal independence statistics of information residuals.
result DIET achieves higher power than other tractable CRTs on synthetic and real benchmarks.
This research designs a data-driven partition to test independence between continuous variables.
problem Testing independence between continuous random variables.
method Empirical log-likelihood statistic and data-driven tree-structured partition.
result Strongly consistent test of independence over probability families.
New algorithm reduces conditional independence tests needed for causal discovery.
problem Efficiently infer causal relations from observational data.
method Established an algorithm with complexity pO(s) tests. result Achieves exponent-optimality up to a logarithmic factor in terms of conditional independence tests.
This work identifies redundant tests in conditional-independence-based discovery that can improve graphical model accuracy.
problem Reliability and sensitivity of conditional-independence-based discovery algorithms.
method Analysis of redundant tests and their impact on error detection and correction.
result Redundant tests can improve graphical model accuracy but not all are beneficial.
New test detects independence in streaming data, adapting to data complexity.
problem Independence testing in streaming data with adaptive stopping.
method Sequential kernelized independence tests using betting principles.
result Valid inference in streaming data with improved power.
Testing (conditional) independence of multivariate random variables is a task central to statistical inference and modelling in general - though unfortunately one for which to date there does not exist a practicable workflow. State-of-art workflows suffer from the need for heuristic or subjective manual choices, high c…
Sequential Kernel-based Conditional Independence Testing via Adaptive Betting
problem Testing conditional independence
method Testing-by-betting on an adaptively optimized Kernel Conditional Independence statistic
result Significantly reduces Type I error inflation while preserving high power
Model-X test detects conditional independence in streaming data.
problem Detecting conditional independence in data streams with arbitrary dependency.
method Sequential testing inspired by model-X and testing by betting.
result Significantly reduces type-I error rate and enhances data efficiency.
Unified framework for structure learning via conditional independence testing.
problem Optimal structure learning and conditional independence testing.
method Established a fundamental connection and reduction between structure learning and conditional independence testing.
result Optimal rates for structure learning are determined by conditional independence testing rates.
The study examines property testing and estimation under non-identically distributed samples, finding necessary and sufficient sample complexities.
problem Property testing and estimation under non-identically distributed samples.
method Analysis of distributional property testing and estimation in settings with heterogeneous entities.
result Necessary and sufficient sample complexities for property testing and estimation under non-identically distributed samples.
Sequential tests for two-sample and independence testing using betting strategies.
problem Testing sequential data for two-sample and independence without kernel selection issues.
method Prediction-based betting strategies that adaptively determine distribution and joint distribution.
result Prediction-based tests outperform kernel-based approaches in high-dimensional or structured data settings.