Measuring conditional independence is one of the important tasks in statistical inference and is fundamental in causal discovery, feature selection, dimensionality reduction, Bayesian network learning, and others. In this work, we explore the connection between conditional independence measures induced by distances on …
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New test for conditional independence using kernel embeddings.
This work develops a non-parametric test for relational independence in non-i.i.d. data.
A statistical test of independence may be constructed using the Hilbert-Schmidt Independence Criterion (HSIC) as a test statistic. The HSIC is defined as the distance between the embedding of the joint distribution, and the embedding of the product of the marginals, in a Reproducing Kernel Hilbert Space (RKHS). It has …
Discusses MultiFIT for multivariate dependence, comparing it to HSIC tests.
Maximum mean discrepancy (MMD), also called energy distance or N-distance in statistics and Hilbert-Schmidt independence criterion (HSIC), specifically distance covariance in statistics, are among the most popular and successful approaches to quantify the difference and independence of random variables, respectively. T…
Determinantal point process have recently been used as models in machine learning and this has raised questions regarding the characterizations of conditional independence. In this paper we investigate characterizations of conditional independence. We describe some conditional independencies through the conditions on t…
This work improves fair tensor decomposition using a kernel criterion.
Sequential Kernel-based Conditional Independence Testing via Adaptive Betting
Efficiently fine-tunes patient-independent seizure detection models with tensor kernel machine.
The generalization performance of kernel methods is largely determined by the kernel, but common kernels are stationary thus input-independent and output-independent, that limits their applications on complicated tasks. In this paper, we propose a powerful and efficient spectral kernel learning framework and learned ke…
A new kernel-based CI test improves on existing methods.
Development of metrics for structural data-generating mechanisms is fundamental in machine learning and the related fields. In this paper, we give a general framework to construct metrics on random nonlinear dynamical systems, defined with the Perron-Frobenius operators in vector-valued reproducing kernel Hilbert space…
We show that the error probability of reconstructing kernel matrices from Random Fourier Features for the Gaussian kernel function is at most , where is the number of random features and is the diameter of the data domain. We also provide an information-theoretic method-independen…
Representations of probability measures in reproducing kernel Hilbert spaces provide a flexible framework for fully nonparametric hypothesis tests of independence, which can capture any type of departure from independence, including nonlinear associations and multivariate interactions. However, these approaches come wi…
New research optimizes HSIC estimation rate for translation-invariant kernels.
FastKCI speeds up KCI tests for causal inference on large datasets.
New test detects independence in streaming data, adapting to data complexity.
Kernel-based tests detect dependencies in multivariate time series, including stationary and non-stationary data.
Gaussian processes adapted for Riemannian manifolds using gauge-independent kernels.
Conditional independence testing is an important problem, especially in Bayesian network learning and causal discovery. Due to the curse of dimensionality, testing for conditional independence of continuous variables is particularly challenging. We propose a Kernel-based Conditional Independence test (KCI-test), by con…
Kernelized cumulants improve statistical analysis in high-dimensional spaces.
Distance-based tests, also called "energy statistics", are leading methods for two-sample and independence tests from the statistics community. Kernel-based tests, developed from "kernel mean embeddings", are leading methods for two-sample and independence tests from the machine learning community. A fixed-point transf…
Faster convergence of kernel mean embeddings using variance information.
MixCIT tests conditional independence for mixed data types efficiently and reliably.
A new kernel test avoids permutations for independence testing.
Independent component analysis (ICA) is a method for recovering statistically independent signals from observations of unknown linear combinations of the sources. Some of the most accurate ICA decomposition methods require searching for the inverse transformation which minimizes different approximations of the Mutual I…
We introduce kernel nonparametric tests for Lancaster three-variable interaction and for total independence, using embeddings of signed measures into a reproducing kernel Hilbert space. The resulting test statistics are straightforward to compute, and are used in powerful interaction tests, which are consistent against…
New insights into CI tests reveal key factors for practical performance.
New statistics improve kernel independence testing efficiency.
New theoretical tools simplify kernel-based tests analysis.
Paper introduces EO_k for quantifying accuracy-fairness trade-offs in FRL.
Paper introduces a new test for conditional independence using weighted partial copulas.
Entropy analysis via kernel methods for probabilistic inference.
We introduce the blind subspace deconvolution (BSSD) problem, which is the extension of both the blind source deconvolution (BSD) and the independent subspace analysis (ISA) tasks. We examine the case of the undercomplete BSSD (uBSSD). Applying temporal concatenation we reduce this problem to ISA. The associated `high …
New method tests conditional independence using spectral representations.
New bounds for KRR condition number reveal overfitting phenomena.
Kernel methods and MLPs perform similarly to linear models in high dimensions.
Efficient tests for various statistical problems using incomplete U-statistics.
We provide a unifying framework linking two classes of statistics used in two-sample and independence testing: on the one hand, the energy distances and distance covariances from the statistics literature; on the other, distances between embeddings of distributions to reproducing kernel Hilbert spaces (RKHS), as establ…
A wild bootstrap method for nonparametric hypothesis tests based on kernel distribution embeddings is proposed. This bootstrap method is used to construct provably consistent tests that apply to random processes, for which the naive permutation-based bootstrap fails. It applies to a large group of kernel tests based on…
Invariant kernels reduce rank and improve generalization across dimensions.
New unsupervised learning technique learns independent kernels for better machine learning tasks.
Unified framework for optimal kernel tests across MMD, HSIC, and KSD.
The article introduces practical estimators for kernel discrepancies.
We introduce the Mondrian kernel, a fast random feature approximation to the Laplace kernel. It is suitable for both batch and online learning, and admits a fast kernel-width-selection procedure as the random features can be re-used efficiently for all kernel widths. The features are constructed by sampling trees via a…
Sequential tests for two-sample and independence testing using betting strategies.
We investigate the problem of testing whether random variables, which may or may not be continuous, are jointly (or mutually) independent. Our method builds on ideas of the two variable Hilbert-Schmidt independence criterion (HSIC) but allows for an arbitrary number of variables. We embed the -dimensional joint …