Discusses MultiFIT for multivariate dependence, comparing it to HSIC tests.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
MULTIFIT tests independence between two random vectors using multiscale Fisher's test.
The K-sample testing problem involves determining whether K groups of data points are each drawn from the same distribution. Analysis of variance is arguably the most classical method to test mean differences, along with several recent methods to test distributional differences. In this paper, we demonstrate the existe…
Unified framework combines dependent microbiome tests.
We describe a novel non-parametric statistical hypothesis test of relative dependence between a source variable and two candidate target variables. Such a test enables us to determine whether one source variable is significantly more dependent on a first target variable or a second. Dependence is measured via the Hilbe…
Temporal data are increasingly prevalent in modern data science. A fundamental question is whether two time series are related or not. Existing approaches often have limitations, such as relying on parametric assumptions, detecting only linear associations, and requiring multiple tests and corrections. While many non-p…
This work improves independence tests for high-dimensional data.
Proposes a new test for validating multivariate dynamic regression models.
e-LOND algorithm controls FDR in online testing with arbitrary dependencies.
We revisit the Kolmogorov-Smirnov and Cramér-von Mises goodness-of-fit (GoF) tests and propose a generalisation to identically distributed, but dependent univariate random variables. We show that the dependence leads to a reduction of the "effective" number of independent observations. The generalised GoF tests are not…
Information theory provides ideas for conceptualising information and measuring relationships between objects. It has found wide application in the sciences, but economics and finance have made surprisingly little use of it. We show that time series data can usefully be studied as information -- by noting the relations…
We consider the problem of asynchronous online testing, aimed at providing control of the false discovery rate (FDR) during a continual stream of data collection and testing, where each test may be a sequential test that can start and stop at arbitrary times. This setting increasingly characterizes real-world applicati…
Bayesian test assesses dependence between mixed data types.
Max-rank improves multiple testing in conformal prediction.
We propose three measures of mutual dependence between multiple random vectors. All the measures are zero if and only if the random vectors are mutually independent. The first measure generalizes distance covariance from pairwise dependence to mutual dependence, while the other two measures are sums of squared distance…
New findings control FDR for online testing methods under positive dependence.
We review the main "omnibus procedures" for goodness-of-fit testing for copulas: tests based on the empirical copula process, on probability integral transformations, on Kendall's dependence function, etc, and some corresponding reductions of dimension techniques. The problems of finding asymptotic distribution-free te…
Measuring conditional dependence is an important topic in statistics with broad applications including graphical models. Under a factor model setting, a new conditional dependence measure based on projection is proposed. The corresponding conditional independence test is developed with the asymptotic null distribution …
We propose a new multivariate dependency measure. It is obtained by considering a Gaussian kernel based distance between the copula transform of the given d-dimensional distribution and the uniform copula and then appropriately normalizing it. The resulting measure is shown to satisfy a number of desirable properties. …
New tests for conditional copulas based on decision trees.
The study identifies extremal dependence in financial markets using a bootstrap-based testing procedure.
DIET tests conditional independence using marginal dependence measures of residual information.
Test for quasi-independence in ordered time data.
Adaptive sequential testing optimizes epidemic control by learning optimal test strategies.
A clustering method for multivariate populations with similar dependence structures.
The statistical comparison of multiple algorithms over multiple data sets is fundamental in machine learning. This is typically carried out by the Friedman test. When the Friedman test rejects the null hypothesis, multiple comparisons are carried out to establish which are the significant differences among algorithms. …
Testing two potentially multivariate variables for statistical dependence on the basis finite samples is a fundamental statistical challenge. Here we explore a family of tests that adapt to the complexity of the relationship between the variables, promising robust power across scenarios. Building on the distance correl…
We propose a nonparametric test of independence, termed optHSIC, between a covariate and a right-censored lifetime. Because the presence of censoring creates a challenge in applying the standard permutation-based testing approaches, we use optimal transport to transform the censored dataset into an uncensored one, whil…
The paper introduces tests for high-dimensional independence using maximum and average distance correlations.
Fractal analysis is carried out on the stock market indices of seven European countries and the US. We find evidence of long range dependence in the log return series of the Mibtel (Italy) and the PX Glob (Czech Republic). Long range dependence implies that predictable patterns in the log returns do not dissipate quick…
Identifying dependency in multivariate data is a common inference task that arises in numerous applications. However, existing nonparametric independence tests typically require computation that scales at least quadratically with the sample size, making it difficult to apply them to massive data. Moreover, resampling i…
New method tests DAGs without assuming linear or independent data.
A new computationally efficient dependence measure, and an adaptive statistical test of independence, are proposed. The dependence measure is the difference between analytic embeddings of the joint distribution and the product of the marginals, evaluated at a finite set of locations (features). These features are chose…
CSD improves goodness-of-fit testing for higher-order dependence.
New method tests Granger non-causality in panel data with cross-sectional dependencies.
We study 'meta-dependence' in conditional independence tests across different empirical distributions.
Paper proposes a differentially private test for joint dependence among random vectors.
Paper proposes a chi-square test for distance correlation.
Test-asset construction affects factor model performance.
New metric predicts neural network reliability under novel conditions.
New algorithm detects tensor dependence structure alterations efficiently.
Deep-learning method improves hypothesis testing for independence.
Simple technique turns any adversarial attack into a universal one using few test examples.
An approach is proposed to determine structural shift in time-series assuming non-linear dependence of lagged values of dependent variable. Copulas are used to model non-linear dependence of time series components.
In this paper we investigate the adaptive market efficiency of the agricultural commodity futures market, using a sample of eight futures contracts. Using a battery of nonlinear tests, we uncover the nonlinear serial dependence in the returns series. We run the Hinich portmanteau bicorrelation test to uncover the momen…
The paper develops robust tests for detecting independence in synchronous stochastic systems with finite sample guarantees.
A new test detects noise in graph data, useful for forecasting.
Conditional independence testing is a fundamental problem underlying causal discovery and a particularly challenging task in the presence of nonlinear and high-dimensional dependencies. Here a fully non-parametric test for continuous data based on conditional mutual information combined with a local permutation scheme …