Cheap permutation tests speed up distribution testing without sacrificing accuracy.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new permutation method improves two-sample testing power.
A new method reduces computational costs for testing RF variable importance measures.
New tests detect high-order interactions without permutations.
TRIP detects unreliable feature importance scores in random forests.
New method for accurate permutation inference in CCA.
This paper introduces differentially private permutation tests for hypothesis testing.
A permutation-based SW test achieves minimax-optimal power for two-sample testing.
A new kernel test avoids permutations for independence testing.
Enhances GNNs by capturing node relationships, outperforming 2-WL test.
Permutation testing is a non-parametric method for obtaining the max null distribution used to compute corrected -values that provide strong control of false positives. In neuroimaging, however, the computational burden of running such an algorithm can be significant. We find that by viewing the permutation testing …
Paper proposes a differentially private test for joint dependence among random vectors.
We tackle permutation in linear regression with a new inference framework.
A new method improves feature importance and model stress-testing reliability.
A new method uses vectorized summaries of persistence diagrams for efficient hypothesis testing.
Distance correlation has gained much recent attention in the data science community: the sample statistic is straightforward to compute and asymptotically equals zero if and only if independence, making it an ideal choice to discover any type of dependency structure given sufficient sample size. One major bottleneck is…
Error bounds based on worst likely assignments use permutation tests to validate classifiers. Worst likely assignments can produce effective bounds even for data sets with 100 or fewer training examples. This paper introduces a statistic for use in the permutation tests of worst likely assignments that improves error b…
Recently, the method of b-bit minwise hashing has been applied to large-scale linear learning and sublinear time near-neighbor search. The major drawback of minwise hashing is the expensive preprocessing cost, as the method requires applying (e.g.,) k=200 to 500 permutations on the data. The testing time can also be ex…
Conditional independence testing is a fundamental problem underlying causal discovery and a particularly challenging task in the presence of nonlinear and high-dimensional dependencies. Here a fully non-parametric test for continuous data based on conditional mutual information combined with a local permutation scheme …
Multiple hypothesis testing is a significant problem in nearly all neuroimaging studies. In order to correct for this phenomena, we require a reliable estimate of the Family-Wise Error Rate (FWER). The well known Bonferroni correction method, while simple to implement, is quite conservative, and can substantially under…
A new test statistic speeds up MMD while maintaining power.
Proposes PEMI for online selective conformal prediction with asymmetric rules.
Conditional independence testing is a key problem required by many machine learning and statistics tools. In particular, it is one way of evaluating the usefulness of some features on a supervised prediction problem. We propose a novel conditional independence test in a predictive setting, and show that it achieves bet…
A new test validates ensemble models against the null hypothesis.
We present a novel algorithm, Westfall-Young light, for detecting patterns, such as itemsets and subgraphs, which are statistically significantly enriched in one of two classes. Our method corrects rigorously for multiple hypothesis testing and correlations between patterns through the Westfall-Young permutation proced…
In literature there are several studies on the performance of Bayesian network structure learning algorithms. The focus of these studies is almost always the heuristics the learning algorithms are based on, i.e. the maximisation algorithms (in score-based algorithms) or the techniques for learning the dependencies of e…
Recently many efforts have been made to incorporate persistence diagrams, one of the major tools in topological data analysis (TDA), into machine learning pipelines. To better understand the power and limitation of persistence diagrams, we carry out a range of experiments on both graph data and shape data, aiming to de…
New statistics improve kernel independence testing efficiency.
We describe a simple, efficient, permutation based procedure for selecting the penalty parameter in the LASSO. The procedure, which is intended for applications where variable selection is the primary focus, can be applied in a variety of structural settings, including generalized linear models. We briefly discuss conn…
A new nonparametric test measures dependence between variables using decision trees.
HOoD detects near-out-of-distribution groups in correlated biomedical assays.
Sample efficiency and scalability to a large number of agents are two important goals for multi-agent reinforcement learning systems. Recent works got us closer to those goals, addressing non-stationarity of the environment from a single agent's perspective by utilizing a deep net critic which depends on all observatio…
Improves A/B testing power using a two-armed bandit framework.
We propose a nonparametric test of independence, termed optHSIC, between a covariate and a right-censored lifetime. Because the presence of censoring creates a challenge in applying the standard permutation-based testing approaches, we use optimal transport to transform the censored dataset into an uncensored one, whil…
Deep neural networks have demonstrated cutting edge performance on various tasks including classification. However, it is well known that adversarially designed imperceptible perturbation of the input can mislead advanced classifiers. In this paper, Permutation Phase Defense (PPD), is proposed as a novel method to resi…
Least-squares models such as linear regression and Linear Discriminant Analysis (LDA) are amongst the most popular statistical learning techniques. However, since their computation time increases cubically with the number of features, they are inefficient in high-dimensional neuroimaging datasets. Fortunately, for k-fo…
CIT and CIF improve feature selection for downstream prediction.
We present an efficient algorithm for simultaneously training sparse generalized linear models across many related problems, which may arise from bootstrapping, cross-validation and nonparametric permutation testing. Our approach leverages the redundancies across problems to obtain significant computational improvement…
We study the problem of independence testing given independent and identically distributed pairs taking values in a -finite, separable measure space. Defining a natural measure of dependence as the squared -distance between a joint density and the product of its marginals, we first show that there is…
Unified framework uses all data to improve multiple testing efficiency.
A new MMD-based test combines kernels for two-sample testing without splitting data.
Permutation-valued features arise in a variety of applications, either in a direct way when preferences are elicited over a collection of items, or an indirect way in which numerical ratings are converted to a ranking. To date, there has been relatively limited study of regression, classification, and testing problems …
Proposes Population Difference Criterion for visually observed subpopulation differences.
New method corrects correlation bias in feature importance.
A new metric assesses causal graphs using node permutations to detect inconsistencies.
To date, testing interactions in high dimensions has been a challenging task. Existing methods often have issues with sensitivity to modeling assumptions and heavily asymptotic nominal p-values. To help alleviate these issues, we propose a permutation-based method for testing marginal interactions with a binary respons…
New framework compares credal sets for hypothesis testing with epistemic uncertainty.
The paper improves Fisher-Pitman tests for Poisson mixtures, detecting autism-related genes.