A graph-based method for two-sample testing across connected nodes.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This work develops a non-parametric test for relational independence in non-i.i.d. data.
Tests for equivariance in non-parametric regression models.
Near-optimal tests and confidence sequences for non-parametric data.
New test detects when generative models memorize training data.
Unified data representation learning improves non-parametric two-sample testing.
Develops non-parametric tests for group symmetry in data.
A new framework improves kernel Stein discrepancy tests for validating distributions.
Deep neural nets optimize kernel parameters for non-parametric two-sample tests.
We propose a method for learning Markov network structures for continuous data without invoking any assumptions about the distribution of the variables. The method makes use of previous work on a non-parametric estimator for mutual information which is used to create a non-parametric test for multivariate conditional i…
New test assesses probabilistic model calibration without expensive approximations.
Constraint-based causal discovery (CCD) algorithms require fast and accurate conditional independence (CI) testing. The Kernel Conditional Independence Test (KCIT) is currently one of the most popular CI tests in the non-parametric setting, but many investigators cannot use KCIT with large datasets because the test sca…
New tests detect high-order interactions without permutations.
Improved change point detection using matched filters for non-parametric tests.
This paper offers a general and comprehensive definition of the day-of-the-week effect. Using symbolic dynamics, we develop a unique test based on ordinal patterns in order to detect it. This test uncovers the fact that the so-called "day-of-the-week" effect is partly an artifact of the hidden correlation structure of …
Adversarially robust machine learning has received much recent attention. However, prior attacks and defenses for non-parametric classifiers have been developed in an ad-hoc or classifier-specific basis. In this work, we take a holistic look at adversarial examples for non-parametric classifiers, including nearest neig…
Paper proposes a new framework for hypothesis testing in imaging.
A tractable pseudo-metric for non-parametric distributions via SPD geometry.
This study examines when non-parametric methods are robust to adversarial examples.
KQT-EWMA monitors multivariate data streams online with flexible and practical change detection.
We propose a new non parametric technique to estimate the CALL function based on the superhedging principle. Our approach does not require absence of arbitrage and easily accommodates bid/ask spreads and other market imperfections. We prove some optimal statistical properties of our estimates. As an application we firs…
In this paper, we propose the exponential Levy neural network (ELNN) for option pricing, which is a new non-parametric exponential Levy model using artificial neural networks (ANN). The ELNN fully integrates the ANNs with the exponential Levy model, a conventional pricing model. So, the ELNN can improve ANN-based model…
Develops a goodness-of-fit test for self-exciting processes.
One of the fundamental problems in supervised classification and in machine learning in general, is the modelling of non-parametric invariances that exist in data. Most prior art has focused on enforcing priors in the form of invariances to parametric nuisance transformations that are expected to be present in data. Le…
Study evaluates two-sample tests for validating generative models in high dimensions.
The paper proposes a new evaluation framework for causal inference models.
This paper investigates the utilization of maximum and average distance correlations for multivariate independence testing. We characterize their consistency properties in high-dimensional settings with respect to the number of marginally dependent dimensions, compare the advantages of each test statistic, examine thei…
A new non parametric approach to the problem of testing the independence of two random process is developed. The test statistic is the Hilbert Schmidt Independence Criterion (HSIC), which was used previously in testing independence for i.i.d pairs of variables. The asymptotic behaviour of HSIC is established when compu…
We present an arbitrage-free non-parametric yield curve prediction model which takes the full (discretized) yield curve as state variable. We believe that absence of arbitrage is an important model feature in case of highly correlated data, as it is the case for interest rates. Furthermore, the model structure allows t…
We introduce a general non-parametric independence test between right-censored survival times and covariates, which may be multivariate. Our test statistic has a dual interpretation, first in terms of the supremum of a potentially infinite collection of weight-indexed log-rank tests, with weight functions belonging to …
Early stopping of iterative algorithms is an algorithmic regularization method to avoid over-fitting in estimation and classification. In this paper, we show that early stopping can also be applied to obtain the minimax optimal testing in a general non-parametric setup. Specifically, a Wald-type test statistic is obtai…
Calculating similarities between objects defined by many heterogeneous data modalities is an important challenge in many multimedia applications. We use a multi-modal topic model as a basis for defining such a similarity between objects. We propose to compare the resulting similarities from different model realizations…
Temporal data are increasingly prevalent in modern data science. A fundamental question is whether two time series are related or not. Existing approaches often have limitations, such as relying on parametric assumptions, detecting only linear associations, and requiring multiple tests and corrections. While many non-p…
Unified CI test for categorical and ordinal data maintains power in high dimensions.
We present a dual-view mixture model to cluster users based on their features and latent behavioral functions. Every component of the mixture model represents a probability density over a feature view for observed user attributes and a behavior view for latent behavioral functions that are indirectly observed through u…
Information divergence functions play a critical role in statistics and information theory. In this paper we show that a non-parametric f-divergence measure can be used to provide improved bounds on the minimum binary classification probability of error for the case when the training and test data are drawn from the sa…
New method tests causal association using noise contrastive backdoor adjustment.
New theoretical tools simplify kernel-based tests analysis.
Bayesian non-parametric model selects latent dimensions automatically.
We propose a new multivariate dependency measure. It is obtained by considering a Gaussian kernel based distance between the copula transform of the given d-dimensional distribution and the uniform copula and then appropriately normalizing it. The resulting measure is shown to satisfy a number of desirable properties. …
We address the problem of non-parametric multiple model comparison: given candidate models, decide whether each candidate is as good as the best one(s) or worse than it. We propose two statistical tests, each controlling a different notion of decision errors. The first test, building on the post selection inference…
LCIT tests conditional independence using latent representations.
Distance correlation has gained much recent attention in the data science community: the sample statistic is straightforward to compute and asymptotically equals zero if and only if independence, making it an ideal choice to discover any type of dependency structure given sufficient sample size. One major bottleneck is…
We introduce a novel non-parametric methodology to test for the dynamical time evolution of the lag-lead structure between two arbitrary time series. The method consists in constructing a distance matrix based on the matching of all sample data pairs between the two time series. Then, the lag-lead structure is searched…
This paper improves AI defenses against network attacks using ML and adversarial learning.
Conditional independence testing is a fundamental problem underlying causal discovery and a particularly challenging task in the presence of nonlinear and high-dimensional dependencies. Here a fully non-parametric test for continuous data based on conditional mutual information combined with a local permutation scheme …
A novel feature selection method using noise-based hypothesis testing improves feature selection accuracy.
We describe a novel non-parametric statistical hypothesis test of relative dependence between a source variable and two candidate target variables. Such a test enables us to determine whether one source variable is significantly more dependent on a first target variable or a second. Dependence is measured via the Hilbe…