The study examines property testing and estimation under non-identically distributed samples, finding necessary and sufficient sample complexities.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper tests DPPs for diversity models, distinguishing them from other distributions.
We propose a new setting for testing properties of distributions while receiving samples from several distributions, but few samples per distribution. Given samples from distributions, , we design testers for the following problems: (1) Uniformity Testing: Testing whether all the 's are …
Wide class of elliptically contoured distributions is a popular model of stock returns distribution. However the important question of adequacy of the model is open. There are some results which reject and approve such model. Such results are obtained by testing some properties of elliptical model for each pair of stoc…
New method relaxes TV distance for two-sample testing without distributional assumptions.
New test for conditional independence using GNNs avoids estimating conditional distributions.
Optimal testing of discrete distributions with high probability, achieving sample complexity bounds.
We study 'meta-dependence' in conditional independence tests across different empirical distributions.
Algorithm distinguishes light-tailed from non-light-tailed distributions.
Polynomial delay algorithm tests causal models with hidden variables.
New DP training ensures models behave similarly at training and test time.
We study three fundamental statistical-learning problems: distribution estimation, property estimation, and property testing. We establish the profile maximum likelihood (PML) estimator as the first unified sample-optimal approach to a wide range of learning tasks. In particular, for every alphabet size and desired…
New tests for distributional causal effects using improved kernel estimators.
Develops a two-sample test using projected Wasserstein distance to handle high-dimensional data.
Cheap permutation tests speed up distribution testing without sacrificing accuracy.
There has been significant study on the sample complexity of testing properties of distributions over large domains. For many properties, it is known that the sample complexity can be substantially smaller than the domain size. For example, over a domain of size , distinguishing the uniform distribution from distrib…
In this work, we consider the sample complexity required for testing the monotonicity of distributions over partial orders. A distribution over a poset is monotone if, for any pair of domain elements and such that , . To understand the sample complexity of this problem, we intro…
A new framework improves kernel Stein discrepancy tests for validating distributions.
We study the problem of hypothesis testing between two discrete distributions, where we only have access to samples after the action of a known reversible Markov chain, playing the role of noise. We derive instance-dependent minimax rates for the sample complexity of this problem, and show how its dependence in time is…
Unified theoretical guarantees for distribution-free changepoint detection and testing.
Paper revisits pre-validation method, improving hypothesis testing.
Paper introduces ITD for detecting distributional changes in decentralized learning environments.
We develop a simple test for deviations from power law tails, which is based on the asymptotic properties of the empirical distribution function. We use this test to answer the question whether great natural disasters, financial crashes or electricity price spikes should be classified as dragon kings or 'only' as black…
We derive a new discrepancy statistic for measuring differences between two probability distributions based on combining Stein's identity with the reproducing kernel Hilbert space theory. We apply our result to test how well a probabilistic model fits a set of observations, and derive a new class of powerful goodness-o…
Survey on statistical inference under memory constraints.
Boosts kernel two-sample test power with multiple kernels.
This paper uses multivariate probability models to assess financial system risks.
Improved change point detection using matched filters for non-parametric tests.
Develops tests for conditional symmetry under group actions.
A new test statistic measures discrepancy between conditional distributions.
A family of maximum mean discrepancy (MMD) kernel two-sample tests is introduced. Members of the test family are called Block-tests or B-tests, since the test statistic is an average over MMDs computed on subsets of the samples. The choice of block size allows control over the tradeoff between test power and computatio…
Optimized testing of discrete distributions using predicted data.
Paper resolves open problems on sample complexity in binary hypothesis testing.
Randomization tests rely on simple data transformations and possess an appealing robustness property. In addition to being finite-sample valid if the data distribution is invariant under the transformation, these tests can be asymptotically valid under a suitable studentization of the test statistic, even if the invari…
Evaluation and validation of complicated control systems are crucial to guarantee usability and safety. Usually, failure happens in some very rarely encountered situations, but once triggered, the consequence is disastrous. Accelerated Evaluation is a methodology that efficiently tests those rarely-occurring yet critic…
Are two sets of observations drawn from the same distribution? This problem is a two-sample test. Kernel methods lead to many appealing properties. Indeed state-of-the-art approaches use the distance between kernel-based distribution representatives to derive their test statistics. Here, we show that distan…
Study on testing two populations with confounders.
Nonparametric tests via kernel embedding of distributions have witnessed a great deal of practical successes in recent years. However, statistical properties of these tests are largely unknown beyond consistency against a fixed alternative. To fill in this void, we study here the asymptotic properties of goodness-of-fi…
Enhances scenario approach for certifying design properties post-design.
New sampling and identity-testing methods for mixtures of distributions that don't satisfy approximate tensorization of entropy.
The statistical properties of the return intervals between successive 1-min volatilities of 30 liquid Chinese stocks exceeding a certain threshold are carefully studied. The Kolmogorov-Smirnov (KS) test shows that 12 stocks exhibit scaling behaviors in the distributions of for different thresholds . …
In the spirit of the emergent field of econophysics, a goodness-of-fit test for the Power-Law distribution, based on the Empirical Distribution Function (EDF) is presented, and related problems are discussed. An analysis of the tail behaviour of the daily logarithmic variation of the Mexican Stock Market Index (IPC), s…
The paper introduces localized conformal p-values for conditional testing problems.
As all physical adaptive quantum-enhanced metrology schemes operate under noisy conditions with only partially understood noise characteristics, so a practical control policy must be robust even for unknown noise. We aim to devise a test to evaluate the robustness of AQEM policies and assess the resource used by the po…
Structure discovery in graphical models is the determination of the topology of a graph that encodes conditional independence properties of the joint distribution of all variables in the model. For some class of probability distributions, an edge between two variables is present if and only if the corresponding entry i…
We propose two nonparametric statistical tests of goodness of fit for conditional distributions: given a conditional probability density function and a joint sample, decide whether the sample is drawn from for some density . Our tests, formulated with a Stein operator, can be applied to any…
The -generalised distribution fits daily stock returns well.
The paper presents new metrics to quantify and test for (i) the equality of distributions and (ii) the independence between two high-dimensional random vectors. We show that the energy distance based on the usual Euclidean distance cannot completely characterize the homogeneity of two high-dimensional distributions in …