Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

4.8%9.6%14.4%19.2% · Nov 201819922001200920182026
48 results for classifier two-sample tests

Conformal C2ST turns weak classifiers into reliable two-sample tests.

problem Determining if two distributions are identical using weak classifiers.
method Developed conformal variants of the C2ST to convert any classifier scores into reliable p-values.
result Even weak classifiers can yield powerful and reliable two-sample tests.

C2ST uses classifiers to test if two datasets are from the same distribution.

problem Assessing if two datasets are from the same distribution.
method Construct a dataset by pairing examples from each dataset with positive or negative labels, then classify and evaluate accuracy.
result C2ST learns a suitable representation, has a simple null distribution, and can interpret differences between datasets.

Improves two-sample hypothesis testing using kernel divergences and scoring rules.

problem Two-sample hypothesis testing in machine learning.
method Proposes Kernel Scoring Rules and Divergences, including the Maximum Mean Discrepancy.
result Kernel Score provides more information about embedded distributions than Maximum Mean Discrepancy.

Deep neural nets optimize kernel parameters for non-parametric two-sample tests.

problem Determining if two samples come from the same distribution.
method Deep kernels trained to maximize test power, adapting to distribution smoothness and shape.
result Deep kernels outperform simpler kernels in high dimensions and complex data.

Kernel method tests if two sets of data are from the same distribution.

problem Two-sample hypothesis testing for high dimensional data with small samples.
method One-class set classification using Set Kernel and one-class SVM.
result The method achieves zero type-I and type-II error on all cancer gene expression data sets.

The paper analyzes a neural network two-sample test using kernel analysis.

problem Determining if two datasets come from the same distribution.
method Time-analysis on a neural tangent kernel (NTK) two-sample test, extending to realistic neural network dynamics.
result Training times needed to detect deviations are well-separated in null and alternative hypothesis scenarios.

In clinical and neuroscientific studies, systematic differences between two populations of brain networks are investigated in order to characterize mental diseases or processes. Those networks are usually represented as graphs built from neuroimaging data and studied by means of graph analysis methods. The typical mach…

2015-11-19abs ↗pdf ↗

A nonparametric two-sample test using a parametric integral probability metric

problem Detecting distributional differences between two independent samples
method Propose a new two-sample test statistic based on a newly introduced integral probability metric (IPM)
result Establish theoretical guarantees for the associated two-sample testing procedure

Test assesses if a linear classifier is random or significant.

problem Determining if a linear classifier captures meaningful differences between classes.
method Proposes a homogeneity test related to linear separability, establishes upper bounds for p-values.
result Upper bounds for p-values are highly accurate for normally distributed samples.

Proposes counterfactual explanations for deep two-sample tests on high-dimensional data.

problem Limited interpretability of deep two-sample tests on high-dimensional data.
method Combines diffusion autoencoder and pretrained deep two-sample test model to generate counterfactuals.
result Counterfactual transformations increase p-values, indicating closer distribution similarity.

Optimal tests for nonparametric one- and two-sample testing are derived using MMD and KSD.

problem Developing optimal tests for nonparametric one- and two-sample testing.
method Using Sanov's theorem and Maximum Mean Discrepancy (MMD), the optimal error exponents are derived for one-sample tests. For two-sample tests, the quadratic-time Kernel Stein Discrepancy (KSD) is shown to achieve the optimal type-II error exponent.
result Achievement of optimal error exponents for nonparametric one- and two-sample testing in the universal setting.

New kernel tests detect differences between distributions exponentially quickly.

problem Characterize the asymptotic performance of kernel two-sample tests.
method Established exponentially consistent kernel two-sample tests for unknown distributions.
result Exponential decay rate of type-II error probability is optimal and independent of kernels.

Meta two-sample testing uses auxiliary data to quickly find powerful tests from limited samples.

problem Challenges in identifying powerful kernels for distinguishing complex distributions with limited data.
method Introduces meta two-sample testing (M2ST) to leverage abundant auxiliary data on related tasks.
result Proposed algorithms improve over baselines and identify powerful tests from scarce observations.

The paper designs tests for comparing ranked preference data and finds significant differences.

problem Comparing pairwise comparison and ranking data in various applications.
method Developed two-sample tests for pairwise comparison and ranking data, proving upper and lower bounds.
result Upper and lower bounds show tightness of the proposed tests, and significant differences in preferences were found.

New test detects differences in heterogeneous datasets.

problem Detecting differences between two samples with unknown heterogeneity.
method Developed a nonparametric testing procedure that handles latent heterogeneity through a composite null.
result The test accurately detects differences in the presence of unknown heterogeneity.

Sequential tests for two-sample and independence testing using betting strategies.

problem Testing sequential data for two-sample and independence without kernel selection issues.
method Prediction-based betting strategies that adaptively determine distribution and joint distribution.
result Prediction-based tests outperform kernel-based approaches in high-dimensional or structured data settings.

Study on kernel tests for high-dimensional data, focusing on MMD and CLT.

problem Asymptotic behavior of kernel two-sample tests in high dimensions and large samples.
method Maximum mean discrepancy (MMD) with isotropic kernels, deriving asymptotic expansions and CLT.
result Interplay between moment discrepancy and dimension-and-sample orders in kernel tests.

When data analysts train a classifier and check if its accuracy is significantly different from chance, they are implicitly performing a two-sample test. We investigate the statistical properties of this flexible approach in the high-dimensional setting. We prove two results that hold for all classifiers in any dimensi…

2016-02-06abs ↗pdf ↗

New deep learning methods improve estimation and GOF assessment for large-scale IFA.

problem Estimating and assessing goodness-of-fit for large-scale confirmatory IFA models.
method Extended deep learning algorithm for parameter estimation and simulation-based tests for GOF assessment.
result Proposed methods provide comparable estimates and detect latent dimensionality misspecification.

A permutation-based SW test achieves minimax-optimal power for two-sample testing.

problem Nonparametric two-sample testing using the sliced Wasserstein distance.
method Proposes a permutation-based SW test and analyzes its performance.
result Achieves minimax separation rate n1/2n^{-1/2} over multinomial and bounded-support alternatives.

Develops a two-sample test using projected Wasserstein distance to handle high-dimensional data.

problem Testing whether two high-dimensional samples come from the same distribution.
method Optimal projection to find a low-dimensional linear mapping that maximizes the Wasserstein distance between projected probability distributions.
result Characterizes the convergence rate of the projected Wasserstein distance and presents practical algorithms.

Optimal tests for goodness of fit and two-sample problems using MMD and KSD.

problem Asymptotically optimal tests for goodness of fit and two-sample problems.
method Maximum Mean Discrepancy (MMD) and Kernel Stein Discrepancy (KSD) based tests.
result Optimal tests achieve the maximum exponential decay rate under specific conditions.

New method relaxes TV distance for two-sample testing without distributional assumptions.

problem Challenges in certifying equality or providing tight bounds on TV distance for two distributions.
method Examined blurred total variation distance, a relaxation of TV distance.
result Provided theoretical guarantees for upper and lower bounds on blurred TV distance.