AutoML simplifies two-sample tests for detecting distribution shifts.
problem Detecting distribution shifts between datasets.
method Uses mean discrepancy of a witness function with squared loss minimization.
result AutoML simplifies and improves two-sample testing performance.
Study evaluates ML methods for two-sample testing with right-censored data.
problem Evaluating ML methods for two-sample testing with right-censored data.
method Developed and compared several ML-based methods with classical tests.
result Proposed methods outperform classical tests in terms of statistical power.
A test for comparing large random graphs based on network statistics.
problem Comparing friendship networks on Facebook and LinkedIn.
method General principle for two-sample hypothesis testing based on concentration of network statistics.
result A consistent two-sample test that is minimax optimal for certain network statistics.
The paper analyzes a neural network two-sample test using kernel analysis.
problem Determining if two datasets come from the same distribution.
method Time-analysis on a neural tangent kernel (NTK) two-sample test, extending to realistic neural network dynamics.
result Training times needed to detect deviations are well-separated in null and alternative hypothesis scenarios.
A graph-based method for two-sample testing across connected nodes.
problem Identifying nodes where two probability distributions differ significantly.
method Collaborative non-parametric two-sample testing (CTST) framework.
result CTST outperforms independent node tests by leveraging graph structure.
Study on testing two populations with confounders.
problem Determining if two populations have the same distribution after accounting for confounding factors.
method Introduce two general frameworks for conditional two-sample testing.
result Demonstrated the power and validity of the proposed frameworks.
A nonparametric two-sample test using a parametric integral probability metric
problem Detecting distributional differences between two independent samples
method Propose a new two-sample test statistic based on a newly introduced integral probability metric (IPM)
result Establish theoretical guarantees for the associated two-sample testing procedure
Kernel method tests if two sets of data are from the same distribution.
problem Two-sample hypothesis testing for high dimensional data with small samples.
method One-class set classification using Set Kernel and one-class SVM.
result The method achieves zero type-I and type-II error on all cancer gene expression data sets.
A neural network-based two-sample test improves classification accuracy.
problem Differentiating between two sub-exponential densities.
method Difference of logit function from trained classification neural network.
result Network complexity scales with intrinsic dimensionality for low-dimensional manifolds.
Efficiently tests two distributions with few label queries.
problem Two-sample test with limited label information.
method Three-stage framework: classifier training, bimodal query, FR test.
result Significantly reduces Type II error compared to uniform querying.
Semi-supervised method boosts two-sample testing with covariate data.
problem Two-sample testing with covariate information.
method Semi-supervised kernel test with asymptotic normality.
result Higher asymptotic power compared to existing methods.
Kernel test evaluates dynamical system data streams.
problem Evaluate if data streams from dynamical systems are from the same distribution.
method Proposes a novel kernel two-sample test for dynamical systems, addressing independence and autocorrelation challenges.
result Data-driven method with theoretical guarantees for anomaly detection.
Conformal C2ST turns weak classifiers into reliable two-sample tests.
problem Determining if two distributions are identical using weak classifiers.
method Developed conformal variants of the C2ST to convert any classifier scores into reliable p-values.
result Even weak classifiers can yield powerful and reliable two-sample tests.
Optimal tests for goodness of fit and two-sample problems using MMD and KSD.
problem Asymptotically optimal tests for goodness of fit and two-sample problems.
method Maximum Mean Discrepancy (MMD) and Kernel Stein Discrepancy (KSD) based tests.
result Optimal tests achieve the maximum exponential decay rate under specific conditions.
A novel kernel learning framework detects abrupt changes in time series data.
problem Detecting abrupt changes in time series data with fewer assumptions.
method KL-CPD, a novel kernel learning framework that optimizes a lower bound of test power via an auxiliary generative model.
result Significantly outperformed other state-of-the-art methods in benchmark datasets and simulation studies.
New test detects differences in heterogeneous datasets.
problem Detecting differences between two samples with unknown heterogeneity.
method Developed a nonparametric testing procedure that handles latent heterogeneity through a composite null.
result The test accurately detects differences in the presence of unknown heterogeneity.
New framework compares credal sets for hypothesis testing with epistemic uncertainty.
problem Comparing distributions with partial ignorance and epistemic uncertainty.
method Credal two-sample testing framework for convex sets of probability measures.
result Direct integration of epistemic uncertainty in hypothesis testing.
E-C2ST uses E-values for high-dimensional data two-sample tests.
problem Statistical testing for high-dimensional data.
method Combines split likelihood ratio tests and predictive independence tests, using E-values for anytime-valid sequential tests.
result E-C2ST achieves enhanced statistical power by partitioning datasets into multiple batches.
The paper designs tests for comparing ranked preference data and finds significant differences.
problem Comparing pairwise comparison and ranking data in various applications.
method Developed two-sample tests for pairwise comparison and ranking data, proving upper and lower bounds.
result Upper and lower bounds show tightness of the proposed tests, and significant differences in preferences were found.
Unified framework for global and local two-sample conditional distribution testing.
problem Testing equality of two conditional distributions.
method Distance and kernel methods, conditional U-statistics, local bootstrap.
result Developed reliable global and local tests.
Develops a two-sample test using projected Wasserstein distance to handle high-dimensional data.
problem Testing whether two high-dimensional samples come from the same distribution.
method Optimal projection to find a low-dimensional linear mapping that maximizes the Wasserstein distance between projected probability distributions.
result Characterizes the convergence rate of the projected Wasserstein distance and presents practical algorithms.
Meta two-sample testing uses auxiliary data to quickly find powerful tests from limited samples.
problem Challenges in identifying powerful kernels for distinguishing complex distributions with limited data.
method Introduces meta two-sample testing (M2ST) to leverage abundant auxiliary data on related tasks.
result Proposed algorithms improve over baselines and identify powerful tests from scarce observations.
New graph tests improve on existing methods for comparing large graphs.
problem Comparing large graphs from different sources.
method Proposed new tests based on asymptotic distributions.
result New tests are computationally less expensive and more reliable.
Proposes a new log-rank test using RKHSs for robust two-sample analysis.
problem Two-sample problem in right-censored data.
method Test statistic based on supremum of RKHS-weighted log-rank tests.
result Proposed test is omnibus for a specific family of RKHSs.
Paper bounds error in kernel MMD test using reference set.
problem Bounding error in kernel MMD test for two sample problem.
method Quantifies error using heat diffusion and reference point weights.
result Provides non-asymptotic, non-probabilistic error bound.
Develops a test for comparing linear models without assuming sparsity.
problem Testing equality of regression slopes in high-dimensional models.
method TIERS framework, self-normalization, ADDS estimator, plug-in approach.
result Robust test for equality of regression slopes under weak conditions.
A density ratio is defined by the ratio of two probability densities. We study the inference problem of density ratios and apply a semi-parametric density-ratio estimator to the two-sample homogeneity test. In the proposed test procedure, the f-divergence between two probability densities is estimated using a density-r…
Sequential tests for two-sample and independence testing using betting strategies.
problem Testing sequential data for two-sample and independence without kernel selection issues.
method Prediction-based betting strategies that adaptively determine distribution and joint distribution.
result Prediction-based tests outperform kernel-based approaches in high-dimensional or structured data settings.
Deep neural networks improve two-sample testing.
problem Efficiently distinguishing between two unknown distributions.
method Deep learning representations for two-sample testing.
result Significant reduction in type-2 error rate compared to existing methods.
Proposes counterfactual explanations for deep two-sample tests on high-dimensional data.
problem Limited interpretability of deep two-sample tests on high-dimensional data.
method Combines diffusion autoencoder and pretrained deep two-sample test model to generate counterfactuals.
result Counterfactual transformations increase p-values, indicating closer distribution similarity.
A permutation-based SW test achieves minimax-optimal power for two-sample testing.
problem Nonparametric two-sample testing using the sliced Wasserstein distance.
method Proposes a permutation-based SW test and analyzes its performance.
result Achieves minimax separation rate n−1/2 over multinomial and bounded-support alternatives. Proposes a method to select variables for kernel two-sample tests.
problem Determining whether two samples have the same distribution using informative variables.
method A framework based on kernel maximum mean discrepancy (MMD) for selecting a subset of variables.
result The sample size requirements for the three kernels depend on the number of selected variables, not the data dimension.
A new witness two-sample test improves data efficiency and power.
problem Nonparametric two-sample testing.
method Optimizes kernel and defines weights and basis points using training data.
result The new test is consistent, has well-controlled type-I error, and has comparable or higher power.
Optimizes two-sample tests for non-Euclidean domains using spectral regularization.
problem Optimizing two-sample tests for non-Euclidean domains.
method Spectral regularization of MMD test to achieve minimax optimality.
result Proposes a spectral regularization method that improves test optimality.
Paper proposes a scoring function for detecting anomalies in large datasets.
problem Detecting outliers in large, feature-rich datasets.
method Binary classification problem with a two-sample linear rank statistic.
result Empirical results show the effectiveness of the proposed scoring function.
Deep neural nets optimize kernel parameters for non-parametric two-sample tests.
problem Determining if two samples come from the same distribution.
method Deep kernels trained to maximize test power, adapting to distribution smoothness and shape.
result Deep kernels outperform simpler kernels in high dimensions and complex data.
C2ST uses classifiers to test if two datasets are from the same distribution.
problem Assessing if two datasets are from the same distribution.
method Construct a dataset by pairing examples from each dataset with positive or negative labels, then classify and evaluate accuracy.
result C2ST learns a suitable representation, has a simple null distribution, and can interpret differences between datasets.
Optimal tests for nonparametric one- and two-sample testing are derived using MMD and KSD.
problem Developing optimal tests for nonparametric one- and two-sample testing.
method Using Sanov's theorem and Maximum Mean Discrepancy (MMD), the optimal error exponents are derived for one-sample tests. For two-sample tests, the quadratic-time Kernel Stein Discrepancy (KSD) is shown to achieve the optimal type-II error exponent.
result Achievement of optimal error exponents for nonparametric one- and two-sample testing in the universal setting.
New method controls FDR for sparse GLMs, identifying positive and negative relationships.
problem Sparse GLMs with high-dimensional data and varying sample size.
method Debiased-Lasso estimator and CLIME method for precision matrix estimation.
result Asymptotically controls directional FDR and FDV for sparse GLMs.
New kernel tests detect differences between distributions exponentially quickly.
problem Characterize the asymptotic performance of kernel two-sample tests.
method Established exponentially consistent kernel two-sample tests for unknown distributions.
result Exponential decay rate of type-II error probability is optimal and independent of kernels.
Two-sample tests improve on existing methods for microtubule data.
problem Testing differences between two groups of filament data.
method Optimal lifts and manifold stability theorem applied to microtubule data.
result New tests outperform existing methods on simulated and real data.
Develops a new test for comparing two groups' densities, showing minimax optimality.
problem Comparing probability densities between two groups.
method Probabilistic tensor product smoothing spline framework for joint density modeling; penalized likelihood ratio test for interaction testing.
result Proposed test is minimax optimal and outperforms conventional approaches.
A test for comparing function samples using MMD.
problem Testing if two functional data samples come from the same distribution.
method Maximum Mean Discrepancy (MMD) for functional data, with theoretical scaling analysis.
result The proposed test is effective and robust to functional reconstructions.
Study on kernel tests for high-dimensional data, focusing on MMD and CLT.
problem Asymptotic behavior of kernel two-sample tests in high dimensions and large samples.
method Maximum mean discrepancy (MMD) with isotropic kernels, deriving asymptotic expansions and CLT.
result Interplay between moment discrepancy and dimension-and-sample orders in kernel tests.
A new method optimizes MMD test power by dynamically selecting kernels, overcoming traditional trade-offs.
problem Fixed kernels fail to distinguish certain distributions, leading to overfitting and variance collapse.
method Complexity-Penalized MMD (CP-MMD) criterion, derived from concentration inequality, optimizes kernel selection.
result CP-MMD maximizes true test power while ensuring unconditional Type-I validity, matching or exceeding state-of-the-art performance.
Improved two-sample testing using L1 geometry for analytic kernels.
problem Detecting differences between distributions.
method Use L1 distance between kernel-based distribution representatives to improve testing power. result Better detection of differences between distributions using L1 norm. New method relaxes TV distance for two-sample testing without distributional assumptions.
problem Challenges in certifying equality or providing tight bounds on TV distance for two distributions.
method Examined blurred total variation distance, a relaxation of TV distance.
result Provided theoretical guarantees for upper and lower bounds on blurred TV distance.
Efficiently tests two distributions using Nyström approximation of MMD.
problem Testing whether two sets of data are from the same distribution in large-scale scenarios.
method Nyström approximation of maximum mean discrepancy (MMD) for scalable testing.
result Finite-sample bound on power of the test for sufficiently separated distributions.