New method detects concept drifts with fewer labels.
problem Real-world data drifts over time, affecting model performance.
method Hierarchical Hypothesis Testing with Request-and-Reverify strategy.
result Significant reduction in label requests with improved performance.
Bayesian model compares classifier accuracies across multiple datasets.
problem Shortcomings of null hypothesis significance tests in comparing classifier accuracies.
method Bayesian hierarchical model analyzing cross-validation results.
result Posterior probability of classifier accuracies being equivalent or different.
Randomized hierarchical clustering tests for stability and detects clusters.
problem Greedy hierarchical clustering's sensitivity to data perturbations.
method Randomization scheme and p-values at each node.
result Valid hypothesis testing procedures for clustering results.
Paper detects and adapts to concept drifts in streaming data.
problem Concept drifts deteriorate classification performance over time.
method Hierarchical Hypothesis Testing (HHT) framework for detection and adaptation.
result HLFR detects and adapts to various concept drift types.
Hierarchical geodesic model for analyzing shapes on manifolds.
problem Analyzing temporal observations on manifold-valued data.
method Adapted functional-based metric for efficiency; variational time discretization of geodesics.
result Performed hypothesis tests and estimated mean trends in longitudinal analysis.
NeurT-FDR controls FDR by incorporating feature hierarchy.
problem Controlling FDR in complex, large-scale hypothesis testing problems.
method NeurT-FDR uses a neural network to parametrize test-level covariates and a regression framework to adjust feature hierarchy.
result NeurT-FDR makes substantially more discoveries than competitive baselines.
Proposes selective inference for testing differences in means between clusters.
problem Inflated type I error rate when testing differences in means between clusters.
method Selective inference approach to control selective type I error rate.
result Controls selective type I error rate by accounting for data-driven cluster definition.
Identifies important features and their resolution for complex models.
problem Characterizing complex learned models' decision-making across instance distributions.
method Model-agnostic approach using hypothesis testing and feature groups.
result Determines important features and their resolution levels for model accuracy.
New method tests weighted networks without thresholding, improving accuracy.
problem Testing and anomaly detection on weighted network data.
method Hierarchical Bayesian hypothesis testing framework for weighted networks.
result Method shows lower Type I error and higher statistical power compared to alternatives.
HyPE improves sample efficiency in DRL by discovering objects and hierarchies of skills.
problem Poor sample efficiency in DRL methods, especially in complex tasks.
method HyPE algorithm that discovers objects and generates hypotheses about their controllability, learning a hierarchy of skills.
result HyPE learns high-scoring policies an order of magnitude faster than state-of-the-art methods.
Recent advances in statistical theory, together with advances in the computational power of computers, provide alternative methods to do mass-univariate hypothesis testing in which a large number of univariate tests, can be properly used to compare MEEG data at a large number of time-frequency points and scalp location…
Extends linear representation hypothesis to categorical and hierarchical concepts in LLMs.
problem Representing concepts without natural contrasts in large language models.
method Formalizes linear representation hypothesis for categorical and hierarchical concepts, proving relationships between concept hierarchy and representation geometry.
result Validated theoretical results on large language models, estimating representations for 900+ concepts.
New method uses LLMs to generate detailed scientific hypotheses.
problem Generating detailed, actionable scientific hypotheses from coarse initial directions.
method Hierarchical search method that incrementally adds details to hypotheses.
result Hierarchical search method consistently outperforms strong baselines on expert-annotated hypotheses.
New framework for domain adaptation using hierarchical optimal transport.
problem Improving domain adaptation when source and target data distributions differ.
method Proposes a new theoretical framework and hierarchical Wasserstein distance.
result Provides more explicit generalization bounds and aligns specific structures for successful adaptation.
A framework for hypothesis testing on attributed graphs using sampling.
problem Statistical testing on graph data, especially large attributed graphs.
method Sampling-based framework with PHASE and PHASEopt for accurate and efficient hypothesis testing.
result PHASE and PHASEopt improve accuracy and efficiency of hypothesis testing in attributed graphs.
This paper proves the necessity and effectiveness of learning the prior in VAEs.
problem Aggregated posterior may not match unit Gaussian prior, leading to poor variational inference.
method Proves necessity and effectiveness of learning the prior, analyzes why it's needed, and proposes hypothesis.
result Learning the prior can improve reconstruction loss and achieve comparable test NLL to deep hierarchical VAEs.
The paper sets thresholds for testing correlation in hypergraphs, distinguishing between independent and correlated states.
problem Testing correlation between two hypergraphs under different models.
method Derives sharp information-theoretic thresholds for distinguishing between null and alternative hypotheses.
result The testing threshold decreases as the hypergraph's uniformity (m) increases, making correlation testing easier for higher uniformity.
Study robust hypothesis testing under Hellinger distance, proving lower bounds and providing tests.
problem Testing close variants of specified distributions robustly to Hellinger distance.
method Lower bound on slack factor, testing with Hellinger balls, symmetric chi-squared distance analysis.
result Lower bound on slack factor quantifies robustness under misspecification.
Paper resolves open problems on sample complexity in binary hypothesis testing.
problem Open problems in distributed simple binary hypothesis testing under information constraints.
method One-shot lower bound on Bayes error, streamlined sample complexity formula, reverse data-processing inequality.
result Optimally tight sample complexity bounds for communication-constrained simple binary hypothesis testing.
hyppo simplifies multivariate hypothesis testing in Python.
problem Inconsistent multivariate hypothesis testing interfaces in Python.
method Unified library for multivariate testing procedures.
result Easy-to-use and flexible for future extensions.
Transforms any test into anytime-valid with sample savings.
problem Sequential data invalidates classical test guarantees.
method Predicts test outcomes to create anytime-valid stopping rules.
result Ensures Type-I error control and near-optimal power.
Paper proposes a new framework for hypothesis testing in imaging.
problem Challenges in hypothesis testing for imaging data.
method Combines self-supervised imaging, vision-language models, and non-parametric hypothesis testing.
result Demonstrates improved power and robust error control in image-based phenotyping.
Paper proposes a robust hypothesis testing method using Sinkhorn distance.
problem Hypothesis testing for small samples.
method Data-driven approach using Sinkhorn uncertainty sets.
result The method provides a more flexible detector compared to Wasserstein robust test.
Diffusion models generalize better with hierarchical data structure and regularization.
problem Understanding generalization in diffusion models with finite data.
method Analyzing diffusion models through data covariance spectra and developing a theoretical framework based on linear neural networks.
result Generalization in diffusion models improves with hierarchical data structure and regularization.
Study hypothesis testing under quantized samples with communication constraints, achieving near-optimal sample complexity.
problem Optimizing hypothesis testing with quantized samples and communication constraints.
method Developed a polynomial-time algorithm achieving near-optimal sample complexity under communication constraints.
result Achieved near-optimal sample complexity under communication constraints, with a logarithmic factor increase over unconstrained setting.
New private algorithm for sequential hypothesis testing with privacy and error rate guarantees.
problem Privacy protection in sequential hypothesis testing for sensitive data.
method Renyi differential privacy, Wald's Sequential Probability Ratio Test (SPRT).
result Private algorithm with strong privacy guarantees and theoretical performance analysis.
Robust test for distributions under Hellinger distance, simpler than optimal tests.
problem Testing and estimating distributions robustly under Hellinger distance.
method Simple robust hypothesis test with optimal sample complexity, robust to Hellinger distance perturbations.
result Empirically demonstrated robustness and power of the test on canonical distributions.
Paper optimizes hypothesis verification in sequential experiments.
problem Maximizing confidence in a verified hypothesis after exploration.
method Formulated as a confidence maximization problem in a POMDP, characterized optimal solutions, and proposed a heuristic.
result Heuristic performs better than existing methods in some scenarios.
Neural networks improve language comprehension by testing and refining hypotheses.
problem Improving language comprehension models through more sophisticated reasoning.
method Memory augmented neural networks with a hypothesis testing loop.
result Achieved state-of-the-art results on language comprehension benchmarks.
Unified Bayesian framework improves clinical trial hypothesis testing.
problem Lack of transparency and inability to quantify evidence in traditional P-values.
method Interval null hypothesis framework combined with Bayes factor-based tests.
result Bayesian interval hypothesis testing ensures frequentist error control and interpretability.
Study on hypothesis testing games with adversarial classification, showing convergence rates.
problem Adversarial classification in hypothesis testing.
method Mixed strategy Nash equilibria analysis, concentration phenomena examination.
result Exponential rates of convergence of classification errors at equilibrium.
The problem of multiple hypothesis testing arises when there are more than one hypothesis to be tested simultaneously for statistical significance. This is a very common situation in many data mining applications. For instance, assessing simultaneously the significance of all frequent itemsets of a single dataset entai…
New framework compares credal sets for hypothesis testing with epistemic uncertainty.
problem Comparing distributions with partial ignorance and epistemic uncertainty.
method Credal two-sample testing framework for convex sets of probability measures.
result Direct integration of epistemic uncertainty in hypothesis testing.
New framework improves text watermark detection under imperfect pseudorandomness.
problem Structured dependence in generated text from language models causes Type I error control issues.
method Hierarchical two-layer partition, minimal units, non-asymptotic efficiency measure, minimax hypothesis testing.
result Closed-form optimal rules for watermark detection under imperfect pseudorandomness.
Develops GLRT for defending against adversarial attacks in hypothesis testing.
problem Adversarial attacks on machine learning models causing misclassification.
method Generalized likelihood ratio test applied to composite hypothesis testing problem.
result GLRT approach yields competitive robustness-accuracy tradeoff under various attacks.
Efficiently annotates hierarchical structure in images using 2AFC testing and deep metric learning.
problem Lack of efficient methods for hierarchical annotation of high-dimensional data like images.
method Two-alternative-forced-choice (2AFC) testing and deep metric learning for embedding data in semantic space.
result Successfully hierarchically clusters data, achieving finer granularity than original labels.
Kernel tests for set-valued data improve hypothesis testing accuracy.
problem Testing distributions of sets with varying sizes, noise, and nuisance variability.
method Interpreting sets as samples from latent distributions and using kernel methods for testing.
result Kernel tests outperform traditional methods in synthetic and real-world experiments.
Framework for online hypothesis testing across various data types.
problem Testing various nonparametric hypotheses in data streams.
method Unified framework using operators on data distributions, leveraging ML models.
result Efficient, adaptive, and error-controlled sequential tests.
New method tests linear hypotheses in high-dimensional models without sparsity assumptions.
problem Testing linear hypotheses in high-dimensional models without restrictive assumptions.
method Proposes a test based on restructured regression with transformed and augmented features.
result Asymptotically exact control on Type I error without sparsity assumptions.
Proposes a new model for testing causal structural priors and synthesizing data.
problem Testing and synthesizing causal structural priors using nonparametric knowledge and neural networks.
method Causal Structural Hypothesis Testing (C-SHT) and Causal Structural Variational Hypothesis Testing (C-SVHT) using deep neural networks.
result Demonstrates out-of-distribution generalization error as a proxy for causal structural prior hypothesis testing.
This paper explains differential privacy through hypothesis testing and analyzes its relaxations.
problem Understanding the hypothesis testing interpretation of differential privacy.
method Identifying conditions for a statistical divergence to satisfy a similar interpretation and analyzing relaxations of differential privacy based on Renyi divergence.
result Improved conversion rules between differential privacy and its relaxations based on Renyi divergence.
Formula derived for sample complexity in binary hypothesis testing.
problem Determine the minimum number of samples to distinguish between two distributions.
method Developed a formula for sample complexity in both prior-free and Bayesian settings, using Jensen-Shannon and Hellinger divergences.
result Formula characterizes sample complexity for a wide range of error parameters, up to multiplicative constants.
Study hypothesis testing for noisy Markov chain samples.
problem Hypothesis testing between two discrete distributions via noisy Markov chain samples.
method Derive instance-dependent minimax rates and analyze spectral properties of the Markov chain.
result Wide statistical window in sample complexity for different initial distributions.
Develops a hypothesis testing framework for generalized Thurstone models.
problem Determining whether pairwise comparison data fits a generalized Thurstone model.
method Introduces separation distance and derives upper and lower bounds for testing.
result Critical threshold for testing depends on observation graph topology and scales as Θ((nk)−1/2) for complete graphs. New framework for valid hypothesis testing in complex data settings.
problem Challenges in classical hypothesis testing frameworks.
method Add and subtract external noise to partition data, orthogonalize, and test hypotheses.
result Valid hypothesis tests can be conducted under minimal assumptions.
A heuristic framework tests the multi-manifold hypothesis in empirical data.
problem Overestimation of parameters in global linear models.
method Heuristic multiscale framework using spline-interpolated manifolds.
result Validates the multi-manifold hypothesis in empirical data.
This paper resolves the test for Markov regime switching models' regime number.
problem Testing the number of regimes in Markov regime switching models.
method Derives the asymptotic distribution of the likelihood ratio test statistic.
result Establishes the asymptotic validity of the parametric bootstrap.
The paper develops a robust test for nonlinear effects using Gaussian processes.
problem Detecting nonlinear interactions between continuous features.
method Hypothesis test based on Gaussian processes, robust to kernel mis-specification, and using ensemble estimators.
result Demonstrates interesting connections between machine learning and statistical inference.