Valid p-value for bounded random variables without distributional assumptions.
problem Calibration of predictive algorithms in a distribution-free setting.
method Built a super-uniform p-value based on a concentration inequality.
result Super-uniform p-value is tighter than existing alternatives.
Unified theoretical guarantees for distribution-free changepoint detection and testing.
problem Distribution-free changepoint inference with finite-sample validity and consistency.
method Distribution-free changepoint localization using conformal p-values with theoretical guarantees.
result Unified distribution-free guarantees for changepoint detection, localization, and testing.
The paper introduces localized conformal p-values for conditional testing problems.
problem Addressing conditional testing problems in statistics.
method Localized conformal p-values defined by inverting prediction intervals.
result Proposes procedures for conditional outlier detection and label screening with FDR and FWER control.
The paper develops p-values for outlier detection using conformal inference.
problem Detecting outliers in nonparametric data sets.
method Conformal inference framework for constructing marginally valid but mutually dependent p-values.
result Valid p-values for outlier detection with conditional independence and marginal false discovery rate control.
OptCS optimizes model selection after conformal inference, controlling FDR and power loss.
problem Challenges in model selection for conformal inference, especially when limited labeled data and many model choices are available.
method OptCS framework that allows valid statistical testing after flexible data-driven model optimization, using novel multiple testing procedures.
result Valid conformal p-values constructed despite substantial data reuse, maintaining FDR control.
A communication-efficient method controls FDR in network settings.
problem Controlling FDR in networks with limited communication.
method Sample-and-Forward: a flexible procedure for multihop networks.
result Nodes can control FDR without sharing p-values, achieving power and FDR control.
A new method for anomaly detection adapts to local non-stationarity in low-data regimes.
problem Adapting conformal anomaly detection to handle distribution shifts in real-world data.
method Proposes a continuous inference relaxation using continuous weighted kernel density estimation to decouple local adaptation from tail resolution.
result Restores detection capabilities and statistical power in low-data regimes while maintaining valid error control.
AdaPT-GMM improves multiple testing power with covariates.
problem Powerful and robust multiple testing with covariates.
method Covariate-assisted Gaussian mixture model with adaptive thresholding.
result AdaPT-GMM delivers high power in various scenarios.
Conformal C2ST turns weak classifiers into reliable two-sample tests.
problem Determining if two distributions are identical using weak classifiers.
method Developed conformal variants of the C2ST to convert any classifier scores into reliable p-values.
result Even weak classifiers can yield powerful and reliable two-sample tests.
Paper proposes AdaDetect for FDR-controlled novelty detection.
problem Semi-supervised novelty detection with probabilistic classification.
method Data-adaptive learning of transformation to control FDR.
result Control of false discovery rate on detected novelties.
CROC identifies the earliest-changing stream as the root cause in multi-stream data.
problem Distribution-free root cause analysis in multi-stream data with unknown distributional changes.
method Conformal p-values and finite-sample valid confidence sets.
result CROC efficiently isolates the root cause under minimal assumptions.
We present the expected values from p-value hacking as a choice of the minimum p-value among m independents tests, which can be considerably lower than the "true" p-value, even with a single trial, owing to the extreme skewness of the meta-distribution. We first present an exact probability distribution (meta-distrib…
Selective inference controls Type I error in k-means clustering tests.
problem Inflated Type I error in classical hypothesis tests for k-means clusters.
method Selective inference approach to control Type I error.
result Proposes a computable finite-sample p-value for selective inference.
GAIF enhances online multiple testing with feedback, improving statistical power.
problem Sequential online multiple testing with delayed feedback.
method GAIF framework using dynamic threshold adjustment and feedback-driven model selection.
result Improves statistical power through feedback-driven model selection.
Estimates peeking effects in p-values to correct bias.
problem Data peeking biases reported p-values downward.
method Develops mechanisms to estimate running extrema of test statistics.
result Corrects bias in p-values due to peeking.
Proposes a statistical test for VAE-based anomaly detection reliability.
problem Ensuring reliability of anomaly detection in high-stakes applications.
method Variance Autoencoder (VAE) Test based on selective inference.
result Validates VAE-based anomaly detection with p-values controlling false detection probability.
We improve prediction set coverage by assigning weights to individual sets.
problem Aggregating multiple prediction sets weakens overall coverage guarantee.
method Propose a framework for weighted aggregation of prediction sets.
result Achieve tighter coverage bounds that interpolate between 1−2α and 1−α guarantees. The PC algorithm allows investigators to estimate a complete partially directed acyclic graph (CPDAG) from a finite dataset, but few groups have investigated strategies for estimating and controlling the false discovery rate (FDR) of the edges in the CPDAG. In this paper, we introduce PC with p-values (PC-p), a fast al…
The paper provides high-probability bounds on false discovery proportions in conformal inference.
problem Existing methods fail to provide high-probability bounds on the realized false discovery proportion.
method Constructing a high-probability envelope for the empirical distribution function of null conformal p-values by sampling from their joint distribution.
result Establishes finite-sample, distribution-free upper bounds on the FDP that hold simultaneously over all possible rejection thresholds.
aLTT selects hyperparameters efficiently with statistical guarantees.
problem Statistical validity and efficiency in hyperparameter selection.
method Sequential data-dependent multiple hypothesis testing with early termination.
result Reduces testing rounds while maintaining statistical validity.
New method converts p-values to e-values for more efficient CP and aggregation.
problem Limitations of existing p-to-e calibrators in CP setting.
method Proposes a novel P2E calibrator for set-preserving calibration.
result Significant efficiency gains over existing p-to-e calibrators.
We propose a differentiable sigmoid function for efficient p-value calculation in clustering.
problem Efficient and accurate p-value calculation for clustering algorithms.
method Designed a differentiable sigmoid function to approximate the Dip-p-value transformation.
result Accelerates computation and integrates well with gradient descent-based learning schemes.
Proposes mCS for multivariate selection with FDR control.
problem Selecting high-quality candidates from multivariate datasets.
method Introduces regional monotonicity and multivariate nonconformity scores.
result Significantly improves selection power with FDR control.
We propose a general method for constructing hypothesis tests and confidence sets that have finite sample guarantees without regularity conditions. We refer to such procedures as "universal." The method is very simple and is based on a modified version of the usual likelihood ratio statistic, that we call "the split li…
We provide a new local class-purity theorem for Lipschitz continuous DNN classifiers. In addition, we discuss how to achieve classification margin for training samples. Finally, we describe how to compute margin p-values for test samples.
Many model selection algorithms produce a path of fits specifying a sequence of increasingly complex models. Given such a sequence and the data used to produce them, we consider the problem of choosing the least complex model that is not falsified by the data. Extending the selected-model tests of Fithian et al. (2014)…
The paper develops distribution-free methods for ordinal classification.
problem Constructing valid prediction sets for ordinal classification problems.
method Leveraging conformal prediction and multiple testing with FWER control.
result The proposed methods achieve satisfactory levels of marginal and class-specific conditional coverages.
Paper develops a new method for open-set and imbalanced classification with valid prediction sets.
problem Tackles open-set and imbalanced classification with new prediction methods.
method Develops a new family of conformal p-values and a selective sample splitting algorithm.
result Valid prediction sets with valid coverage in open-set scenarios and informative predictions under extreme class imbalance.
Null-Calibrated Conformal Selection via Target-Membership Scores
problem Identifying test candidates whose unknown responses fall in a target region while controlling the false discovery rate
method Membership-score-based conformal selection
result Finite-sample valid null p-values
Structural equation models and Bayesian networks have been widely used to study causal relationships between continuous variables. Recently, a non-Gaussian method called LiNGAM was proposed to discover such causal models and has been extended in various directions. An important problem with LiNGAM is that the results a…
This paper improves prediction accuracy for multi-input classification tasks using p-value aggregation.
problem Generating accurate predictive sets with guaranteed coverage for multi-input classification tasks.
method Integrates p-values from each observation to reduce the size of the predicted label set while maintaining class-conditional coverage.
result The method reduces the size of the predicted label set while preserving the required coverage guarantee.
Conformal Test Martingales can be 'blind' to significant changes in data distribution.
problem The converse of exchangeability does not hold, leading to potential blindness of CTMs.
method Explicit construction of A-cryptic change-point using bivariate Gaussian distributions. result CTMs can be perfectly cryptic to a significant change in marginal means.
Statistical guarantees for hyperparameter selection
problem Hyperparameter selection in AI systems
method Learn-then-test framework
result Provable reliability and safety
Unified Bayesian framework improves clinical trial hypothesis testing.
problem Lack of transparency and inability to quantify evidence in traditional P-values.
method Interval null hypothesis framework combined with Bayes factor-based tests.
result Bayesian interval hypothesis testing ensures frequentist error control and interpretability.
A method to split a data point into two parts that individually cannot reconstruct the whole, but together can.
problem Splitting a single data point into two parts such that neither can reconstruct the whole but together can.
method Borrowing ideas from Bayesian inference to achieve a continuous analog of data splitting.
result A method to achieve data fission, enabling post-selection inference in finite samples.
Assigning significance in high-dimensional regression is challenging. Most computationally efficient selection algorithms cannot guard against inclusion of noise variables. Asymptotically valid p-values are not available. An exception is a recent proposal by Wasserman and Roeder (2008) which splits the data into two pa…
Proposes a method to quantify the reliability of salient regions in deep learning models using p-values.
problem Difficulty in assessing the reliability of saliency maps generated by deep learning models.
method Proposes a selective inference framework to quantify the reliability of salient regions as selected hypotheses by deep learning models.
result The method can provably control the probability of false positive detections of salient regions.
Image segmentation is one of the most fundamental tasks of computer vision. In many practical applications, it is essential to properly evaluate the reliability of individual segmentation results. In this study, we propose a novel framework to provide the statistical significance of segmentation results in the form of …
Machine-learned anomaly detection in new-physics searches needs calibration and look-elsewhere correction
problem Machine-learned anomaly detection in new-physics searches
method Conformal prediction layer
result Calibrated significance with distribution-free guarantees
This paper introduces an approach for detecting differences in the first-order structures of spatial point patterns. The proposed approach leverages the kernel mean embedding in a novel way by introducing its approximate version tailored to spatial point processes. While the original embedding is infinite-dimensional a…
A method selects candidates based on predictions with statistical control.
problem Screening candidates for resource-intensive steps like hiring or drug discovery.
method Wraps around any prediction model to produce a subset of candidates with controlled false selection rate.
result Empirically demonstrates selection of candidates whose predictions exceed a data-dependent threshold.
The paper confirms two groups of gamma-ray bursts using a new nonparametric metric.
problem Determining the number of inherent groups in gamma-ray bursts.
method A new nonparametric interpoint distance-based measure, combined with clustering methods.
result Confirms two groups of short and long gamma-ray bursts.
New methods improve anomaly detection with reduced false positives.
problem Effective anomaly detection with controlled error rates.
method Leave-one-out-, bootstrap-, and cross-conformal anomaly detection methods.
result Improved anomaly detection with reduced false positives.
Logistic regression is used thousands of times a day to fit data, predict future outcomes, and assess the statistical significance of explanatory variables. When used for the purpose of statistical inference, logistic models produce p-values for the regression coefficients by using an approximation to the distribution …
New method tests Granger non-causality in panel data with cross-sectional dependencies.
problem Testing Granger non-causality in panel data with cross-sectional dependencies.
method Proposes a new approach to aggregate p-values from panel members to test Granger non-causality, showing lower FDR.
result Our approach discovers true causal relations in panel data, unlike state-of-the-art methods.
E-values enhance conformal prediction methods.
problem Distribution-free uncertainty quantification.
method Reformulation of conformal prediction using e-values.
result E-values offer new theoretical and practical capabilities.
Logistic regression for brain imaging without p-values.
problem Computing the distribution of random field suprema is hard.
method Uses logistic regression for brain network classification.
result Performs classification at each edge level without preselected features.
Nonparametric IPSS selects features with false discovery control.
problem Feature selection in high-dimensional data with theoretical false discovery control.
method Integrated Path Stability Selection (IPSS) applied to nonparametric feature importance scores.
result IPSS accurately controls false discovery rate and detects more true positives than existing methods.