Valid inference from data and predictions.
problem Valid statistical inference with machine learning predictions.
method Framework for valid inference using machine learning predictions.
result Valid confidence intervals without assumptions on predictions.
Paper extends prediction-powered inference using conformal prediction for robust and valid imputation.
problem Safe use of black-box ML models for imputing missing data with strong guarantees.
method Connecting prediction-powered inference with conformal prediction for valid and additional guarantees.
result First general prediction-powered procedure for e-values operating off-line.
Predictive e-values enhance statistical inference across various tasks.
problem Insufficient data limits traditional statistical inference.
method Apply prediction-powered inference to e-values.
result Every e-value-based inference has a prediction-powered counterpart.
PPBoot simplifies prediction-powered inference.
problem Prediction-powered inference problems.
method Bootstrap-based method for arbitrary estimation problems.
result PPBoot often performs nearly identically to PPI(++).
FPPI selectively uses predictions to improve inference efficiency.
problem Improving statistical inference with limited labeled data and heterogeneous prediction quality.
method Filtered Prediction-Powered Inference (FPPI) framework.
result FPPI achieves strictly improved asymptotic efficiency compared to existing methods.
FAB-PPI uses prior knowledge to improve prediction-powered inference.
problem Improving statistical inference with machine learning predictions.
method Informing PPI with prior knowledge on prediction quality.
result FAB-PPI improves inference accuracy and confidence intervals.
PPI++ uses machine learning predictions to improve inference from small datasets.
problem Efficient inference from small labeled datasets with high-quality predictions.
method Adapts prediction-powered inference (PPI) to compute confidence sets for any parameter dimensionality.
result Improves classical intervals using only labeled data, always yielding better results.
Extends PPI to sequential setting, improving inference over time.
problem Sequential data growth with unlabelled data.
method Prediction-powered confidence sequence procedures using Ville's inequality and the method of mixtures.
result Asymptotically valid uniformly over time, accommodating prior knowledge.
StratPPI improves prediction-powered inference with stratified sampling.
problem Improving statistical estimates with limited human-labeled data.
method Combining small human-labeled data with large automatic-labeled data, stratifying data for tighter confidence intervals.
result StratPPI provides substantially tighter confidence intervals than unstratified approaches.
A framework uses a mixture of predictors for semi-supervised inference.
problem Limited labeled data, abundant unlabeled data.
method Mixture of Experts (MOE) for semi-supervised inference.
result MOE-powered inference framework achieves smallest possible variance.
Improved local multivariable regression for better inference with limited data.
problem Limited sample size hampers local polynomial/multivariable regression.
method Prediction-Powered Inference (PPI) algorithm for local multivariable regression.
result Significantly reduces estimation variance without increasing error.
PPI uses predictions and weighting to infer from partially labeled data.
problem Valid inference with partially labeled data.
method Combines model-based predictions with bias correction from labeled data, using Horvitz-Thompson and Hájek corrections.
result IPW-adjusted PPI with estimated propensities performs similarly to known-probability case.
Study improves treatment effect estimation using unlabeled covariates.
problem Estimating treatment effects with limited labeled data.
method Developed efficiency bounds and estimators for semi-supervised setting.
result Estimators using unlabeled covariates have lower asymptotic variance.
Generalizes prediction-powered inference for binary classifier evaluation.
problem Evaluation of binary classifiers with partially observed outcomes.
method Generalizes PPI to any regular asymptotically linear estimator and proposes modified estimators for covariate shift.
result PPI can be a computationally-simple alternative to existing methods, achieving no greater than the semi-parametric efficiency lower bound in certain scenarios.
AM-PPI uses multiple predictors to reduce label cost in healthcare AI.
problem Reduces label cost in post-deployment monitoring of healthcare AI.
method Combines model predictions with a small labeled sample, routing each instance to a cost-appropriate subset of predictors.
result Produces narrower confidence intervals than single-predictor methods.
PPI uses proxy data to improve inference from limited labels across related tasks.
problem Statistical inference with limited labels across multiple related tasks.
method Prediction-powered inference framework that uses cross-task recalibration to improve power and accuracy.
result Cross-task recalibration can substantially reduce confidence interval widths when labels are scarce.
PAS improves estimation of multiple means using ML predictions and shrinkage.
problem Improving statistical estimates with limited gold-standard data and noisy ML predictions.
method Prediction-Powered Adaptive Shrinkage (PAS) that combines PPI with empirical Bayes shrinkage.
result PAS adapts to the reliability of ML predictions and outperforms traditional methods in large-scale applications.
Unified approach combines prediction-powered inference and variance reduction for semi-supervised optimization.
problem Scarcity of labeled data in semi-supervised optimization.
method PPI-SVRG, combining PPI and SVRG methods.
result Unified convergence bound with improved performance under label scarcity.
New method uses machine learning to improve statistical inference.
problem Performing inference on conditional functionals with scarce labeled data.
method Combines localization with prediction-based variance reduction.
result Valid and sharp confidence intervals for conditional functionals.
Cross-prediction improves inference from small labeled datasets.
problem Valid inference from small labeled datasets with imperfect predictions.
method Imputes missing labels via machine learning and debiases predictions.
result Inferences achieve desired error probability and are more powerful.
PPI uses predictions to improve inference from incomplete data.
problem Incomplete or costly-to-measure outcomes in research fields.
method Leverages large unlabeled datasets for improved statistical efficiency with bias correction.
result PPI variants produce tighter confidence intervals than complete-case analysis.
Paper establishes statistical inference for performative predictions.
problem Dynamic influence of predictions on their targets.
method End-to-end framework for estimation and inference under performativity.
result Established central limit theorem for performative settings.
Over the years, ensemble methods have become a staple of machine learning. Similarly, generalized linear models (GLMs) have become very popular for a wide variety of statistical inference tasks. The former have been shown to enhance out- of-sample predictive power and the latter possess easy interpretability. Recently,…
PPI uses survey sampling methods for inference, bridging ML and statistics.
problem Combining machine learning predictions with small labeled data for valid inference.
method Equivalence of PPI estimators to survey sampling methods.
result PPI estimators are algebraically equivalent to survey sampling methods.
New method uses predictions to infer causal effects without labeled data.
problem Data labeling costs limit causal inference experiments.
method Prediction-Powered Causal Inferences (PPCI) using conditional calibration and transfer constraints.
result Valid causal inference achieved on experiments with no human annotations.
Calibrated Prediction-Powered Inference improves semisupervised mean estimation by calibrating prediction scores.
problem Semisupervised mean estimation with a small labeled sample and a large unlabeled sample, and miscalibrated prediction models.
method Calibrated Prediction-Powered Inference (Calibeating) post-hoc calibrates the prediction score on the labeled sample before using it for semisupervised estimation.
result Calibrated Prediction-Powered Inference can improve the original score both as a predictor of the outcome and as a regression adjustment for semisupervised inference.
We rebias estimates to improve interval calibration and prediction accuracy.
problem Constructing accurate intervals for noisy and biased estimates.
method Empirical Bayes rebiasing strategy that learns bias distribution from data.
result Substantial precision gains in prediction-powered inference.
Prediction-powered causal inference achieves smaller asymptotic variance than traditional methods.
problem Estimating causal and structural parameters in a semi-supervised setting.
method Combining efficient influence function with debiased machine learning and semi-supervised Riesz regression.
result Asymptotic variances of estimators match the derived efficiency bound.
Improved statistical inference for expensive data using machine learning predictions.
problem Statistical inference under adaptive two-phase multiwave sampling with expensive measurements.
method Multiwave Predict-Then-Debias estimator combining proxy information and expensive measurements.
result Valid estimators and confidence intervals for M-estimation under adaptive sampling.
New method improves deep learning models' uncertainty estimates.
problem Overconfidence in deep learning predictions.
method Develops a novel training algorithm using conformal inference.
result Produces more reliable uncertainty estimates without sacrificing accuracy.
PPI++ outperforms gold-standard labels only if pseudo-labels are highly correlated.
problem Optimizing statistical estimation using noisy pseudo-labels.
method Exact finite-sample analysis of PPI++ on mean estimation problem.
result PPI++ has provably worse estimation error than gold-standard labels alone in some settings.
New method uses AI predictions as cheaper alternatives to expensive outcomes.
problem Using expensive outcomes for statistical inference.
method Recalibrated prediction-powered inference using machine learning techniques.
result Significant gains in effective sample size over existing PPI proposals.
Improves risk control in predictions using semi-supervised calibration.
problem Noisy hyper-parameter tuning from limited labeled data.
method Semi-supervised calibration using unlabeled data to tune hyper-parameters rigorously.
result Improves prediction accuracy without sacrificing statistical validity.
New method combines multiple datasets to estimate ATE with valid confidence intervals.
problem Combining multiple observational datasets to estimate ATE with valid confidence intervals.
method Prediction-powered inferences to shrink CIs and provide valid CIs.
result Valid confidence intervals for ATE from multiple datasets.
New methods improve causal inference generalization using trial and observational data.
problem Limited trial data makes generalizing causal inferences to target populations statistically infeasible.
method Develops algorithms that combine trial and observational data to estimate complex nuisance functions.
result Improves generalization of causal inferences when the additional observational study is high-quality.
Generative Augmented Inference improves AI-generated data for causal inference.
problem Challenges in using AI-generated annotations for reliable causal inference.
method Generative Augmented Inference (GAI) treats AI outputs as informative features for learning true labels, flexibly modeling the relationship using nonparametric methods.
result GAI significantly reduces estimation error and improves confidence interval quality compared to human-only and PPI-based methods.
Hybrid Amortized Inference improves PPG model interpretability.
problem Tension between PPG biomarker accuracy and clinical interpretability.
method Introduces PPGen for biophysical PPG signal-physiological parameter relation, and HAI for fast, robust estimation.
result Hybrid Amortized Inference accurately infers physiological parameters from PPG signals.
New definition reveals encoding explanations that retain predictive power.
problem Challenges in evaluating and identifying encoding explanations.
method Developed a definition of encoding based on conditional dependence.
result Existing evaluation scores do not rank non-encoding explanations correctly, but STRIPE-X does.
The paper proposes a method to reliably select design algorithms for machine learning-guided design tasks.
problem Choosing the right design algorithm for machine learning-guided design tasks.
method Combining designs' predicted property values with held-out labeled data to reliably forecast characteristics of the label distributions produced by different design algorithms.
result The method is guaranteed to return design algorithms that yield successful label distributions.
Bayesian neural networks can be partially stochastic without losing predictive power.
problem The necessity of fully stochastic parameters in Bayesian neural networks.
method Theoretical and empirical investigation of partially stochastic networks compared to fully stochastic ones.
result Expressive predictive distributions require only small amounts of stochasticity, and partially stochastic networks can match or outperform fully stochastic networks.
SAGE quantifies feature importance in machine learning models.
problem Understanding the role of individual features in complex models.
method Formalizing predictive power through model-based and universal measures, and introducing SAGE for efficient calculation.
result SAGE assigns more accurate feature importance values than other methods.
This research evaluates measures of dependence for financial time-series data.
problem Accurately preparing time series data and selecting an appropriate measure of dependence is challenging.
method Review and establishment of a comprehensive analysis framework for shaping time-series data and evaluating measures of dependence.
result A method, framework, and example for selecting and evaluating a suitable measure of dependence are presented.
New method handles missing data using AI for efficient inference.
problem Parameter estimation and inference with blockwise missing data.
method Tractable solution using AI models and semiparametric theory.
result IBM(RAY) and IBM(Adaptive) estimators achieve efficiency gains.
New methods improve inference with scarce labels using regression.
problem Efficient inference with limited labeled data.
method Relates PPI++ to ordinary least squares regression and uses robust regressors.
result Improved variance in estimators for few-label scenarios.
C-PP-COAD detects anomalies with limited real data, reducing dependency on real calibration data.
problem Limited real calibration data for online anomaly detection.
method Context-aware prediction-powered conformal online anomaly detection (C-PP-COAD).
result Significantly reduces dependency on real calibration data without compromising FDR control.
This paper provides a guide to feature importance methods for better scientific inference.
problem Limited understanding of data-generating process due to opaque ML model mechanisms.
method Comprehensive review and new proofs of global feature importance methods.
result Facilitates a thorough understanding and concrete recommendations for FI methods.
A new variational method improves deep neural network inference.
problem Overparametrized deep neural networks struggle with variational approximations.
method A novel variational family with two independent linear subspaces.
result State-of-the-art performance across various tasks and datasets.
New analysis shows ROI's predictive power for stock returns weakens significantly.
problem The predictive power of retail order imbalance (ROI) for future stock returns.
method Replicated Boehmer et al. (2021) using a more recent period and analyzed the effect of using alternative quote midpoint (QMP) method.
result Past ROI can no longer predict weekly returns on large-cap stocks, and the long-short strategy based on past ROI is no longer profitable.