Valid p-value for bounded random variables without distributional assumptions.
problem Calibration of predictive algorithms in a distribution-free setting.
method Built a super-uniform p-value based on a concentration inequality.
result Super-uniform p-value is tighter than existing alternatives.
New method relaxes TV distance for two-sample testing without distributional assumptions.
problem Challenges in certifying equality or providing tight bounds on TV distance for two distributions.
method Examined blurred total variation distance, a relaxation of TV distance.
result Provided theoretical guarantees for upper and lower bounds on blurred TV distance.
This paper offers a distribution-free method for post-detection changepoint localization.
problem Locating the exact time of a change in distribution after a sequential detection procedure.
method A distribution-free framework using conformal test martingales for sequential change detection and post-detection inference.
result Valid post-detection coverage guarantees and non-asymptotic bounds on confidence set size.
New trade-off found in bandit problems with unknown range.
problem Stochastic bandit problems with unknown range.
method Exhibit a strategy achieving new trade-off between distribution-dependent and distribution-free regret bounds.
result Achieves the rates for regret indicated by the new trade-off.
New algorithm learns disjunctions faster than previous methods.
problem Learning Boolean disjunctions in the agnostic PAC model.
method Developed an agnostic learner with complexity 2 i l d e O ( n 1 / 3 ) 2^{ ilde{O}(n^{1/3})} 2 i l d e O ( n 1/3 ) . result First separation between SQ and CSQ models in distribution-free agnostic learning.
Research in reinforcement learning has produced algorithms for optimal decision making under uncertainty that fall within two main types. The first employs a Bayesian framework, where optimality improves with increased computational time. This is because the resulting planning task takes the form of a dynamic programmi…
Study robust linear regression without distributional assumptions for heavy-tailed responses.
problem Linear regression with heavy-tailed responses and no distributional assumptions.
method Combining truncated least squares, median-of-means, and aggregation theory to construct a non-linear estimator.
result Achieves excess risk of order d / n d/n d / n with optimal sub-exponential tail. We consider K K K -armed stochastic bandits and consider cumulative regret bounds up to time T T T . We are interested in strategies achieving simultaneously a distribution-free regret bound of optimal order K T \sqrt{KT} K T and a distribution-dependent regret that is asymptotically optimal, that is, matching the κ ln T κ\ln T κ ln T lower b…
This work examines fundamental limits in model falsification without assuming specific distributions.
problem Establishing lower bounds on model class risk in distribution-free settings.
method Model-agnostic fundamental hardness result for constructing lower bounds on test error.
result No positive lower bound on model class risk is possible in certain settings.
The paper develops a method to create non-asymptotic confidence ellipsoids for linear regression without strong noise distribution assumptions.
problem Constructing reliable confidence regions for linear regression with finite sample sizes and general noise distributions.
method The paper introduces the SPS EOA algorithm to create non-asymptotically guaranteed confidence ellipsoids for linear regression problems.
result The sizes of SPS outer ellipsoids are shown to decrease at the optimal rate for linear regression problems.
Study three types of uncertainty quantification for binary classification without distributional assumptions.
problem Uncertainty quantification for binary classification in a distribution-free setting.
method Established theorems connecting calibration, confidence intervals, and prediction sets for score-based classifiers.
result Distribution-free calibration is only possible using scoring functions that partition feature space into countably many sets.
Measuring mutual information from finite data is difficult. Recent work has considered variational methods maximizing a lower bound. In this paper, we prove that serious statistical limitations are inherent to any method of measuring mutual information. More specifically, we show that any distribution-free high-confide…
The paper develops distribution-free methods for ordinal classification.
problem Constructing valid prediction sets for ordinal classification problems.
method Leveraging conformal prediction and multiple testing with FWER control.
result The proposed methods achieve satisfactory levels of marginal and class-specific conditional coverages.
New bounds for agnostic learning with average smoothness.
problem Distribution-free nonparametric regression with average smoothness.
method Distribution-free uniform convergence bounds and agnostic learning algorithm.
result Distribution-free uniform convergence bounds for average-smoothness classes in the agnostic setting.
Locus scores predictions for risk, reducing large-loss events.
problem Deployment cost from inaccurate predictions, especially large losses.
method Distribution-free loss-scale reliability score using any predictive distribution.
result Reduces large-loss frequency compared to standard heuristics.
Characterizes distribution-free rates in unbalanced classification problems.
problem Minimizing error under two different distributions in unbalanced settings.
method Characterizes minimax rates over all pairs of distributions using a geometric condition.
result Identifies a dichotomy between hard and easy classes based on a three-points-separation condition.
Enhances scenario approach for certifying design properties post-design.
problem Certifying additional useful properties in designs not considered during the design phase.
method Two-level framework of appropriateness: baseline and post-design. Distribution-free upper bounds on risk derived.
result Distribution-free upper bounds on the risk of failing to meet post-design appropriateness.
The paper proposes a method for distribution-free prediction sets that adapt to unknown temporal changes.
problem Distribution-free prediction sets require reliable calibration data, which is often unavailable in real-world settings with temporal changes.
method The method selects an adaptive window to construct prediction sets, optimizing a bias-variance tradeoff.
result The method provides sharp coverage guarantees and is shown to be adaptive to temporal drift through numerical experiments.
The paper provides high-probability bounds on false discovery proportions in conformal inference.
problem Existing methods fail to provide high-probability bounds on the realized false discovery proportion.
method Constructing a high-probability envelope for the empirical distribution function of null conformal p-values by sampling from their joint distribution.
result Establishes finite-sample, distribution-free upper bounds on the FDP that hold simultaneously over all possible rejection thresholds.
Algorithm learns without knowing distribution, reducing error.
problem Sequential prediction with adversarial injections and abstentions.
method Boosting procedure of weak learners for general VC classes.
result Sublinear error guarantees for general VC classes.
Unified theoretical guarantees for distribution-free changepoint detection and testing.
problem Distribution-free changepoint inference with finite-sample validity and consistency.
method Distribution-free changepoint localization using conformal p-values with theoretical guarantees.
result Unified distribution-free guarantees for changepoint detection, localization, and testing.
The paper tackles uncertainty quantification for classification under label shift without assuming i.i.d. data.
problem Uncertainty quantification for classification under label shift in non-i.i.d. settings.
method The paper uses conformal prediction and post-hoc binning for distribution-free UQ, and reweights these methods for label shift.
result The reweighted methods improve UQ performance under label shift, preserving coverage and calibration.
A new pricing strategy learns customer valuations without noise distribution knowledge.
problem Setting optimal prices for products based on customer valuations with unknown noise.
method Developed a novel perturbed linear bandit framework to learn both contextual functions and market noise.
result Proved sub-linear regret bound and demonstrated superior performance on simulations and real data.
ERAPS builds prediction sets for time-series data.
problem Uncertainty quantification in complex machine learning methods for time-series data.
method ERAPS is an ensemble-based framework for constructing prediction sets for time-series data, allowing unknown dependencies within features and responses.
result ERAPS demonstrates valid marginal and conditional coverage and yields smaller prediction sets than competing methods.
The paper explores how over-parameterized linear regression models generalize without violating learning theory principles.
problem Understanding how over-parameterized linear regression models generalize without violating learning theory principles.
method The paper uses the predictive normalized maximum likelihood (pNML) learner to investigate the minimum norm solution of over-parameterized linear regression models.
result The model generalizes well when the test sample lies in a subspace spanned by eigenvectors associated with large eigenvalues of the training data.
FedFaiREE addresses fairness in decentralized learning with small samples.
problem Ensuring fairness in decentralized federated learning with limited data.
method FedFaiREE is a post-processing algorithm for distribution-free fair learning in decentralized settings with small samples.
result FedFaiREE provides theoretical guarantees for both fairness and accuracy in decentralized environments.
Novel concentration inequalities are obtained for the missing mass, i.e. the total probability mass of the outcomes not observed in the sample. We derive distribution-free deviation bounds with sublinear exponents in deviation size for missing mass and improve the results of Berend and Kontorovich (2013) and Yari Saeed…
New framework controls statistical dispersion for high-stakes applications.
problem Understanding and controlling the dispersion of loss distributions in high-stakes applications.
method Simple yet flexible framework for distribution-free control of statistical dispersion measures.
result Proposed methods control statistical dispersion measures with societal implications.
We develop an online learning method for prediction, which is important in problems with large and/or streaming data sets. We formulate the learning approach using a covariance-fitting methodology, and show that the resulting predictor has desirable computational and distribution-free properties: It is implemented onli…
CREDO assesses decision optimality under uncertainty without assuming a model.
problem Uncertainty in decision-making without reliable quantification of optimality.
method CREDO uses the inverse feasible region and conformal prediction balls to estimate decision optimality probability.
result CREDO provides accurate, efficient, and reliable evaluations of decision optimality.
Large language models can't efficiently reason conditionally in a distribution-free setting.
problem Impossibility of conditional PAC-efficient reasoning in large language models.
method Proof of impossibility in a distribution-free setting for non-atomic input spaces.
result Any algorithm achieving conditional PAC efficiency must defer to the expert model with high probability.
This paper improves model selection with cross-validation using domain knowledge.
problem Improving model selection with cross-validation risk estimation.
method Establishes distribution-free deviation bounds using VC dimension, formalizes Learning Spaces based on domain knowledge.
result Enhanced generalization through selection of candidate models based on domain knowledge.
New algorithm for precise changepoint localization without assumptions.
problem Offline changepoint localization in arbitrary distributions.
method Distribution-free algorithm CONformal CHangepoint localization (CONCH) using exchangeability arguments.
result Derives principled score functions for informative and small confidence sets with normalized length shrinking to zero.
Conformal prediction provides distribution-free uncertainty quantification for black-box models.
problem Uncertainty quantification for high-risk machine learning applications.
method Conformal prediction creates valid uncertainty sets without distributional assumptions.
result Sets contain the ground truth with a specified probability, e.g., 90%.
Improves conditional coverage of regression models using conformal prediction.
problem Lack of conditional coverage guarantees in conformal prediction methods.
method Proposes a novel algorithm to train a regression function to improve conditional coverage after split conformal prediction.
result Establishes an upper bound for miscoverage gap and proposes an end-to-end algorithm to control it.
CROC identifies the earliest-changing stream as the root cause in multi-stream data.
problem Distribution-free root cause analysis in multi-stream data with unknown distributional changes.
method Conformal p-values and finite-sample valid confidence sets.
result CROC efficiently isolates the root cause under minimal assumptions.
New tests for binary classification regression functions without distribution assumptions.
problem Testing regression functions in binary classification without distributional assumptions.
method Conditional kernel mean embeddings and resampling-based framework.
result Distribution-free hypothesis tests with exact type I error control.
We introduce a modular framework for market making. It combines cost-function based automated market makers with bandit algorithms. We obtain worst-case profits guarantee's relative to the best in hindsight within a class of natural "overround" cost functions . This combination allow us to have distribution-free guaran…
Mack's estimator improves chain ladder prediction for large exposure insurance models.
problem Uncertainty quantification in compound Poisson loss models.
method Large exposure asymptotics applied to Mack's estimator.
result Chain ladder prediction uncertainty can be quantified without model assumptions.
FaiREE provides fair classification with guarantees for small datasets.
problem Fairness in classification often requires large sample sizes and distributional assumptions.
method FaiREE offers finite-sample and distribution-free fairness guarantees.
result FaiREE achieves optimal accuracy and satisfies various fairness notions.
Conformal prediction offers distribution-free inference for complex models.
problem Traditional predictive inference methods are limited by assumptions about data distributions and model details.
method Conformal prediction uses symmetry assumptions and treats learning algorithms as black boxes.
result Conformal prediction provides exact finite-sample guarantees, even under limited assumptions.
This work presents the concept of kernel mean embedding and kernel probabilistic programming in the context of stochastic systems. We propose formulations to represent, compare, and propagate uncertainties for fairly general stochastic dynamics in a distribution-free manner. The new tools enjoy sound theory rooted in f…
Method predicts NAFLD risk with high accuracy and distribution-free coverage guarantees.
problem Insufficient population-level screening tools for NAFLD.
method Gradient-boosted decision trees with conformal prediction.
result Method achieves AUROC of 0.912 internally and 0.891 externally, superior to other models.
Paper develops a neural network method for censored survival analysis.
problem Distribution-free quantile prediction for censored survival data.
method Develops a novel neural network algorithm for simultaneous quantile optimization.
result The algorithm produces better calibrated quantiles on real datasets.
The study bounds the stability of Gaussian mixtures under small perturbations.
problem Stability of Gaussian mixtures under small changes in distribution.
method Deriving an explicit bound on parameter stability of spherical Gaussian Mixture Models (sGMM) in a pre-defined model class.
result Upper bound on parameter distance of close sGMMs to the original sGMM, dependent only on the original model.
Unified framework for fair classification with group-blindness/awareness guarantees.
problem Challenges in enforcing fairness and group-blindness in binary classification.
method Unified framework based on post-processing procedure, applicable to various group fairness notions.
result Minimax rate-optimality of the proposed algorithm with controlled excess risk.
The paper develops bounds for predictive values in binary classification.
problem Lack of confidence intervals for positive and negative predictive values.
method Bi-criterion framework and distribution-free large deviation and uniform convergence bounds.
result New bounds for predictive values without relying on concentration inequalities.
UCB algorithm adapted for large-scale, non-sub-Gaussian problems.
problem Selecting the best alternative from a large set of options with non-sub-Gaussian performance distributions.
method Adapted UCB algorithm for non-sub-Gaussian settings, focusing on sample size and meta-UCB selection.
result UCB algorithms can achieve sample optimality in large-scale, non-sub-Gaussian problems.