The network jackknife provides conservative variance estimates for network statistics.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Jackknife variance estimation validated for generalized U-statistics.
Extends Infinitesimal Jackknife for model covariance, enhancing ensemble model analysis.
We study the variability of predictions made by bagged learners and random forests, and show how to estimate standard errors for these methods. Our work builds on variance estimates for bagging proposed by Efron (1992, 2012) that are based on the jackknife and the infinitesimal jackknife (IJ). In practice, bagged predi…
Paper proposes diagnostics for error and variance estimation in randomized matrix computations.
We provide additional statistical background for the methodology developed in the clinical analysis of knee osteoarthritis in "A Precision Medicine Approach to Develop and Internally Validate Optimal Exercise and Weight Loss Treatments for Overweight and Obese Adults with Knee Osteoarthritis" (Jiang et al. 2020). Jiang…
The infinitesimal jackknife (IJ) has recently been applied to the random forest to estimate its prediction variance. These theorems were verified under a traditional random forest framework which uses classification and regression trees (CART) and bootstrap resampling. However, random forests using conditional inferenc…
Cluster jackknife improves inference for staggered DID methods.
Discriminative jackknife estimates deep learning uncertainty.
This article proposes a generalisation of the delete- jackknife to solve hyperparameter selection problems for time series. I call it artificial delete- jackknife to stress that this approach substitutes the classic removal step with a fictitious deletion, wherein observed datapoints are replaced with artificial …
Modified jackknife method improves predictive inference for time series data.
Prediction intervals in supervised Machine Learning bound the region where the true outputs of new samples may fall. They are necessary in the task of separating reliable predictions of a trained model from near random guesses, minimizing the rate of False Positives, and other problem-specific tasks in applied Machine …
Proposes a method to make statistical inferences robust in spatially dependent settings with missing at random labels.
Enhances polynomial chaos models with uncertainty intervals.
Cross validation (CV) and the bootstrap are ubiquitous model-agnostic tools for assessing the error or variability of machine learning and statistical estimators. However, these methods require repeatedly re-fitting the model with different weighted versions of the original dataset, which can be prohibitively time-cons…
In this paper we propose using the principle of boosting to reduce the bias of a random forest prediction in the regression setting. From the original random forest fit we extract the residuals and then fit another random forest to these residuals. We call the sum of these two random forests a \textit{one-step boosted …
Adaptive method improves prediction intervals with global coverage guarantees and local error distribution.
The paper proposes a method to improve fairness in machine learning models without refitting.
Frequentist method estimates uncertainty in RNNs without altering architecture.
JAWS audits predictive uncertainty under covariate shift using jackknife+ weighted methods.
We consider the performance of the bootstrap in high-dimensions for the setting of linear regression, where but is not close to zero. We consider ordinary least-squares as well as robust regression methods and adopt a minimalist performance requirement: can the bootstrap give us good confidence intervals fo…
Random forests have proven to be reliable predictive algorithms in many application areas. Not much is known, however, about the statistical properties of random forests. Several authors have established conditions under which their predictions are consistent, but these results do not provide practical estimates of ran…
New algorithm uses control variates to improve multi-armed bandit performance.
Adaptive classification methods ensure correct prediction intervals.
Paper proposes CIV estimator for categorical instruments in small sample settings.
New methods improve prediction intervals across multiple environments.
High-dimensional regression models struggle with resampling methods.
The paper derives uniform stability-based coverage bounds for conformal prediction methods.
Paper accelerates conformal prediction by using approximate leave-one-out estimators.
Two methods improve Gaussian process predictive distributions' calibration.
Random forests are stable and provide reliable prediction intervals.
coverforest speeds up conformal predictions for random forests.
Deep latent variable models have become a popular model choice due to the scalable learning algorithms introduced by (Kingma & Welling, 2013; Rezende et al., 2014). These approaches maximize a variational lower bound on the intractable log likelihood of the observed data. Burda et al. (2015) introduced a multi-sample v…
We use statistical learning methods to construct an adaptive state estimator for nonlinear stochastic systems. Optimal state estimation, in the form of a Kalman filter, requires knowledge of the system's process and measurement uncertainty. We propose that these uncertainties can be estimated from (conditioned on) past…
MAGIC method optimally estimates model predictions changes.
Study evaluates posterior covariance matrix W for frequentist evaluation of Bayesian estimators.
Develops asymptotic theory for deep Cox models to enable valid inference.
A hybrid algorithm fuses significance-based splitting with honest sample-splitting for estimating heterogeneous treatment effects.
A new method reduces bootstrap simulation cost and improves accuracy.
The weighted nearest neighbors (WNN) estimator has been popularly used as a flexible and easy-to-implement nonparametric tool for mean regression estimation. The bagging technique is an elegant way to form WNN estimators with weights automatically generated to the nearest neighbors; we name the resulting estimator as t…
Assessing heterogeneous treatment effects has become a growing interest in advancing precision medicine. Individualized treatment effects (ITE) play a critical role in such an endeavor. Concerning experimental data collected from randomized trials, we put forward a method, termed random forests of interaction trees (RF…
There has been increasing interest in modelling survival data using deep learning methods in medical research. Current approaches have focused on designing special cost functions to handle censored survival data. We propose a very different method with two steps. In the first step, we transform each subject's survival …
Paper proposes a new method to improve BART model predictions outside training data range.
There has been an increasing interest in testing the equality of large Pearson's correlation matrices. However, in many applications it is more important to test the equality of large rank-based correlation matrices since they are more robust to outliers and nonlinearity. Unlike the Pearson's case, testing the equality…
Conformal prediction is a popular tool for providing valid prediction sets for classification and regression problems, without relying on any distributional assumptions on the data. While the traditional description of conformal prediction starts with a nonconformity score, we provide an alternate (but equivalent) view…
Rectified decision trees improve machine learning interpretability and effectiveness.
New algorithms delete user data from machine learning models efficiently.
A new method controls risk for set predictors using cross-validation.