New methods for CI testing under model misspecification.
problem Challenges in CI testing with misspecified models.
method Proposes new approximations and upper bounds for testing errors of regression-based CI tests.
result Introduces the Rao-Blackwellized Predictor Test (RBPT) robust against misspecified inductive biases.
DSM on manifolds removes singularities and computes small-noise expansions.
problem DSM on manifolds with singular noise.
method Rao-Blackwellized score matching, nearest-point projection, intrinsic Riemannian score.
result Canonical target equals intrinsic Riemannian score up to a small correction.
Paper improves Gumbel-Softmax estimator variance reduction.
problem Challenges in gradient estimation for models with discrete latent variables.
method Rao-Blackwellization applied to straight-through Gumbel-Softmax estimator.
result Reduces mean squared error and variance of Gumbel-Softmax estimator.
Derives new equations for stochastic volatility models.
problem Modeling local-stochastic-volatility models and their derivatives.
method Conditional forward equation, Dupire stochastic PDE, rolling expiry vanilla option SPDE.
result New equations for LSV models and their derivatives.
We wish to compute the gradient of an expectation over a finite or countably infinite sample space having K≤∞ categories. When K is indeed infinite, or finite but very large, the relevant summation is intractable. Accordingly, various stochastic gradient estimators have been proposed. In this paper, we de…
Proposes a method to stabilize Black Box Variational Inference using the James-Stein estimator.
problem Stability issues and fine-tuning required in basic Black Box Variational Inference.
method Reframe stochastic gradient ascent as multivariate estimation problem using James-Stein estimator.
result Provides a simpler method with consistent performance in terms of model fit and convergence time.
A new variational method for SSMs improves inference efficiency.
problem Hard variational inference for state space models.
method Proposes variational marginal particle filter (VMPF) based on Rao-Blackwellization.
result VMPF provides tighter variational bounds and sometimes benefits from unbiased reparameterization.
The article compares predictor importance in classification problems with categorical outcomes.
problem Comparing predictor importance in classification problems with categorical response variables.
method The approach is based on the categorical Gini correlation (CGC) and tests differences in CGCs across predictor groups.
result The proposed methodology accommodates predictors of arbitrary and unequal dimensions and allows for dependence between predictor groups.
Adversarial training achieves optimal test error for shallow networks.
problem Achieving optimal adversarial test error for general data distributions.
method Applying new Rademacher complexity bounds and properties of optimal adversarial predictors.
result Adversarial training can achieve optimal adversarial test error for general data distributions.
New sampler reduces MCMC complexity for Bayesian variable selection.
problem High-dimensional Bayesian variable selection with high computation complexity.
method Variable-complexity subset weighted-Tempered Gibbs Sampler (wTGS) with Rao-Blackwellized estimator.
result Variances of Rao-Blackwellized estimator are smaller than those of subset wTGS.
New estimator reduces variance in discrete random variables.
problem Estimating gradients for discrete random variables with reduced variance.
method Sampling without replacement and Rao-Blackwellization.
result Our estimator is the most consistent gradient estimator across different entropy settings.
Improves survey sampling with unbiased machine learning methods.
problem Design-consistent model-assisted estimation lacks a general theory for machine learning.
method Proposes a subsampling Rao-Blackwell method for design-unbiased estimation.
result Yields efficiency gains over standard methods while ensuring valid estimation.
Faced with distribution shift between training and test set, we wish to detect and quantify the shift, and to correct our classifiers without test set labels. Motivated by medical diagnosis, where diseases (targets) cause symptoms (observations), we focus on label shift, where the label marginal p(y) changes but the …
Fibonacci Ensembles use Fibonacci weights to improve ensemble learning, inspired by natural growth patterns.
problem Improving ensemble learning methods to enhance model performance and interpretability.
method Introduces Fibonacci weights and a recursive ensemble dynamic to reduce variance and enrich representational depth.
result Fibonacci weighting can match or improve upon uniform averaging in ensemble learning experiments.
Partition functions of probability distributions are important quantities for model evaluation and comparisons. We present a new method to compute partition functions of complex and multimodal distributions. Such distributions are often sampled using simulated tempering, which augments the target space with an auxiliar…
The study finds a trade-off between model size, test loss, and training loss for linear predictors.
problem Finding the optimal balance between model size, test loss, and training loss for linear predictors.
method Established an algorithm and distribution-independent trade-off using non-asymptotic analysis.
result Models with low test loss are either classical (close to noise level training loss) or modern (large number of parameters).
RF models implicitly regularize kernel methods as feature count increases.
problem Understanding implicit regularization in RF models.
method Random matrix theory applied to Gaussian RF models and KRR.
result The average RF predictor is close to a KRR predictor with an effective ridge.
Proposes a new probabilistic framework for domain generalization.
problem Learning predictors robust to unseen domain shifts.
method Quantile Risk Minimization (QRM) and Empirical QRM (EQRM) algorithms.
result Empirical QRM outperforms state-of-the-art baselines on various datasets.
This paper continues study, both theoretical and empirical, of the method of Venn prediction, concentrating on binary prediction problems. Venn predictors produce probability-type predictions for the labels of test objects which are guaranteed to be well calibrated under the standard assumption that the observations ar…
New method calibrates uncertainty estimates for image classifiers without labeled data.
problem Uncertainty estimates for modern classifiers are unreliable without labeled calibration data.
method Calibrates uncertainty estimates using unlabeled examples for distribution shifts.
result Proposes a method that provides excellent uncertainty estimates under natural distribution shifts.
New gradient estimators for discrete variables improve model training.
problem Training models with discrete latent variables is challenging due to high gradient variance.
method Introduced novel gradient estimators based on importance sampling and statistical couplings, extending to categorical variables.
result Proposed gradient estimators outperform previous methods in systematic experiments.
Method constructs confidence regions for linear models with arbitrary predictors.
problem Constructing confidence regions for linear models with non-linear predictors.
method Mixed Integer Linear Programming for constraints.
result Empty confidence regions for hypothesis testing.
Causal predictors don't generalize better across domains than non-causal predictors.
problem How well do causal predictors generalize across different domains?
method 16 prediction tasks on tabular datasets, selecting causal features.
result Causal predictors do not outperform non-causal predictors in domain generalization.
New bounds show linear predictors rarely overfit with certain optimization methods.
problem Bounding test error for linear predictors with stochastic optimization methods.
method Coupling argument for fixed point methods like stochastic and batch mirror descent.
result Locally-adapted rates that depend on predictor properties, not global problem structure.
Paper revisits pre-validation method, improving hypothesis testing.
problem Improving hypothesis testing in pre-validated models with different feature dimensions.
method Extended problem formulation, analytical distribution, and bootstrap procedure.
result Proposed analytical distribution and bootstrap procedure for pre-validated predictors.
Autoencoder performance is predicted by eigenvalues of weight matrices.
problem Predicting an autoencoder's generalization ability without dataset knowledge.
method Analyze Jacobian matrices' eigenvalues to bound mean squared errors.
result Eigenvalues are good predictors of MSE on test points.
LESS combines local predictors for subsets to learn from heterogeneous input-output pairs.
problem Learning from heterogeneous input-output pairs in populations with varied behavior.
method LESS algorithm: generates subsets, trains local predictors, combines them.
result LESS is highly competitive compared to state-of-the-art methods.
This paper presents Sparse Partitioning, a Bayesian method for identifying predictors that either individually or in combination with others affect a response variable. The method is designed for regression problems involving binary or tertiary predictors and allows the number of predictors to exceed the size of the sa…
We introduce a dynamic mechanism for the solution of analytically-tractable substructure in probabilistic programs, using conjugate priors and affine transformations to reduce variance in Monte Carlo estimators. For inference with Sequential Monte Carlo, this automatically yields improvements such as locally-optimal pr…
New test for SGD in binary classification reduces computation time.
problem Determining optimal stopping for SGD in binary classification.
method Proposes a new, simple, computationally inexpensive termination criterion for SGD.
result Termination criterion reduces expected misclassification probability.
Blackwell's theorems influence modern AI through information compression and decision making.
problem Information compression and decision making under uncertainty.
method Theorems developed in the 1940s and 1950s, applied to modern AI.
result Blackwell theorems remain relevant and influence modern AI subfields.
New theory validates the use of invariant predictors for OOD generalization.
problem Ensuring predictors generalize well across unseen environments.
method Developed new theoretical conditions and derived an Inter Gradient Alignment algorithm.
result Validated the necessity of invariant predictors for OOD optimality.
Efficiency criteria improve conformal predictors' performance.
problem Improving the performance of conformal predictors.
method Learning classifiers by minimizing observed fuzziness as a training objective function.
result Conformal predictors trained by minimizing observed fuzziness perform better than traditional ones.
Adaptive kernels from neural networks improve model performance.
problem Improving neural network performance through adaptive kernels.
method Deriving adaptive kernels from infinite-width neural networks using feature learning and gradient flow training.
result Adaptive kernels achieve lower test loss compared to traditional kernels.
This paper tackles continuous covariate shift by adaptively training predictors.
problem Continuous covariate shift where input distributions change over time.
method Online density ratio estimation method to adaptively train predictors.
result Excess risk guarantee for the predictor through dynamic regret bound.
While the use of deep learning in drug discovery is gaining increasing attention, the lack of methods to compute reliable errors in prediction for Neural Networks prevents their application to guide decision making in domains where identifying unreliable predictions is essential, e.g. precision medicine. Here, we prese…
Derives new equations for volatility models and option pricing.
problem Modeling and pricing options in local-stochastic-volatility models.
method Develops conditional forward equations and Dupire stochastic PDEs.
result Derives new SPDE for vanilla options.
Proposes BSSP to stabilize predictions in biased data.
problem Distribution shift between training and test data causes prediction instability.
method Balance-subsampled stable prediction (BSSP) algorithm based on fractional factorial design.
result Significantly improves prediction stability across unknown test data.
We resolve the open problem of optimal sample complexity for multicalibration and deterministic predictors.
problem Optimal sample complexity for multicalibration and deterministic predictors
method Minimax-optimal multicalibration algorithm and generalization to OI predictors
result Minimax-optimal multicalibration algorithm and deterministic predictors with optimal sample complexity
Paper proposes a method to identify key variables in thick data.
problem Selecting significant variables for response prediction.
method Permutation tests on candidate variables for feature selection.
result The approach outperforms Lasso in feature selection.
New analysis identifies key factors in wildfire-generated thunderstorms.
problem Understanding the causes of pyrocumulonimbus (pyroCb) storms.
method Invariant Causal Prediction, conditional independence test, greedy-ICP search algorithm.
result Identified seven causal predictors for pyroCb formation.
Study evaluates 31 performance predictors in NAS, recommending best for different settings.
problem Understanding and comparing different performance prediction techniques in NAS.
method Analysis of 31 techniques, testing correlation and rank-based measures, speed-up potential.
result Certain predictor families can be combined for better predictive power.
Optimized testing of discrete distributions using predicted data.
problem Testing discrete distributions with reduced sample complexity.
method Adaptable algorithms that use a predicted distribution to reduce sample size.
result Optimal sample complexity improvements with self-adjusting algorithms.
In this article, we derive concentration inequalities for the cross-validation estimate of the generalization error for stable predictors in the context of risk assessment. The notion of stability has been first introduced by \cite{DEWA79} and extended by \cite{KEA95}, \cite{BE01} and \cite{KUNIY02} to characterize cla…
Study explores efficient data division for ICPs.
problem Efficiently dividing limited development data for ICPs.
method Experiments with training, calibration, and test data divisions.
result Allows overlap between training and calibration sets improves efficiency.
SFB uses stable features to adapt unstable ones for better performance.
problem Improving classifier performance on out-of-distribution data by leveraging stable features.
method SFB learns a predictor that separates stable and unstable features, then adapts unstable predictions using stable predictions.
result SFB can learn an asymptotically-optimal predictor without test-domain labels.
A novel feature selection method using noise-based hypothesis testing improves feature selection accuracy.
problem Challenges in feature selection for complex, high-dimensional datasets.
method Introduces multiple random noise features and evaluates feature importance against noise feature maxima using non-parametric bootstrap-based hypothesis testing.
result Outperforms existing methods in simulated and real-world datasets.
DetectShift framework detects and quantifies dataset shifts in various data types.
problem Frequent dataset shifts decrease supervised learning performance.
method DetectShift framework quantifies and tests for multiple dataset shifts in various data types.
result DetectShift framework effectively detects dataset shifts even in higher dimensions.