This paper describes Simpson's paradox, and explains its serious implications for randomised control trials. In particular, we show that for any number of variables we can simulate the result of a controlled trial which uniformly points to one conclusion (such as 'drug is effective') for every possible combination of t…
New estimator improves statistical validity of synthetic data integration.
problem Combining synthetic data generated by large language models with real data for valid inference.
method Generalized method of moments estimator with theoretical guarantees.
result Improves estimates of target parameter through interactions between synthetic and real data.
Used to estimate the risk of an estimator or to perform model selection, cross-validation is a widespread strategy because of its simplicity and its apparent universality. Many results exist on the model selection performances of cross-validation procedures. This survey intends to relate these results to the most recen…
Valid inference from data and predictions.
problem Valid statistical inference with machine learning predictions.
method Framework for valid inference using machine learning predictions.
result Valid confidence intervals without assumptions on predictions.
Research in natural language processing proceeds, in part, by demonstrating that new models achieve superior performance (e.g., accuracy) on held-out test data, compared to previous results. In this paper, we demonstrate that test-set performance scores alone are insufficient for drawing accurate conclusions about whic…
Adaptive auditing improves AI robustness testing with anytime-valid guarantees.
problem Cost and time of annotation limit rigorous AI failure mode characterization.
method Introduces hypothesis testing framework for adaptive audits using SAVI.
result Proves anytime-valid type-I error control and robustness certification.
Neuroimaging research has predominantly drawn conclusions based on classical statistics, including null-hypothesis testing, t-tests, and ANOVA. Throughout recent years, statistical learning methods enjoy increasing popularity, including cross-validation, pattern classification, and sparsity-inducing regression. These t…
New method improves spatial prediction validation accuracy.
problem Validation methods fail for spatial prediction tasks due to mismatch between validation and test locations.
method Proposes a new validation method that adapts existing covariate-shift ideas to spatial settings.
result Proves and demonstrates the new method's superiority in spatial prediction validation.
Predictive models ground many state-of-the-art developments in statistical brain image analysis: decoding, MVPA, searchlight, or extraction of biomarkers. The principled approach to establish their validity and usefulness is cross-validation, testing prediction on unseen data. Here, I would like to raise awareness on e…
We give a necessary and sufficient geometric structural condition for a stable codimension 1 integral varifold on a smooth Riemannian manifold to correspond to an embedded smooth hypersurface away from a small set of generally unavoidable singularities; when this condition is satisfied, the singular set is empty if the…
The paper evaluates index-based allocation policies using data from randomized control trials.
problem Evaluating index-based allocation policies in resource-scarce scenarios.
method Using data from randomized control trials, the paper introduces an efficient estimator and methods for computing asymptotically correct confidence intervals.
result Valid statistical conclusions can be drawn for index-based allocation policies.
Machine learning methods may have the potential to significantly accelerate drug discovery. However, the increasing rate of new methodological approaches being published in the literature raises the fundamental question of how models should be benchmarked and validated. We reanalyze the data generated by a recently pub…
A well-known result asserts that any isometric immersion with flat normal bundle of a Riemannian manifold with constant sectional curvature into a space form is (at least locally) holonomic. In this note, we show that this conclusion remains valid for the larger class of Einstein manifolds. As an application, when assu…
New method uses predictions to infer causal effects without labeled data.
problem Data labeling costs limit causal inference experiments.
method Prediction-Powered Causal Inferences (PPCI) using conditional calibration and transfer constraints.
result Valid causal inference achieved on experiments with no human annotations.
A new model validation framework for agentic AI systems based on POMDPs.
problem Model validation of agentic AI systems.
method A POMDP-based framework for belief-state, forecast, and policy validation.
result The framework decomposes autonomous decision making into information, beliefs, forecasts, actions, and utility.
FreB protocol uses AI to infer hidden parameters with valid confidence regions.
problem Generating biased or overconfident conclusions from AI-generated posterior distributions.
method Frequentist-Bayes (FreB) protocol reshapes AI-generated posterior distributions into valid confidence regions.
result FreB provides valid confidence regions that consistently include true parameters with expected probability.
We propose a stepsize adaptation scheme for stochastic gradient descent. It operates directly with the loss function and rescales the gradient in order to make fixed predicted progress on the loss. We demonstrate its capabilities by conclusively improving the performance of Adam and Momentum optimizers. The enhanced op…
New method for valid inference from ML-predicted data.
problem Invalid scientific conclusions from ML-predicted outcomes.
method Assumption-Lean and Data-Adaptive Post-Prediction Inference (PSPA)
result Valid and powerful inference based on ML-predicted data.
A new method corrects for bias in selecting the best candidate.
problem Bias in selecting the best candidate leads to invalid conclusions.
method Zoom correction, flexible for parametric and nonparametric settings.
result Valid inference on the winner is possible, even with selection bias.
The validity of the Efficient Market Hypothesis has been under severe scrutiny since several decades. However, the evidence against it is not conclusive. Artificial Neural Networks provide a model-free means to analize the prediction power of past returns on current returns. This chapter analizes the predictability in …
A new method uses randomized trials to estimate the strength of unobserved confounding.
problem Unobserved confounding compromises causal conclusions from non-randomized studies.
method Designs a statistical test to detect unobserved confounding strength and estimates a lower bound.
result Estimates an asymptotically valid lower bound on unobserved confounding strength.
STAND-DA improves AD in DA target domains with limited data.
problem Statistical validity of AD after DA with limited data.
method Selective Inference framework for GPU-accelerated p-value computation. result Valid p-values and controlled false positive rate. bioLeak addresses data leakage in biomedical machine learning studies.
problem Data leakage causes optimistic bias in machine learning models for biomedical studies.
method bioLeak provides leakage-aware resampling workflows and model audits in R.
result The package supports various machine learning tasks and can detect leakage mechanisms.
Paper presents a workflow for reliable unsupervised learning in science.
problem Lack of standardization in unsupervised learning workflows for reproducible scientific discoveries.
method Structured workflow including data preparation, modeling, validation, and communication.
result Illustrates the importance of validation in unsupervised learning.
This paper focuses on the stability of the non-arbitrage condition in discrete time market models when some unknown information τ is partially/fully incorporated into the market. Our main conclusions are twofold. On the one hand, for a fixed market S, we prove that the non-arbitrage condition is preserved under a m…
Over the past two decades, several consistent procedures have been designed to infer causal conclusions from observational data. We prove that if the true causal network might be an arbitrary, linear Gaussian network or a discrete Bayes network, then every unambiguous causal conclusion produced by a consistent method f…
Study improves predictive performance testing for high-dimensional data using exhaustive nested cross-validation.
problem Reproducibility issues in K-fold cross-validation for high-dimensional data. method Proposes a novel predictive performance test based on exhaustive nested cross-validation, addressing computational complexity with a closed-form expression.
result Demonstrates the effectiveness of Ridge-based methods in high-dimensional predictive performance testing.
Choosing a reference group in Oaxaca-Blinder decomposition can reverse conclusions.
problem The choice of reference group in Oaxaca-Blinder decomposition can lead to different conclusions.
method The study uses the Oaxaca-Blinder decomposition to investigate how the choice of reference group affects the results.
result The Oaxaca-Blinder decomposition can yield different conclusions based on the choice of reference group.
Deep learning, computational neuroscience, and cognitive science have overlapping goals related to understanding intelligence such that perception and behaviour can be simulated in computational systems. In neuroimaging, machine learning methods have been used to test computational models of sensory information process…
Study finds unsupervised imputation before cross-validation can reduce computational costs without significantly degrading model performance.
problem High computational costs in pipeline modeling algorithms with imputation steps.
method Empirical assessment of unsupervised imputation before vs during cross-validation.
result Reduced variance of imputation before cross-validation leads to lower overall root mean squared error.
Novel strategy benchmarks observational studies against randomized trials.
problem Benchmarking observational studies for treatment effect bias.
method Statistical test for null hypothesis of treatment effect difference.
result Valid lower bound on maximum bias strength for any subgroup.
Develops simple approximations for credit risk estimation.
problem Estimating conditional default probabilities for corporate loans.
method Closed form approximations to maximum likelihood estimator for logit type models.
result The approximations are as good as the MLE under certain conditions.
Benchopt automates machine learning benchmarking across languages and hardware.
problem Limited transparency and tedious re-implementation work in machine learning validation.
method A collaborative framework for automating, reproducing, and publishing optimization benchmarks.
result Demonstrates practical findings that highlight the importance of details in machine learning validation.
Generalization of deep networks has been of great interest in recent years, resulting in a number of theoretically and empirically motivated complexity measures. However, most papers proposing such measures study only a small set of models, leaving open the question of whether the conclusion drawn from those experiment…
Although deep learning models have proven effective at solving problems in natural language processing, the mechanism by which they come to their conclusions is often unclear. As a result, these models are generally treated as black boxes, yielding no insight of the underlying learned patterns. In this paper we conside…
Deep learning detects sleep state fluctuations in neonates from single EEG channel.
problem Monitoring sleep state fluctuations in neonatal intensive care units.
method Deep learning-based algorithm trained on 53 EEG recordings, validated on 30 polysomnography recordings.
result High accuracy (90%) in detecting quiet sleep states from single EEG channel, generalizing well to external dataset.
Study explores efficient data division for ICPs.
problem Efficiently dividing limited development data for ICPs.
method Experiments with training, calibration, and test data divisions.
result Allows overlap between training and calibration sets improves efficiency.
Efficient CV for ESNs improves time series predictions.
problem Lack of CV in time series modeling, especially for ESNs.
method Two-level optimizations for k-fold CV of ESNs. result Proposed CV schemes give better and more stable test performance.
The paper proposes a learning algorithm that improves adaptability and generalization.
problem Improving adaptability and generalization in learning models.
method Learning to meta-learn by meta-finetuning on related tasks before adapting to specific tasks.
result Learning to meta-learn improves adaptability and generalization across various tasks.
Develops methods for statistical inference in high-dimensional linear mixed models.
problem Statistical inference for high-dimensional linear mixed models with large fixed effects and small random effects.
method Inspired by de-biasing penalized estimators, corrects a `naive' ridge estimator to build asymptotically valid confidence intervals.
result Demonstrates that the proposed method outperforms those that ignore correlation induced by random effects.
Rigidity theorem for critical points of Allen-Cahn equation on S³.
problem Rigidity of critical points with low Morse index on S³.
method Analysis of nullity and symmetries of critical points, Frankel-type theorem for nodal sets.
result Critical points with index five are symmetric and vanish on a Clifford torus, realizing the fifth width of the min-max spectrum.
BayesFlow trains neural networks for fast Bayesian inference.
problem Fast Bayesian inference for complex models.
method Amortized neural networks for intractable posterior distributions.
result Fast inference through pre-trained neural networks.
Background: Parkinson's disease (PD) is a prevalent long-term neurodegenerative disease. Though the diagnostic criteria of PD are relatively well defined, the current medical imaging diagnostic procedures are expertise-demanding, and thus call for a higher-integrated AI-based diagnostic algorithm. Methods: In this pape…
Traditional statistical theory assumes that the analysis to be performed on a given data set is selected independently of the data themselves. This assumption breaks downs when data are re-used across analyses and the analysis to be performed at a given stage depends on the results of earlier stages. Such dependency ca…
pmsims R package uses Gaussian process for flexible sample size estimation in clinical models.
problem Determining adequate sample size for clinical prediction models.
method Simulation-based Gaussian process search for flexible sample size estimation.
result Gaussian process-based method produces more stable sample size estimates, especially in challenging settings.
AI tool automates blood segmentation from head CT scans after SAH.
problem Accurate volumetric assessment of SAH patients for clinical and prognostic implications.
method Transformer-based Swin UNETR architecture for noncontrast CT scans.
result High accuracy and robust performance across internal and external validation cohorts.
Analyzes adversarial training's impact on loss landscape, proposing PAS to improve model performance.
problem Challenges in optimizing models under adversarial training due to loss landscape properties.
method Analytical studies of adversarial loss functions, numerical analyses, PAS strategy.
result Adversarial training impairs optimization, but PAS strategy improves model performance.
CVTMLE improves statistical inference in settings of positivity or Donsker class violations.
problem Inference issues in causal inference due to data sparsity or near-positivity violations.
method Cross-validation of TMLE (CVTMLE) to improve performance in settings of positivity or Donsker class violations.
result CVTMLE vastly improves confidence interval coverage without affecting bias, especially in small sample sizes and near-positivity violations.