This paper describes Simpson's paradox, and explains its serious implications for randomised control trials. In particular, we show that for any number of variables we can simulate the result of a controlled trial which uniformly points to one conclusion (such as 'drug is effective') for every possible combination of t…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New estimator improves statistical validity of synthetic data integration.
Used to estimate the risk of an estimator or to perform model selection, cross-validation is a widespread strategy because of its simplicity and its apparent universality. Many results exist on the model selection performances of cross-validation procedures. This survey intends to relate these results to the most recen…
Valid inference from data and predictions.
Research in natural language processing proceeds, in part, by demonstrating that new models achieve superior performance (e.g., accuracy) on held-out test data, compared to previous results. In this paper, we demonstrate that test-set performance scores alone are insufficient for drawing accurate conclusions about whic…
Adaptive auditing improves AI robustness testing with anytime-valid guarantees.
Neuroimaging research has predominantly drawn conclusions based on classical statistics, including null-hypothesis testing, t-tests, and ANOVA. Throughout recent years, statistical learning methods enjoy increasing popularity, including cross-validation, pattern classification, and sparsity-inducing regression. These t…
New method improves spatial prediction validation accuracy.
Predictive models ground many state-of-the-art developments in statistical brain image analysis: decoding, MVPA, searchlight, or extraction of biomarkers. The principled approach to establish their validity and usefulness is cross-validation, testing prediction on unseen data. Here, I would like to raise awareness on e…
We give a necessary and sufficient geometric structural condition for a stable codimension 1 integral varifold on a smooth Riemannian manifold to correspond to an embedded smooth hypersurface away from a small set of generally unavoidable singularities; when this condition is satisfied, the singular set is empty if the…
The paper evaluates index-based allocation policies using data from randomized control trials.
Machine learning methods may have the potential to significantly accelerate drug discovery. However, the increasing rate of new methodological approaches being published in the literature raises the fundamental question of how models should be benchmarked and validated. We reanalyze the data generated by a recently pub…
A well-known result asserts that any isometric immersion with flat normal bundle of a Riemannian manifold with constant sectional curvature into a space form is (at least locally) holonomic. In this note, we show that this conclusion remains valid for the larger class of Einstein manifolds. As an application, when assu…
New method uses predictions to infer causal effects without labeled data.
A new model validation framework for agentic AI systems based on POMDPs.
FreB protocol uses AI to infer hidden parameters with valid confidence regions.
We propose a stepsize adaptation scheme for stochastic gradient descent. It operates directly with the loss function and rescales the gradient in order to make fixed predicted progress on the loss. We demonstrate its capabilities by conclusively improving the performance of Adam and Momentum optimizers. The enhanced op…
New method for valid inference from ML-predicted data.
A new method corrects for bias in selecting the best candidate.
The validity of the Efficient Market Hypothesis has been under severe scrutiny since several decades. However, the evidence against it is not conclusive. Artificial Neural Networks provide a model-free means to analize the prediction power of past returns on current returns. This chapter analizes the predictability in …
A new method uses randomized trials to estimate the strength of unobserved confounding.
STAND-DA improves AD in DA target domains with limited data.
bioLeak addresses data leakage in biomedical machine learning studies.
Paper presents a workflow for reliable unsupervised learning in science.
This paper focuses on the stability of the non-arbitrage condition in discrete time market models when some unknown information is partially/fully incorporated into the market. Our main conclusions are twofold. On the one hand, for a fixed market , we prove that the non-arbitrage condition is preserved under a m…
Over the past two decades, several consistent procedures have been designed to infer causal conclusions from observational data. We prove that if the true causal network might be an arbitrary, linear Gaussian network or a discrete Bayes network, then every unambiguous causal conclusion produced by a consistent method f…
Study improves predictive performance testing for high-dimensional data using exhaustive nested cross-validation.
Choosing a reference group in Oaxaca-Blinder decomposition can reverse conclusions.
Deep learning, computational neuroscience, and cognitive science have overlapping goals related to understanding intelligence such that perception and behaviour can be simulated in computational systems. In neuroimaging, machine learning methods have been used to test computational models of sensory information process…
Study finds unsupervised imputation before cross-validation can reduce computational costs without significantly degrading model performance.
Novel strategy benchmarks observational studies against randomized trials.
This paper concerns the development of an inferential framework for high-dimensional linear mixed effect models. These are suitable models, for instance, when we have repeated measurements for subjects. We consider a scenario where the number of fixed effects is large (and may be larger than ), but the n…
Benchopt automates machine learning benchmarking across languages and hardware.
Generalization of deep networks has been of great interest in recent years, resulting in a number of theoretically and empirically motivated complexity measures. However, most papers proposing such measures study only a small set of models, leaving open the question of whether the conclusion drawn from those experiment…
Although deep learning models have proven effective at solving problems in natural language processing, the mechanism by which they come to their conclusions is often unclear. As a result, these models are generally treated as black boxes, yielding no insight of the underlying learned patterns. In this paper we conside…
Deep learning detects sleep state fluctuations in neonates from single EEG channel.
Study explores efficient data division for ICPs.
Efficient CV for ESNs improves time series predictions.
The paper proposes a learning algorithm that improves adaptability and generalization.
Rigidity theorem for critical points of Allen-Cahn equation on S³.
BayesFlow trains neural networks for fast Bayesian inference.
Background: Parkinson's disease (PD) is a prevalent long-term neurodegenerative disease. Though the diagnostic criteria of PD are relatively well defined, the current medical imaging diagnostic procedures are expertise-demanding, and thus call for a higher-integrated AI-based diagnostic algorithm. Methods: In this pape…
Traditional statistical theory assumes that the analysis to be performed on a given data set is selected independently of the data themselves. This assumption breaks downs when data are re-used across analyses and the analysis to be performed at a given stage depends on the results of earlier stages. Such dependency ca…
pmsims R package uses Gaussian process for flexible sample size estimation in clinical models.
AI tool automates blood segmentation from head CT scans after SAH.
Analyzes adversarial training's impact on loss landscape, proposing PAS to improve model performance.
CVTMLE improves statistical inference in settings of positivity or Donsker class violations.
Objective: To evaluate unsupervised clustering methods for identifying individual-level behavioral-clinical phenotypes that relate personal biomarkers and behavioral traits in type 2 diabetes (T2DM) self-monitoring data. Materials and Methods: We used hierarchical clustering (HC) to identify groups of meals with simila…