Algorithm detects unmeasured confounding in observational data.
problem Estimating treatment effects in observational studies with untestable conditions.
method Two-stage procedure that detects dependencies between causal mechanisms.
result Algorithm efficiently detects confounding on simulated and semi-synthetic data.
VerifAI toolkit improves neural network-based aircraft taxiing system safety.
problem Improving safety of autonomous aircraft taxiing systems using neural networks.
method Unified approach to formal analysis and retraining of AI systems, including falsification, debugging, and retraining.
result Improved neural network performance and reduced failure cases in aircraft taxiing system.
Proposes a falsification framework to test algorithmic discriminant validity.
problem Unintended model behavior in predictive algorithms.
method Falsification framework based on statistical tests comparing prediction losses across outcomes.
result Establishes discriminant validity for some outcomes but not others.
Audit financial machine learning workflows to detect spurious predictability.
problem Spurious predictability in financial machine learning models.
method Falsification audit testing predictive workflows against synthetic environments.
result Many apparent financial predictions are artifacts, not genuine.
Study finds no statistically significant trading edge in MNQ futures signals from OHLCV data.
problem Testing intraday momentum signals from OHLCV data in MNQ futures under realistic execution constraints.
method 947 trading days of five-minute data, 14 signal families evaluated, strict institutional criteria applied.
result No signal satisfies all criteria simultaneously, gross edge insufficient to overcome costs.
Proposes a test to ensure predictive algorithms predict intended outcomes better than unintended ones.
problem Unintended model behavior leading to prediction of unintended outcomes.
method Falsification framework using nonparametric hypothesis testing to compare prediction losses across outcomes.
result Establishes discriminant validity with respect to gender but not race in an admissions setting.
This work examines fundamental limits in model falsification without assuming specific distributions.
problem Establishing lower bounds on model class risk in distribution-free settings.
method Model-agnostic fundamental hardness result for constructing lower bounds on test error.
result No positive lower bound on model class risk is possible in certain settings.
This paper considers the problem of detection in distributed networks in the presence of data falsification (Byzantine) attacks. Detection approaches considered in the paper are based on fully distributed consensus algorithms, where all of the nodes exchange information only with their neighbors in the absence of a fus…
A new method predicts causal relationships without joint data.
problem Falsifying causal discovery algorithms without ground truth.
method Leave-One-Variable-Out (LOVO) prediction for causal graphs.
result LOVO method correlates prediction error with causal algorithm accuracy.
Tests validity of DML estimators without assumptions.
problem Validating DML estimators without making assumptions.
method Develops tests to falsify assumptions for DML estimators.
result Falsifies assumptions for DML estimators with non-trivial power.
We information-theoretically reformulate two measures of capacity from statistical learning theory: empirical VC-entropy and empirical Rademacher complexity. We show these capacity measures count the number of hypotheses about a dataset that a learning algorithm falsifies when it finds the classifier in its repertoire …
Efficiently generates models resistant to falsification.
problem Creating models that cannot be disproven by tests.
method Exploits connections between high-dimensional multicalibration and expected variational inequality problems to develop an efficient algorithm.
result First to efficiently produce online outcome indistinguishable generative models resistant to infinite classes of tests.
This chapter studies emerging cyber-attacks on reinforcement learning (RL) and introduces a quantitative approach to analyze the vulnerabilities of RL. Focusing on adversarial manipulation on the cost signals, we analyze the performance degradation of TD(λ) and Q-learning algorithms under the manipulation. For TD($…
The standard taxonomy of predictive uncertainty is inconsistent with standard measures.
problem Uncertainty taxonomy and measure inconsistency
method Proof of inconsistency
result Uncertainty is not reducible to data collection
New method falsifies causal discovery results without ground truth.
problem Evaluation of causal discovery algorithms without ground truth data.
method Detects incompatibilities between causal graphs learned on different subsets of variables.
result Detection of incompatibilities can falsify wrongly inferred causal relations.
The paper argues that machine learning is a falsificationist process.
problem The role of falsification in machine learning is underexplored.
method The paper presents a falsificationist account of artificial neural networks, emphasizing empirical risk minimization and implicit regularization.
result Artificial neural networks can be seen as a falsificationist process, rejecting inadequate prediction rules.
A new metric assesses causal graphs using node permutations to detect inconsistencies.
problem Quantifying the goodness of causal graphs and distinguishing them from random graphs.
method Constructing a baseline through node permutations and comparing inconsistencies.
result The proposed metric can distinguish between true and wrong causal graphs.
There are (at least) three approaches to quantifying information. The first, algorithmic information or Kolmogorov complexity, takes events as strings and, given a universal Turing machine, quantifies the information content of a string as the length of the shortest program producing it. The second, Shannon information…
Survey of algorithms for testing AI-driven CPS safety.
problem Testing AI-driven CPS for safety in complex environments.
method Survey of applied algorithms for safety validation.
result Survey of existing tools and techniques for safety validation.
New method tightens bounds on causation probabilities using independent datasets.
problem Challenging point identification of causation probabilities without strong assumptions.
method Imposes counterfactual consistency between SCMs constructed from independent datasets and uses conditional mutual information.
result Significantly tighter bounds on causation probabilities are established.
New methods improve robust decision-making under uncertainty in off-policy evaluation.
problem Statistical uncertainty and causal considerations in off-policy evaluation.
method Marginal Ratio (MR) estimator, Conformal Off-Policy Prediction (COPP), causal bounds.
result Improved robustness and uncertainty quantification in off-policy decision-making.
Faster Tsetlin Machines use clause indexing to speed inference and learning.
problem Overfitting and slow inference in Tsetlin Machines.
method Introduced a look-up table that indexes clauses based on feature falsification, enabling faster evaluation of clauses.
result Up to 15 times faster classification and three times faster learning on MNIST and Fashion-MNIST.
GE finds failures in autonomous systems without domain heuristics.
problem Finding failures in autonomous systems without domain-specific heuristics.
method Adaptive stress testing using go-explore (GE) algorithm.
result GE finds failures in scenarios other RL techniques cannot solve.
Causal Bayesian networks interpret actions as interventions to connect models to real-world outcomes.
problem Connecting causal model predictions to real-world outcomes.
method Formal framework to interpret actions as interventions and prove impossibility results.
result No non-circular interpretation exists that satisfies natural desiderata without violating some.
A new method estimates treatment effects across multiple studies considering differences.
problem Estimating treatment effects across multiple studies with varying conditions.
method The multi-study R-learner framework that accounts for between-study heterogeneity.
result The multi-study R-learner is more efficient and normal than existing methods in the presence of heterogeneity.
Study shows convergence of Fubini-Study currents to equilibrium metrics on Kähler manifolds.
problem Convergence of Fubini-Study currents to equilibrium metrics in Kähler geometry.
method Analysis of continuous Hermitian metrics and their Fubini-Study currents on line bundles.
result The scaled difference between Fubini-Study currents and equilibrium metrics converges to zero in the sense of currents.
Boosting strategies for merging vs. ensembling studies analyzed.
problem Deciding between merging and ensembling studies for boosting.
method Analytical transition point and bias-variance decomposition for boosting with linear learners.
result Theoretical guidelines for merging vs. ensembling studies.
A critical decision point when training predictors using multiple studies is whether studies should be combined or treated separately. We compare two multi-study prediction approaches in the presence of potential heterogeneity in predictor-outcome relationships across datasets: 1) merging all of the datasets and traini…
Treatment recommendations within Clinical Practice Guidelines (CPGs) are largely based on findings from clinical trials and case studies, referred to here as research studies, that are often based on highly selective clinical populations, referred to here as study cohorts. When medical practitioners apply CPG recommend…
This article examines five common misunderstandings about case-study research: (1) Theoretical knowledge is more valuable than practical knowledge; (2) One cannot generalize from a single case, therefore the single case study cannot contribute to scientific development; (3) The case study is most useful for generating …
Acute respiratory infections have epidemic and pandemic potential and thus are being studied worldwide, albeit in many different contexts and study formats. Predicting infection from symptom data is critical, though using symptom data from varied studies in aggregate is challenging because the data is collected in diff…
This paper reviewed the machine learning-based studies for quantitative positron emission tomography (PET). Specifically, we summarized the recent developments of machine learning-based methods in PET attenuation correction and low-count PET reconstruction by listing and comparing the proposed methods, study designs an…
Ricci flow simulations show unstable Fubini-Study metrics develop singularities.
problem Understanding the behavior of unstable perturbations in Ricci flow.
method Numerical simulations of Ricci flow starting from unstable Fubini-Study metrics.
result Ricci flow solutions from unstable Fubini-Study metrics develop local singularities.
New method uncovers bias mechanisms in observational studies.
problem Understanding the sources of bias in observational studies.
method Analyzing the relationship between bias magnitude and nuisance function estimators' performance.
result Method can distinguish between common sources of causal bias.
GenAI improves actuarial practices through case studies.
problem Improving actuarial practices using AI.
method Four case studies using LLMs, Retrieval-Augmented Generation, and vision-enabled LLMs.
result GenAI enhances claim cost prediction, market comparisons, and car damage classification.
Proves polynomial injectivity of Fubini-Study map for ample line bundles.
problem Injectivity of Fubini-Study map for ample line bundles.
method Polynomial injectivity proof with polynomial dependence on ample line bundle exponent.
result Quantitative version of injectivity proved, polynomial in ample line bundle exponent.
Optimal ensemble construction improves prediction accuracy for multi-study tasks, especially in pandemic scenarios.
problem Poor out-of-study prediction performance due to heterogeneous datasets.
method Optimal ensemble construction using a two-stage stacking strategy that jointly estimates ensemble weights and study-specific model parameters.
result Our method outperforms multi-study stacking and other standard methods in predicting excess mortality during the pandemic.
Study evaluates machine learning for predicting treatment effects in observational studies.
problem Challenges in measuring treatment effects due to confounding bias in observational studies.
method Simulated two scenarios with and without confounding, using linear and non-linear relationships. Used machine learning models (linear regression, lasso regression, random forest) to predict counterfactuals and treatment effects.
result Machine learning models perform well under linearity but poorly under non-linearity, even in the presence of confounding.
In this paper, as a fundamental study on the theory of Morse functions and their higher dimensional versions or fold maps and applications to geometric theory of manifolds, which were started in 1950s by differential topologists such as Thom and Whitney and have been studied actively, we study algebraic and differentia…
Study evaluates and compares numerical differentiation methods on three case studies.
problem Evaluating and comparing numerical differentiation methods for efficiency.
method Forward, Backward, and Centered Finite-Difference methods applied at two levels of precision.
result Different methods perform differently across case studies, with varying levels of computational cost and accuracy.
Ablation studies show BCF model's propensity score is not essential for treatment effect estimation.
problem Understanding the necessity of propensity score in nonparametric treatment effect estimation.
method Partial ablation studies of Bayesian Causal Forest (BCF) model.
result Excluding estimated propensity score does not affect treatment effect estimation or uncertainty quantification.
Cognitive brain imaging is accumulating datasets about the neural substrate of many different mental processes. Yet, most studies are based on few subjects and have low statistical power. Analyzing data across studies could bring more statistical power; yet the current brain-imaging analytic framework cannot be used at…
Study finds dividend policy has no significant effect on IPO stock prices.
problem Impact of dividend policy on IPO price performance.
method Long-run performance statistics and GARCH model, dummy variable used.
result Dividend policy has no significant effect on IPO stock prices.
Study constant mean curvature tubes in homogeneous spaces.
problem Global geometry of constant mean curvature tubes.
method Screw-motion invariants, foliation, numerical isoperimetric profile.
result Foliation result and embeddedness proof.
Study on the convergence rate of prescribed scalar curvature flow.
problem Prescribing scalar curvature on manifolds.
method Inspired by Yamabe flow convergence rate study, analyze the prescribed scalar curvature flow convergence rate.
result Determine the convergence rate of the prescribed scalar curvature flow.
Study on convergence rate of weighted Yamabe flow.
problem Weighted Yamabe problem on smooth metric measure spaces.
method Weighted Yamabe flow and its convergence rate analysis.
result Study and analysis of convergence rate of the weighted Yamabe flow.
The causal assumptions, the study design and the data are the elements required for scientific inference in empirical research. The research is adequately communicated only if all of these elements and their relations are described precisely. Causal models with design describe the study design and the missing data mech…
Virtual knot theory is a generalization (discovered by the author in 1996) of knot theory to the study of all oriented Gauss codes. (Classical knot theory is a study of planar Gauss codes.) Graph theory studies non-planar graphs via graphical diagrams with virtual crossings. Virtual knot theory studies non-planar Gauss…