Simulation workflow is a top-level model for the design and control of simulation process. It connects multiple simulation components with time and interaction restrictions to form a complete simulation system. Before the construction and evaluation of the component models, the validation of upper-layer simulation work…
Study validates ML-UQ calibration statistics using simulated reference values.
problem Validation of ML-UQ calibration statistics is lacking due to lack of predefined reference values.
method Proposed validation workflow using simulated reference values derived from synthetic datasets.
result Some statistics, like CC and ENCE, are overly sensitive to generative distribution choice.
Validates composite systems using discrepancy propagation.
problem Validation of industrial systems with costly real-world tests.
method Propagates bounds on distributional discrepancy measures through a composite system.
result Derives upper bound on real system failure probability from simulations.
Posterior SBC validates inference conditionally on observed data.
problem Validating inference for specific observed data.
method Simulation-based calibration checking (SBC) adapted to use posterior parameters.
result Validates inference conditionally on observed data.
Approach selects variables and time intervals for comparing high-dimensional time-series data.
problem Comparing high-dimensional time-series data for significant differences.
method Data is split into subintervals, and two-sample tests are performed on each to identify distinguishing variables.
result The approach effectively identifies variables and time intervals where data significantly differs.
Machine learning (especially reinforcement learning) methods for trading are increasingly reliant on simulation for agent training and testing. Furthermore, simulation is important for validation of hand-coded trading strategies and for testing hypotheses about market structure. A challenge, however, concerns the robus…
This paper introduces NPR, a technique to improve Bayesian inference for multi-modal, high-dimensional simulations.
problem Challenges in Bayesian inference for multi-modal, high-dimensional simulations.
method Introduces Neural Posterior Regularization (NPR) to enforce exploration of input parameter space.
result Empirically validated that NPR significantly improves performance on various simulation tasks.
AST provides a method to validate safe autonomy without unsafe simplifications.
problem Validation of safe autonomy in complex systems.
method Adaptive Stress Testing (AST) approach.
result AST can find failures without unsafe simplifications.
fintech-kMC simulates financial platforms for AI/ML model validation.
problem Validation of AI/ML models in real-world financial applications.
method Agent-based model with kinetic Monte Carlo engine.
result Generates realistic synthetic data for testing AI/ML models.
pmsims R package uses Gaussian process for flexible sample size estimation in clinical models.
problem Determining adequate sample size for clinical prediction models.
method Simulation-based Gaussian process search for flexible sample size estimation.
result Gaussian process-based method produces more stable sample size estimates, especially in challenging settings.
Proposes BSI for valid statistical inference on bandit algorithms.
problem Valid statistical inference on bandit algorithms' performance.
method Fits a simulator of the bandit environment from observed data and uses it to estimate mean reward under any policy.
result Proves asymptotically valid confidence intervals and maintains nominal coverage.
The paper develops a faster surrogate model for simulators using hybrid methods.
problem The need for faster validation of automotive technologies using simulators.
method Testing classical methods and building hybrid models combining them.
result A hybrid surrogate model outperforms classical methods in multivariate time series prediction.
New approach uses dynamic programming to efficiently discover failures in autonomous vehicle simulations.
problem Efficiently discovering rare failure events in autonomous vehicle simulations.
method Approximate dynamic programming and scene decomposition to estimate failure distribution.
result Increased number of failures discovered compared to baseline approaches.
We provide a bridge between generative modeling in the Machine Learning community and simulated physical processes in High Energy Particle Physics by applying a novel Generative Adversarial Network (GAN) architecture to the production of jet images -- 2D representations of energy depositions from particles interacting …
The Heston model is validated for option pricing using theoretical derivations and empirical market data.
problem Validating the Heston model for accurate option pricing.
method Theoretical derivations and empirical validations using Monte Carlo simulations and machine learning.
result The Heston model is robust and relevant for current financial markets.
Method estimates uncertainty in spatial predictions by defining an 'area of applicability'.
problem Uncertainty in predictions for new geographic locations far from training data.
method Proposes a dissimilarity index (DI) based on minimum distance to training data, and defines the 'area of applicability' (AOA) using a threshold on DI.
result Prediction error within the AOA is comparable to the cross-validation error of the model, while cross-validation error does not apply outside the AOA.
Paper revisits pre-validation method, improving hypothesis testing.
problem Improving hypothesis testing in pre-validated models with different feature dimensions.
method Extended problem formulation, analytical distribution, and bootstrap procedure.
result Proposed analytical distribution and bootstrap procedure for pre-validated predictors.
This work proposes validation diagnostics for SBI algorithms using Normalizing Flows.
problem Lack of appropriate validation methods for SBI algorithms with complex, high-dimensional data.
method Develops validation diagnostics based on Normalizing Flows with theoretical guarantees.
result Offers theoretical guarantees on consistency of NF-based estimators.
CausalSim corrects bias in trace-driven simulations for more accurate results.
problem Bias in trace-driven simulations due to system conditions during trace collection.
method CausalSim learns a causal model of system dynamics and latent factors from an RCT to remove bias from trace data.
result CausalSim reduces simulation errors by 53% and 61% compared to baselines, providing more accurate insights.
New method finds failures in high-fidelity simulators with fewer steps.
problem Finding failures in high-fidelity simulators is expensive and impractical.
method Adaptive stress testing with backward algorithm adaptation from low-fidelity to high-fidelity.
result Significantly fewer high-fidelity simulation steps needed to find failures.
Survey of algorithms for testing AI-driven CPS safety.
problem Testing AI-driven CPS for safety in complex environments.
method Survey of applied algorithms for safety validation.
result Survey of existing tools and techniques for safety validation.
Waldo method constructs valid confidence regions for simulator-based inference.
problem Constructing valid confidence regions for simulator-based inference with high-dimensional data.
method Reframes Wald test statistic and uses regression-based machinery for Neyman inversion.
result Waldo method produces conditionally valid and precise confidence regions.
Method selects valid IVs from a large set using clustering and test of overidentifying restrictions.
problem Selecting valid instrumental variables from a large set of candidates.
method Agglomerative hierarchical clustering combined with a test of overidentifying restrictions.
result Achieves oracle properties when the largest group of IVs is valid.
Combines trial and observational data to improve policy evaluation.
problem External validity of randomized trial results in target populations.
method Uses covariate data to model trial sampling and certifies policy evaluations.
result Valid trial-based policy evaluations under model miscalibration.
New emulator bridges simulators using conditional optimal transport.
problem Bridging simulators with minimal distortion.
method Flow-based approach to learn likelihood transport, COT-FM for optimal matching.
result Emulator accurately captures full correction between simulators.
The paper validates a centrality measure for financial networks during financial distress.
problem Systemic risk and shock propagation in financial networks.
method Statistical validation method for network centrality measures.
result The proposed centrality measure increases significantly during financial distress.
Study examines if LLMs' trading styles match real market behavior.
problem Lack of behavioral consistency in LLMs' trading strategies.
method Year-long simulations with LLMs, operationalizing behavioral finance drivers, and comparing with financial theory.
result LLMs' strategy switching is only partially consistent with behavioral finance theories.
We develop an approximate formula for evaluating a cross-validation estimator of predictive likelihood for multinomial logistic regression regularized by an ℓ1-norm. This allows us to avoid repeated optimizations required for literally conducting cross-validation; hence, the computational time can be significantl…
Cross-validation is one of the most popular model selection methods in statistics and machine learning. Despite its wide applicability, traditional cross validation methods tend to select overfitting models, due to the ignorance of the uncertainty in the testing sample. We develop a new, statistically principled infere…
New method improves spatial prediction validation accuracy.
problem Validation methods fail for spatial prediction tasks due to mismatch between validation and test locations.
method Proposes a new validation method that adapts existing covariate-shift ideas to spatial settings.
result Proves and demonstrates the new method's superiority in spatial prediction validation.
Study reveals gaps between simulated and real-world treatment effect evaluation metrics.
problem Evaluation of treatment effect estimation models differs between academic and practical settings.
method Comprehensive empirical study comparing semi-simulated benchmarks and real-world datasets.
result Counterfactual metrics do not reliably predict observable metrics, and rankings from simulated benchmarks do not generalize to real-world data.
Algorithm simulates counterfactuals for fairness analysis.
problem Analytical intractability of counterfactuals in conditional distributions.
method Proposes an algorithm using particle filtering for discrete and continuous variables.
result Asymptotically valid inference for counterfactuals.
Study predicts lens performance using neural networks.
problem Predicting visual acuity from lens designs.
method Used a CNN to classify Landolt Cs and validate its ability to predict VA from induced defocus.
result Validation showed consistent offset of +0.20 logMAR from simulated VA, comparable to clinical repeatability.
Bayesian approach improves car-following model calibration and validation with limited data.
problem Difficulties in calibrating car-following models using limited data for work zones.
method Bayesian programming for data analysis and parameter estimation.
result Bayesian methods enhance model calibration and validation accuracy.
Proposes a framework to explain KS deterioration in credit risk models.
problem Inconsistent and ad hoc diagnosis of KS decline in credit risk models.
method Counterfactual diagnostic framework attributing KS decline to sampling variability, portfolio composition, covariate shift, and residual deterioration.
result The proposed approach provides more interpretable and governance-relevant explanations than threshold-based review alone.
A fast bootstrap method estimates cross-validation standard error.
problem Uncertainty quantification in cross-validation estimates.
method Random-effects model to estimate variance component.
result Valid confidence intervals for model performance.
Improved surrogate model for field-valued QoIs using LF and HF simulations.
problem Accurate and efficient modeling of field-valued quantities under uncertain inputs.
method Bifidelity Karhunen-Loève expansion with active learning.
result Consistent improvements in predictive accuracy and sample efficiency.
AutoSimulate efficiently optimizes synthetic data generation.
problem Optimizing synthetic data generation for machine learning.
method Differentiable approximation of the objective function for efficient optimization.
result Significantly faster (up to 50x) and more efficient (up to 30x) synthetic data generation.
Instrumented data enables causal scientific machine learning
problem Insufficient data for causal scientific machine learning
method Instrumented data with explicit model, uncertainty, and counterfactuals
result Supports causal interventions through Pearl's do-operator
New method bounds hardware noise without assumptions.
problem Estimating hardware noise without assumptions.
method Machine Learning and Conformal Prediction.
result Theoretical upper bounds of fidelity.
Random matrix theory is used to assess the significance of weak correlations and is well established for Gaussian statistics. However, many complex systems, with stock markets as a prominent example, exhibit statistics with power-law tails, that can be modelled with Levy stable distributions. We review comprehensively …
Accurate model selection is a fundamental requirement for statistical analysis. In many real-world applications of graphical modelling, correct model structure identification is the ultimate objective. Standard model validation procedures such as information theoretic scores and cross validation have demonstrated poor …
The article proposes a method to make valid insurance claim predictions without relying on specific models.
problem Prediction of insurance claims using statistical models can be unreliable due to model misspecification, selection effects, and lack of finite-sample validity.
method The article employs conformal prediction, a machine learning strategy that is model-free and tuning-parameter-free, ensuring finite-sample validity.
result The proposed method guarantees valid predictions at a pre-assigned coverage probability level and performs well in insurance applications, including meeting Solvency II requirements.
While many statistical models and methods are now available for network analysis, resampling network data remains a challenging problem. Cross-validation is a useful general tool for model selection and parameter tuning, but is not directly applicable to networks since splitting network nodes into groups requires delet…
A method for efficient CV estimates in Bayesian hierarchical models.
problem Computational infeasibility of cross-validation in Bayesian hierarchical regression models.
method Conditioning on variance-covariance parameters to transform CV into an optimization problem.
result Equivalent or improved predictive estimates compared to full cross-validation.
This work improves safety validation of autonomous vehicles by finding interpretable failures.
problem Finding interpretable failures of autonomous systems in simulation.
method Signal temporal logic expressions optimized for high likelihood and human interpretability.
result Our methodology finds more interpretable failures with higher likelihood compared to baseline approaches.
The paper extends conformal risk control to be valid with high probability over a growing calibration dataset.
problem Valid risk control over a growing calibration dataset.
method Quantile-based arguments for anytime-valid control.
result Guarantees remain valid with high probability over a cumulatively growing calibration dataset.
Enhances multi-project scheduling with multiple priority rules.
problem Resource allocation in multi-project scheduling with limited time and resources.
method Simulation-based approach using composite priority rules.
result Increased probability of finding schedules with shortest duration.