Proposes a framework to explain KS deterioration in credit risk models.
problem Inconsistent and ad hoc diagnosis of KS decline in credit risk models.
method Counterfactual diagnostic framework attributing KS decline to sampling variability, portfolio composition, covariate shift, and residual deterioration.
result The proposed approach provides more interpretable and governance-relevant explanations than threshold-based review alone.
A new method controls risk for set predictors using cross-validation.
problem Inefficient set predictors when data limited.
method Cross-validation conformal risk control (CV-CRC).
result CV-CRC offers theoretical guarantees and reduces set size.
The paper validates a centrality measure for financial networks during financial distress.
problem Systemic risk and shock propagation in financial networks.
method Statistical validation method for network centrality measures.
result The proposed centrality measure increases significantly during financial distress.
ECV method optimizes ensemble parameters for randomized ensembles.
problem Efficient tuning of ensemble parameters in randomized ensembles.
method ECV (Extrapolated Cross-Validation) method for tuning ensemble and subsample sizes.
result ECV yields δ-optimal ensembles for squared prediction risk.
We improve prediction risk estimation for large datasets using sketching and ridge regression.
problem Estimating prediction risks for large datasets efficiently and accurately.
method Random matrix theory, generalized cross validation, sketched ridge regression ensembles, and ensemble trick.
result Consistent risk estimation and prediction intervals for large-scale datasets.
New IF method improves accuracy in deep neural networks with noisy data.
problem Inaccurate influence estimates in deep neural networks, especially with noisy data.
method Established a connection between influence estimation error, validation set risk, and sharpness, introducing a novel estimation form for flat validation minima.
result Our novel Influence Function approach provides more accurate influence estimates, validated across various tasks.
A new model validation framework for agentic AI systems based on POMDPs.
problem Model validation of agentic AI systems.
method A POMDP-based framework for belief-state, forecast, and policy validation.
result The framework decomposes autonomous decision making into information, beliefs, forecasts, actions, and utility.
Proposes a method to generate counterfactuals for ensemble models using entropic risk measures.
problem Finding a single counterfactual explanation for an ensemble of models.
method Incorporates entropic risk measure into a constrained optimization to generate counterfactuals valid for an adjustable fraction of models.
result Entropic risk measure allows generation of counterfactuals valid for all models in the ensemble under a limiting case.
The study evaluates financial risk using copulas and statistical tests.
problem Validating bivariate forecasts in risk evaluation.
method Using copulas to characterize dependencies, applying statistical tests to validate forecasts, removing heteroskedasticity.
result A Student copula accurately describes financial time series dependencies.
The paper extends conformal risk control to be valid with high probability over a growing calibration dataset.
problem Valid risk control over a growing calibration dataset.
method Quantile-based arguments for anytime-valid control.
result Guarantees remain valid with high probability over a cumulatively growing calibration dataset.
This paper describes the current taxonomy of model risk, ways for its mitigation and management and the importance of the model validation function in collaboration with other departments to design and implement them.
The paper analyzes the risk of CV-tuned regularized estimators and connects it to SURE.
problem Understanding the risk of CV-tuned regularized estimators.
method Derives asymptotic risk function of CV-tuned estimators and connects it to SURE.
result The risk function provides a more detailed picture of predictive performance than uniform bounds.
CSA fills a gap in RLVR-trained LLM deployment by providing anytime-valid selective risk control.
problem Deployment of RLVR-trained LLMs in regulated organizations requires a safety certificate for every round without waiting for long-run averages.
method CSA uses a (test statistic, validity guarantee, deployment rule) framework to fill the gap, maintaining a Ville-type e-process per threshold on a Bonferroni grid.
result CSA provides the first anytime-valid selective risk control for RLVR-trained LLMs, matching the long-run average certification rate and satisfying pathwise validity and non-refusing deployment on every cell.
Four geometries govern sequential and distribution-free inference.
problem Sequential and distribution-free inference challenges.
method Four distinct admissibility geometries.
result Four classes of admissible procedures are pairwise non-nested.
Cross-validation is the workhorse of modern applied statistics and machine learning, as it provides a principled framework for selecting the model that maximizes generalization performance. In this paper, we show that the cross-validation risk is differentiable with respect to the hyperparameters and training data for …
Financial institutions face new model risks with AI, requiring enhanced model risk management.
problem New model risks from Generative AI applications in financial institutions.
method Enhanced model risk framework with additional testing and controls.
result Financial institutions need to enhance their model risk management for Generative AI applications.
Framework mitigates risk non-monotonicity in high-dimensional predictions.
problem Risk non-monotonicity in high-dimensional predictions.
method Model-agnostic framework using cross-validation and data-driven methodologies (zero- and one-step).
result Modified prediction procedures achieve monotonic asymptotic risk behavior.
This text is a survey on cross-validation. We define all classical cross-validation procedures, and we study their properties for two different goals: estimating the risk of a given estimator, and selecting the best estimator among a given family. For the risk estimation problem, we compute the bias (which can also be …
Study on time-varying APT validity in Japanese stock market.
problem Validity of Arbitrage Pricing Theory (APT) in Japanese stock market over time.
method Rolling window method applied to Fama and MacBeth's two-step regression and Kamstra and Shi's generalized GRS test.
result APT validity is unstable over time in Japanese stock market, influenced by monetary policy and business cycle.
New model solves equity premium puzzle.
problem Equity premium puzzle regarding risk behavior of investors.
method Developed a new tool called the sufficiency factor to analyze risk behavior of investors.
result Validated the new model with a coefficient of relative risk aversion of 1.033526.
This work establishes always-valid risk bounds for online matrix completion.
problem Challenges in establishing always-valid concentration inequalities for online matrix completion.
method Combines non-asymptotic martingale concentration and regularized low-rank matrix regression.
result Establishes always-valid risk bound process for online matrix completion.
This paper improves model selection with cross-validation using domain knowledge.
problem Improving model selection with cross-validation risk estimation.
method Establishes distribution-free deviation bounds using VC dimension, formalizes Learning Spaces based on domain knowledge.
result Enhanced generalization through selection of candidate models based on domain knowledge.
Hybrid framework predicts Arctic permafrost decline, risks infrastructure, and provides tools.
problem Tackles permafrost decline and infrastructure risk assessment in Arctic territories.
method Hybrid physics-machine learning framework integrating 2.9 million observations.
result Projects mean permafrost fraction decline of -20.3 pp under RCP8.5 forcing, with high-risk zones identified.
Corrects GCV for inconsistent risk estimation in finite ensembles of penalized estimators.
problem Inconsistent risk estimation of GCV for finite ensembles of penalized estimators.
method Identifies a correction involving an additional scalar correction based on degrees of freedom adjusted training errors from each ensemble component.
result CGCV maintains computational advantages of GCV and is model-free uniformly consistent for ridge regression.
RandALO speeds up risk estimation for large datasets.
problem Estimating out-of-sample risk for large, high-dimensional models.
method RandALO: a randomized approximate leave-one-out estimator.
result RandALO is a computationally efficient risk estimator in high dimensions.
Analysis of cross-validation for early-stopped gradient descent in high-dimensional regression.
problem Inconsistency of GCV for early-stopped GD in high-dimensional least squares regression.
method Theoretical analysis of GCV and LOOCV applied to early-stopped GD in high-dimensional least squares regression.
result LOOCV converges uniformly to the prediction risk of early-stopped GD, while GCV is generically inconsistent.
Cross-validation under sample selection bias can, in principle, be done by importance-weighting the empirical risk. However, the importance-weighted risk estimator produces sub-optimal hyperparameter estimates in problem settings where large weights arise with high probability. We study its sampling variance as a funct…
New insights into ridge regression with correlated data, improving risk prediction.
problem Understanding and predicting risk in ridge regression with correlated samples.
method Random matrix theory and free probability for asymptotic analysis; modified GCV estimator (CorrGCV) for unbiased prediction.
result GCV estimator fails for out-of-sample risk with correlated data; CorrGCV provides an unbiased estimator.
In this article, we derive concentration inequalities for the cross-validation estimate of the generalization error for empirical risk minimizers. In the general setting, we prove sanity-check bounds in the spirit of \cite{KR99} \textquotedblleft\textit{bounds showing that the worst-case error of this estimate is not m…
New model solves equity premium puzzle with risk aversion coefficient.
problem Equity premium puzzle in financial markets.
method Developed a new model incorporating investor risk behavior, tested with specific coefficients.
result Validated model with empirical studies, confirming coefficient of 1.033526.
Study identifies key ESG variables for assessing financial risk.
problem Assessing financial risk from ESG data with many variables.
method Proposed framework for hierarchical ESG data, selecting relevant variables.
result Selected ESG variables are more relevant to financial risk than aggregated scores.
Study ridge ensembles in proportional feature-to-sample size regime, proving risk equivalence and GCV consistency.
problem Characterizing and optimizing ridge ensembles in proportional feature-to-sample size regimes.
method Proportional asymptotics analysis, GCV for tuning, proving risk equivalence.
result Risk of optimal full ridgeless ensemble matches optimal ridge predictor's risk.
ASRI index detects crypto market risks with high precision and lead time.
problem Detecting systemic risks in cryptocurrency markets.
method Four weighted sub-indices (Stablecoin, DeFi, Contagion, Regulatory) validated against historical crises.
result ASRI detects significant abnormal signals with high statistical significance and lead time.
Extends risk control to adaptive data collection, anytime-valid guarantees.
problem Ensuring safety of machine learning models with critical risk measures.
method Sequential risk controlling prediction sets (RCPS) for adaptive data collection and active labeling.
result Anytime-valid guarantees for risk control in sequential data collection.
Risk bounds for Classification and Regression Trees (CART, Breiman et. al. 1984) classifiers are obtained under a margin condition in the binary supervised classification framework. These risk bounds are obtained conditionally on the construction of the maximal deep binary tree and permit to prove that the linear penal…
SCRIB assigns multiple labels to each example to control class-specific prediction risks.
problem Lack of a sound mechanism to decide when to refrain from predicting in DL classifiers.
method Set-classifier with Class-specific Risk Bounds (SCRIB) that assigns multiple labels to each example and controls class-specific prediction risks.
result SCRIB obtained class-specific risks 35%-88% closer to the target risks than baseline methods.
Model predicts cannabis use disorder risk for adolescents and young adults.
problem Predicting cannabis use disorder progression in adolescents and young adults.
method Bayesian machine learning model trained on longitudinal data.
result Model provides personalized risk assessment with AUC of 0.68-0.75 and E/O ratio of 0.95-1.
This paper explores how train-validation splits help in NAS to prevent overfitting.
problem NAS overfits with train-validation splits and needs better generalization guarantees.
method Established refined properties of validation loss and risk for NAS.
result NAS with train-validation splits can select the most generalizable model.
Bayesian framework forecasts financial tail risks using realized volatility and nonlinear thresholds.
problem Forecasting financial tail risks using realized volatility and nonlinear thresholds.
method Bayesian Markov Chain Monte Carlo method for model estimation; nonlinear threshold regression specification.
result The proposed framework produces competitive tail risk forecasts compared to GARCH and Realized-GARCH models.
Study analyzes smart contract adoption under bounded risk, showing stable adoption but fragile financial outcomes.
problem Understanding smart contract adoption in derivative markets under risk constraints.
method Structural theory linked with simulation and real-world validation.
result Adoption intensity is stable but profitability and service outcomes are sensitive to volatility.
We consider a priori generalization bounds developed in terms of cross-validation estimates and the stability of learners. In particular, we first derive an exponential Efron-Stein type tail inequality for the concentration of a general function of n independent random variables. Next, under some reasonable notion of s…
Study uses copulas and DCC-GARCH for multivariate risk analysis of VaR and CVaR.
problem Multivariate risk analysis for Value at Risk (VaR) and Conditional Value at Risk (CoVaR).
method Copulas and Dynamic Conditional Correlation (DCC)-GARCH models applied to historical financial data.
result Comparison of different copula families for goodness-of-fit and effectiveness.
New method combines experimental and observational data for causal inference.
problem Combining internal validity of experiments and larger sample sizes of observations.
method Empirical risk minimization (ERM) framework with cross-validation.
result Efficacy and reliability demonstrated on real and synthetic data.
Develops a framework to control risk in online learning models.
problem Rigorous uncertainty quantification for online learning models.
method A framework for constructing uncertainty sets that provably control risk.
result Guarantees risk control at any user-specified level even with distribution shifts.
Risk estimation is at the core of many learning systems. The importance of this problem has motivated researchers to propose different schemes, such as cross validation, generalized cross validation, and Bootstrap. The theoretical properties of such estimates have been extensively studied in the low-dimensional setting…
New criterion assesses cluster separability for validation.
problem Validating cluster analysis results and determining the number of clusters.
method Distinguishability criterion, combined loss function-based framework.
result Validated cluster configurations and determined the number of clusters.
New risk measures assess cryptocurrency market vulnerabilities during financial distress.
problem Capturing systemic risk in cryptocurrency markets during financial distress.
method Introducing Vulnerability Conditional Risk Measures (VCoES) and related measures.
result Validated theoretical insights and demonstrated practical relevance in cryptocurrency market.
Develops asymptotic theory for deep Cox models to enable valid inference.
problem Theoretical gaps in deep neural network estimators for Cox models.
method Asymptotic distribution theory linking in-sample optimization error to population risk.
result Pointwise and multivariate asymptotic normality for subsampled ensemble estimators.