ABROCA assesses algorithmic bias, revealing skewed distributions that inflate results.
problem Detecting nuanced performance differences in classifier fairness.
method Study of ABROCA metric's statistical properties under various conditions.
result ABROCA distributions are skewed, inflating results by chance in imbalanced classes.
Satellite imagery improves house price prediction models.
problem Improving accuracy of housing price estimation models.
method Transfer learning from ImageNet-pretrained Inception-v3 model to satellite images.
result Achieved a 10% improvement in R-squared score.
Study GLS estimator properties in multivariate regression with heteroskedastic and autocorrelated errors.
problem Asymptotic properties of GLS estimator in multivariate regression with specific error structures.
method Derive Wald statistics for linear restrictions and assess their performance.
result Wald statistics remain robust to heteroskedasticity and autocorrelation.
This paper presents an assessment of global economic energy potentials for all major natural energy resources. This work is based on both an extensive literature review and calculations using natural resource assessment data. Economic potentials are presented in the form of cost-supply curves, in terms of energy flows …
Framework for assessing fairness across similar predictive models.
problem Fairness in predictive models across different groups.
method Develops a framework for characterizing fairness over the set of good models under selective labels.
result Framework can replace or audit models for better fairness properties.
Paper introduces new risk measures for Kelly criterion.
problem Aggressive Kelly criterion investment strategy.
method Unified approach to risk assessment in Kelly criterion.
result Two new measures for quantifying risk.
The study improves the assessment of fairness in face recognition using ROC curves and statistical guarantees.
problem Improving the assessment of fairness in face recognition systems.
method Proves asymptotic guarantees for empirical ROC curves and fairness metrics, and introduces a recentering technique to avoid bootstrap pitfalls.
result Demonstrates the practical relevance of the methods for assessing fairness in face recognition systems.
Robust method estimates self-similarity for mammogram images, improving cancer detection.
problem Statistical assessment of self-similarity in real data with large mean level shifts.
method Theil-type weighted regression for wavelet-based estimation, compared to OLS and AV.
result Robust approach shows nearly 68% accuracy in cancer vs non-cancer classification.
Solvency II Directive 2009/138/EC requires an insurance and reinsurance undertakings assessment of a Solvency Capital Requirement by means of the so-called "Standard Formula" or by means of partial or full internal models. Focusing on the first approach, the bottom-up aggregation formula proposed by the regulator permi…
We show how risk measures originally defined in a model free framework in terms of acceptance sets and reference assets imply a meaningful underlying probability structure. Hereafter we construct a maximal domain of definition of the risk measure respecting the underlying ambiguity profile. We particularly emphasise li…
E-scores assess LLM outputs for correctness, addressing p-hacking issues.
problem Limited principled mechanisms to assess generative model correctness.
method Use e-values to complement LLM outputs with e-scores, providing flexibility in tolerance levels.
result Achieves guarantees of correctness assessment and upper bounds size distortion.
Improves pre-trial risk assessments by making them safer without changing existing rules.
problem Improving pre-trial risk assessments while maintaining deterministic rules.
method Developed a maximin robust optimization approach to find a safer policy.
result Can safely improve certain components of the risk assessment instrument.
The ROC curve is widely used to assess the quality of prediction/classification/ranking algorithms, and its properties have been extensively studied. The precision-recall (PR) curve has become the de facto replacement for the ROC curve in the presence of imbalance, namely where one class is far more likely than the oth…
The paper analyzes the risk of investing in a basket of 27 cryptocurrencies using statistical distributions.
problem Risk assessment of capital allocation in a basket of cryptocurrencies.
method Used statistical tests to determine the most appropriate distribution (SDI) for modeling returns, and adapted the generalized Pareto distribution for tail risk assessment.
result Found that a combination of stable and generalized Pareto distributions provides a more accurate risk assessment for the basket of cryptocurrencies.
Proposes CCE to assess point-wise reliability of neural network predictions.
problem Overconfidence and misaligned predictive distributions in neural networks.
method Introduces Conditional Congruence (CCE) metric using conditional kernel mean embeddings.
result CCE exhibits correctness, monotonicity, reliability, and robustness in high-dimensional regression tasks.
This paper introduces a new property of estimators of the strength of statistical association, which helps characterize how well an estimator will perform in scenarios where dependencies between continuous and discrete random variables need to be rank ordered. The new property, termed the estimator response curve, is e…
Framework for assessing explainable AI systems.
problem Lack of consensus on explainability properties.
method Survey of literature, development of taxonomy and descriptors.
result Operationalization of the framework in Explainability Fact Sheets.
Verifying probabilistic forecasts for extreme events is a highly active research area because popular media and public opinions are naturally focused on extreme events, and biased conclusions are readily made. In this context, classical verification methods tailored for extreme events, such as thresholded and weighted …
We present a new approach to assessing the robustness of neural networks based on estimating the proportion of inputs for which a property is violated. Specifically, we estimate the probability of the event that the property is violated under an input model. Our approach critically varies from the formal verification f…
The study simplifies assessing overlap in logistic regression models using empirical likelihood.
problem Assessing overlap in multidimensional logistic regression models.
method Translation of Silvapulle's condition to empirical likelihood maximization, mechanized with R code.
result Minimal overlapping structures are cataloged in dimensions less than four, providing rules for higher dimensions.
Study uses deep learning to quickly estimate tissue properties for personalized radio-frequency dosimetry.
problem Time-consuming tissue segmentation limits personalized radio-frequency dosimetry.
method Developed a learning-based approach using Convolutional Neural Networks (CNN) for magnetic resonance images.
result Smooth distribution of dielectric properties improves SAR distribution consistency.
AI systems need reliable testing to ensure safety and trustworthiness.
problem Current AI Act lacks functional trustworthiness for AI systems.
method Define technical application distribution, set risk-based performance, and conduct statistically valid testing.
result Reliable functional trustworthiness is essential for AI systems.
New test assesses probabilistic model calibration without expensive approximations.
problem Assessing calibration of probabilistic models with scores.
method Kernel Calibration Conditional Stein Discrepancy (KCCSD) test using new score-based kernels.
result Control over type-I error with improved scalability and efficiency.
The bootstrap provides a simple and powerful means of assessing the quality of estimators. However, in settings involving large datasets, the computation of bootstrap-based quantities can be prohibitively demanding. As an alternative, we present the Bag of Little Bootstraps (BLB), a new procedure which incorporates fea…
We show how to reduce the problem of computing VaR and CVaR with Student T return distributions to evaluation of analytical functions of the moments. This allows an analysis of the risk properties of systems to be carefully attributed between choices of risk function (e.g. VaR vs CVaR); choice of return distribution (p…
New conditional risk measures called conditional generalized quantiles defined and characterized.
problem Developing new risk measures for dynamic risk assessment.
method Propose and characterize conditional generalized quantiles using expected utility model and equivalent conditions.
result Characterized conditional generalized quantiles as well-defined and equivalent to a conditional first order condition.
Study examines dependence properties of Bayesian neural network units in finite-width networks.
problem Understanding dependence properties of hidden units in practical finite-width Bayesian neural networks.
method Theoretical analysis and empirical evaluation of depth and width impacts.
result Hidden units in finite-width Bayesian neural networks are dependent, contrary to the infinite-width limit assumption.
In this article we propose a novel measure of systemic risk in the context of financial networks. To this aim, we provide a definition of systemic risk which is based on the structure, developed at different levels, of clustered neighbours around the nodes of the network. The proposed measure incorporates the generaliz…
In this paper, as a first step in examining the properties of a feasible portfolio subset that is characterized by budget and risk constraints, we assess the maximum and minimum of the investment concentration using replica analysis. To do this, we apply an analytical approach of statistical mechanics. We note that the…
The paper introduces a US crime index to assess financial losses from property and cyber crimes.
problem Lack of indices evaluating crime's financial impact on investments.
method Developed an index-based insurance portfolio using FBI financial losses data.
result Real estate, ransomware, and government impersonation are major risk contributors.
New method evaluates visual explanations of deep models using adversarial perturbations.
problem Lack of objective evaluation of visual explanations of deep models.
method Proposes an adversarial perturbation approach to evaluate visual explanations of deep models.
result Demonstrates the effectiveness of the proposed approach through comparisons with existing methods.
Graph networks struggle with multi-task learning due to varying property loss surface curvatures.
problem Graph networks underperform in multi-task learning for crystal and molecule properties.
method Assessed curvature of property loss surfaces via spectral properties of Hessians, matrix-free using randomized numerical linear algebra.
result Varying curvature of property loss surfaces explains graph networks' multi-task learning inefficiency.
The average portfolio structure of institutional investors is shown to have properties which account for transaction costs in an optimal way. This implies that financial institutions unknowingly display collective rationality, or Wisdom of the Crowd. Individual deviations from the rational benchmark are ample, which il…
One of the biggest challenges in the research of generative adversarial networks (GANs) is assessing the quality of generated samples and detecting various levels of mode collapse. In this work, we construct a novel measure of performance of a GAN by comparing geometrical properties of the underlying data manifold and …
Persistent homology enhances graph classification by capturing long-range graph properties.
problem Lack of formal assessment of persistent homology in graph learning.
method Brief introduction and theoretical discussion of persistent homology in graph context, followed by empirical analysis.
result Persistent homology improves graph classification, especially for data with prominent topological structures.
Framework assesses autograders' reliability and biases.
problem Mixed reliability and biases in autograders for LLM evaluation.
method Bayesian GLMs to model evaluation outcomes.
result Explicit quantification of scoring differences and biases.
Paper assesses how features influence classification of COVID-19 patients.
problem Evaluating the influence of features on classification problems.
method Introduced a measure of feature influence using Shapley value of cooperative games, with axiomatic characterisation.
result Validated the performance of the measure through experiments on COVID-19 patients.
Bayesian test assesses dependence between mixed data types.
problem Assessing dependence between text, image, and sound data.
method Bayesian kernelised correlation test using Dirichlet process model.
result Demonstrated effectiveness compared to other methods.
The paper introduces CoCoCat bonds for multi-region natural catastrophes, accounting for complex dependencies.
problem Valuation of multi-region contingent convertible bonds under complex dependencies.
method Developed a model accounting for inter-regional dependencies using change-of-measure techniques.
result Significant impact of inter-regional dependencies on CoCoCat bond pricing.
The paper examines statistical properties of IL and LVR in automated market makers.
problem Assessing the performance of automated market makers and their profitability.
method Analysis of random walk properties and statistical integral combined with CFMM mechanics.
result IL and LVR have identical expectation values but different distribution functions for Brownian motion.
Protein structure prediction has been a grand challenge problem in the structure biology over the last few decades. Protein quality assessment plays a very important role in protein structure prediction. In the paper, we propose a new protein quality assessment method which can predict both local and global quality of …
The paper assesses the risk of negative treatment effects using bounds and inference.
problem Risk of negative treatment effects on a significant portion of the population.
method Characterizes tight bounds on the conditional value at risk (CVaR) of the individual treatment effect (ITE) distribution using covariate-conditional average treatment effect (CATE) function.
result Developed a debiasing method to estimate these bounds efficiently from data and construct confidence intervals, even in complex scenarios.
Proposes a decentralized insurance protocol for DeFi.
problem Over-insurance and inefficiencies in DeFi collateral.
method Smart contract-based economic model without external dependencies.
result Solves over-insurance and capital inefficiencies.
Study evaluates and compares traditional and causal machine learning methods for estimating direct price effects of environmental amenities.
problem Estimating direct price effects of environmental amenities in housing markets.
method Empirical Monte Carlo simulation to compare traditional regression and causal machine learning approaches.
result Causal Machine Learning (CML) methods, particularly causal forest DID, perform comparably to generalized DID in most scenarios.
Paper introduces active Bayesian method for assessing black-box classifiers efficiently.
problem Need to assess performance of black-box classifiers reliably with limited labels.
method Develops inference strategies and proposes active Bayesian framework for efficient instance selection.
result Significant gains in performance assessment with fewer labels compared to traditional methods.
Reply to Tetlock et al. on tail risk and probability gap.
problem Expert judgment fails to account for tail risk.
method Comparison of forecasting tournaments and extreme value theory.
result Greater gap between tail expectation and probability properties.
Risk measures such as Expected Shortfall (ES) and Value-at-Risk (VaR) have been prominent in banking regulation and financial risk management. Motivated by practical considerations in the assessment and management of risks, including tractability, scenario relevance and robustness, we consider theoretical properties of…
This work introduces benchmarks for evaluating nanophotonic structures in design simulations.
problem Design and understanding of nanophotonic structures for various applications.
method Development of frameworks and benchmarks for evaluating nanophotonic structures in parametric design problems.
result Strategic use of evaluation fidelity in enhancing structure designs.