Unified calibration metrics improve forecast sharpness and accuracy.
problem Improving the sharpness of probabilistic forecasts while maintaining calibration.
method Kernel-based calibration metrics that unify and generalize existing methods for classification and regression.
result Enhanced calibration, sharpness, and decision-making across various tasks.
This work evaluates and benchmarks calibration metrics for data-driven regression models.
problem Conflicting results from different calibration metrics make it hard to compare and interpret model performance.
method Systematically extracted and benchmarked 14 regression calibration metrics across various data types and recalibration methods.
result Many metrics disagree on the same recalibration result, highlighting the need for careful metric selection.
This paper argues against using calibration metrics for assessing posterior probabilities and proposes expected proper scoring rules instead.
problem The assessment of posterior probabilities generated by machine learning classifiers using calibration metrics is flawed and should be replaced with expected proper scoring rules.
method The paper reviews proper scoring rules from a practical perspective, explains why expected PSRs are a principled measure of posterior quality, and introduces a new calibration metric called calibration loss.
result Calibration loss is superior to expected calibration error and expected score divergence calibration metrics for assessing posterior probabilities.
New decision-theoretic calibration error metric improves prediction reliability.
problem Improving the reliability of predictions for decision-making.
method Proposed Calibration Decision Loss (CDL) and an efficient algorithm to achieve near-optimal CDL.
result Near-optimal CDL guarantees vanishing payoff loss from miscalibration.
This paper reviews metrics to assess AI model calibration accuracy.
problem AI model probabilities do not always match their true accuracy.
method Comprehensive review of 82 probability calibration metrics.
result Identified 4 classifier families and 1 object detection family of metrics.
Reassesses calibration metrics in machine learning models.
problem Inconsistent reporting of calibration metrics in recent literature.
method Calibration-based decomposition of Bregman divergences, visualization of calibration and generalization error.
result New visualization technique for detecting trade-offs between calibration and generalization.
KCal calibrates deep networks by embedding logits in a metric space.
problem Overconfident predictions from DNNs, especially in high-risk applications.
method KCal learns a metric space on the penultimate-layer latent embedding and generates predictions using kernel density estimates.
result KCal provides a provable full calibration guarantee and consistently outperforms baselines.
New framework for evaluating multiclass classifier calibration.
problem Ensuring classifiers are well-calibrated for trustworthy predictions.
method Utility Calibration framework that measures calibration error relative to a utility function.
result Unified and robust interpretation of existing calibration metrics.
A new calibration metric bridges testability and actionability.
problem Combining testability and actionable insights for forecast probabilities.
method Cutoff Calibration Error (CCE) that assesses calibration over intervals of forecasted probabilities.
result Cutoff Calibration Error is both testable and actionable.
Overconfidence and underconfidence in machine learning classifiers is measured by calibration: the degree to which the probabilities predicted for each class match the accuracy of the classifier on that prediction. How one measures calibration remains a challenge: expected calibration error, the most popular metric, ha…
It is often observed that the probabilistic predictions given by a machine learning model can disagree with averaged actual outcomes on specific subsets of data, which is also known as the issue of miscalibration. It is responsible for the unreliability of practical machine learning systems. For example, in online adve…
The paper defines and studies hyperbolicity in calibrated manifolds and derives Schwarz lemmas.
problem Defining and studying hyperbolicity in calibrated manifolds.
method Introducing Rφ-hyperbolicity and φ-hyperbolicity, defining the KR φ-metric, and deriving Schwarz lemmas. result Characterization of φ-hyperbolic domains and extension of Schwarz lemma to calibrated geometries. Machine learning models deployed in real-world applications are often evaluated with precision-based metrics such as F1-score or AUC-PR (Area Under the Curve of Precision Recall). Heavily dependent on the class prior, such metrics make it difficult to interpret the variation of a model's performance over different subp…
New method uses interval-based metric to validate prediction uncertainty in machine learning.
problem Validation of prediction uncertainty in machine learning regression tasks is unreliable due to heavy-tailed distributions.
method Shift from variance-based metrics to interval-based Prediction Interval Coverage Probability (PICP).
result PICP method more quickly and reliably tests prediction intervals than variance-based metrics.
A new metric CKCE improves model calibration comparison.
problem Comparing the calibration of probabilistic models is challenging.
method CKCE based on Hilbert-Schmidt norm of conditional mean operators.
result CKCE provides more consistent and robust model calibration comparisons.
Given a transportation cost c:M×Mˉ→R, optimal maps minimize the total cost of moving masses from M to Mˉ. We find a pseudo-metric and a calibration form on M×Mˉ such that the graph of an optimal map is a calibrated maximal submanifold. We define the mass of space-like current…
Study optimizes tree-based models for better alignment of predicted scores and actual probabilities.
problem Traditional calibration metrics fail to align predicted scores with actual probabilities when score distributions deviate from the underlying data.
method Optimizes tree-based models (Random Forest, XGBoost) using Kullback-Leibler (KL) divergence to minimize the difference between predicted and true probability distributions.
result Optimized tree-based models yield superior alignment between predicted scores and actual probabilities without significant performance loss.
This paper improves confidence measurement in deep metric learning models.
problem Measuring confidence in deep metric learning models is challenging.
method Approximates class distributions using Gaussian kernel smoothing and calibrates the confidence metric.
result Improves generalization and robustness of deep metric learning models.
We study locally conformal calibrated G2-structures whose underlying Riemannian metric is Einstein, showing that in the compact case the scalar curvature cannot be positive. As a consequence, a compact homogeneous 7-manifold cannot admit an invariant Einstein locally conformal calibrated G2-structure unless the…
TCE measures calibration error with a test-based approach.
problem Measuring calibration error of probabilistic binary classifiers.
method TCE uses a novel loss function based on a statistical test.
result TCE offers clear interpretation, consistent scale, and enhanced visual representation.
Survey of methods to calibrate neural network predictions.
problem Ensuring neural networks provide accurate confidence levels.
method Empirical comparison of calibration methods.
result Various techniques for calibrating neural networks.
Survey on assessing and improving classifier calibration for better decision making.
problem Ensuring classifiers correctly quantify prediction uncertainty.
method Overview of principles, methods, and evaluation metrics for calibration.
result New methods and extensions from binary to multiclass settings.
Paper introduces new metrics for evaluating model accuracy.
problem Improving model accuracy and calibration.
method Developed two second-order accuracy metrics with integral and numerical representations.
result Validates model calibration settings and reveals distortions.
Complex classification performance metrics such as the Fβ-measure and Jaccard index are often used, in order to handle class-imbalanced cases such as information retrieval and image segmentation. These performance metrics are not decomposable, that is, they cannot be expressed in a per-example manner, which hinder…
The paper proves geodesics and conic sections are length-minimizing under specific metrics.
problem Finding shortest paths in complex geometries.
method Calibrations and conformal metrics.
result Geodesics and conic sections are length-minimizing.
We show that for a metric space with an even number of points there is a 1-Lipschitz map to a tree-like space with the same matching number. This result gives the first basic version of an unoriented Kantorovich duality. The study of the duality gives a version of global calibrations for 1-chains with coefficients in $…
New method calibrates classifier probabilities with guaranteed coverage.
problem Inaccurate probability estimates by classifiers in high-risk applications.
method Adaptive temperature scaling algorithm for conformal prediction.
result Improves calibration error measures and standard metrics across various tasks.
New tan-concavity property for Lagrangian phase operators helps in studying dHYM metrics.
problem Lack of concavity in Lagrangian phase operator for dHYM metrics.
method Introduce tangent Lagrangian phase flow (TLPF) on almost calibrated (1,1)-forms.
result TLPF exists for all positive time and converges to dHYM metrics under certain conditions.
New findings show that not all area-minimizing surfaces are calibrated, even on complex manifolds.
problem Understanding when area-minimizing surfaces cannot be calibrated.
method Analyzing homology classes and metrics on manifolds to determine if area-minimizers are calibrated.
result Calibrated area-minimizers are non-generic, challenging the common assumption that they are typical.
Study explores calibration properties in neural architectures.
problem Calibration issues in deep neural networks despite improved accuracy.
method Leverages Neural Architecture Search (NAS) to evaluate 117,702 neural networks.
result Identifies key architectural designs beneficial for calibration.
The paper calibrates uncertainty in dropout variational inference models.
problem Miscalibration of model uncertainty in dropout variational inference.
method Logit scaling methods are extended to recalibrate model uncertainty.
result Logit scaling reduces miscalibration, improving reliability of predictions.
This paper improves lottery ticketing by calibrating network confidence.
problem Uncalibrated confidence in lottery tickets leads to overconfidence and poor performance.
method The paper introduces various calibration strategies and explores their impact on lottery tickets.
result Calibration mechanisms consistently improve lottery ticket performance, even under distribution shifts.
Study shows label errors impact model disparity metrics, proposing mitigation methods.
problem Impact of label errors on model disparity metrics.
method Empirical study, characterizing label error effects; proposing estimation and relabeling methods.
result Label errors significantly affect model disparity metrics, particularly for minority groups.
Study of geometric properties of almost calibrated forms on Kähler manifolds.
problem Understanding the geometry of almost calibrated (1,1) forms on compact Kähler manifolds. method Investigates the infinite dimensional Riemannian manifold structure, CAT(0) geodesic metric space, and geodesics of the space of almost calibrated forms.
result The space of almost calibrated forms is an infinite dimensional Riemannian manifold with non-positive sectional curvature and CAT(0) geodesic metric space.
Proposes CCE to assess point-wise reliability of neural network predictions.
problem Overconfidence and misaligned predictive distributions in neural networks.
method Introduces Conditional Congruence (CCE) metric using conditional kernel mean embeddings.
result CCE exhibits correctness, monotonicity, reliability, and robustness in high-dimensional regression tasks.
The paper studies how to extend local calibration pairs to global ones in various situations. As a result, new discoveries involving mass-minimizing properties are exhibited. In particular, we show that a R-homologically nontrivial connected submanifold M of a smooth Riemannian manifold X is homologically…
New method calibrates heterogeneous treatment effect models.
problem Difficulty in estimating and calibrating heterogeneous treatment effects.
method Defined and proposed a robust estimator for HTE calibration, based on doubly robust treatment effect estimators.
result Proposed method evaluates calibration of learned HTE models, addressing overfitting and high-dimensionality.
New method calibrates deep networks by preserving top-k predictions.
problem Calibrated confidence scores for multi-class deep networks to avoid rare mistakes.
method Intra order-preserving functions combined with neural network architecture.
result Outperforms state-of-the-art methods in evaluation metrics.
Variational characterization of calibrated submanifolds in different contexts.
problem Characterize calibrated submanifolds using variational principles.
method Variational approach with special variations of ambient metrics and calibrations.
result Critical points of volume functional correspond to calibrated submanifolds.
We study the uniqueness of minimal submanifolds and the stability of the mean curvature flow in several well-known model spaces of manifolds of special holonomy. These include the Stenzel metric on the cotangent bundle of spheres, the Calabi metric on the cotangent bundle of complex projective spaces, and the Bryant--S…
A new framework for consistent segmentation evaluation reduces operating losses.
problem Inconsistent thresholding-based segmentation methods lead to suboptimal solutions.
method Developed a consistent ranking-based framework (RankDice/RankIoU) using Bayes rules and Dice-/IoU-calibration.
result The proposed framework is Dice-/IoU-calibrated and provides excess risk bounds and convergence rates.
A new framework evaluates LLM calibration in open-ended QA.
problem Evaluating LLM calibration in open-ended QA settings.
method Sem-ECE framework: sampling answers, grouping by semantic classes, and using frequencies as confidence.
result Sem-ECE estimators are unbiased and Sem2 achieves smaller calibration error. This research improves deep neural network calibration using a new loss function.
problem Improving probability calibration in deep neural networks.
method Introduces Focal Calibration Loss (FCL) to minimize Euclidean norm and penalize calibration error.
result FCL achieves state-of-the-art performance in both calibration and accuracy metrics.
The paper addresses calibration in label ranking, a structured prediction task.
problem Calibration in label ranking is not well understood and often poorly calibrated.
method Formalized calibration for label ranking, developed a hierarchy of notions, and empirically evaluated models.
result Popular label ranking models are often poorly calibrated, with differences between sub-ranking and top-k metrics.
This paper tackles non-identifiability in financial market simulations using multivariate time series data.
problem Non-identifiability issue in social simulation models, leading to indistinguishable simulated time series data.
method Proposes a maximization-based aggregation function to form a new calibration objective function using multiple time series features.
result Significant improvements in alleviating non-identifiability and achieving higher simulation fidelity.
This paper improves multi-class calibration methods using mutual information maximization-based binning.
problem Calibration of deep neural network predictions, especially for small prior classes.
method I-Max concept for binning, shared class-wise calibration strategy.
result Improves multi-class ranking and calibration performance using a small calibration set.
Study on estimating conditional risk in machine learning.
problem Estimating expected loss of prediction models given input features.
method Analyzed in classification and regression settings, showing equivalence to standard regression. Developed theoretical insights and empirical validation.
result Conditional risk calibration is distinct from existing uncertainty quantification problems.
We equip many non compact non simply connected surfaces with smooth Riemannian metrics whose isoperimetric profile is smooth, a highly non generic property. The computation of the profile is based on a calibration argument, a rearrangement argument, the Bol-Fiala curvature dependent inequality, together with new result…