We study consistency properties of surrogate loss functions for general multiclass learning problems, defined by a general multiclass loss matrix. We extend the notion of classification calibration, which has been studied for binary and multiclass 0-1 classification problems (and for certain other specific learning pro…
This research analyzes the consistency of convex and nonconvex surrogate losses for adversarially robust classification.
problem Ensuring classifiers are robust to adversarial perturbations.
method Analysis of convex and nonconvex surrogate losses through the lens of calibration.
result No convex surrogate loss is calibrated with respect to the adversarial 0-1 loss for linear models, but nonconvex losses can be calibrated under certain conditions.
New method simplifies checking consistency of differentiable loss functions.
problem Verifying consistency of differentiable loss functions is difficult.
method Developed a new approach called strong indirect elicitation (strong IE) to simplify checking consistency.
result Strong IE is equivalent to calibration for strongly convex, differentiable surrogates.
The study offers new theoretical insights into structured prediction with convex loss minimization.
problem The challenge of structured prediction with efficient convex surrogate loss minimization.
method Constructing a convex surrogate loss and proving tight bounds on the calibration function.
result Formalizes the intuition that some task losses make learning harder than others, and that 0-1 loss is ill-suited for general structured prediction.
Study on calibration and consistency of adversarial surrogate losses.
problem Designing robust classifiers with theoretical guarantees.
method Extensive analysis of H-calibration and H-consistency of adversarial surrogate losses.
result Some convex loss functions and supremum-based convex losses are not H-calibrated for important hypothesis sets.
The study develops a theory for structured prediction using smooth convex surrogates.
problem Developing a theoretical framework for structured prediction.
method Characterizing smooth convex surrogates compatible with task losses and deriving statistical guarantees.
result Derives tight bounds for the calibration function and novel results for existing surrogate frameworks.
This paper develops convex surrogates for optimizing the multi-label F-measure.
problem Optimizing the F-measure for multi-label classification is computationally hard.
method Designing convex surrogate losses calibrated for the F-measure.
result The F-measure for multi-label problems has a rank of at most s2+1. The paper analyzes top-k classification and proposes consistent loss functions.
problem Understanding consistency of top-k classification in challenging tasks.
method Theoretical analysis, defining top-k calibration, proposing new loss functions.
result Proposes a new consistent hinge loss and a top-k calibrated convex loss.
New framework quantifies learning guarantees for inconsistent convex surrogates.
problem Analyzing consistency properties of machine learning methods with inconsistent convex surrogates.
method Extending the framework of Osokin et al. (2017) to inconsistent surrogates, introducing a new lower bound on the calibration function.
result Shows how learning with inconsistent surrogates can have guarantees on sample complexity and optimization difficulty.
This paper characterizes and designs loss functions for robust classification with abstention.
problem Ensuring robustness against adversarial attacks and knowing when to abstain from prediction.
method Proposes adversarial robust reject option loss and characterizes surrogates for calibration.
result Shifted Double Ramp Loss and Shifted Double Sigmoid Loss satisfy the calibration conditions.
We simplify deriving calibration functions for multiclass classification.
problem Deriving calibration functions for multiclass classification is time-consuming and case-specific.
method We introduce a streamlined analysis to derive calibration functions for multiple surrogate losses.
result We provide explicit derivation of calibration functions for various multiclass classification losses.
We present surrogate regret bounds for arbitrary surrogate losses in the context of binary classification with label-dependent costs. Such bounds relate a classifier's risk, assessed with respect to a surrogate loss, to its cost-sensitive classification risk. Two approaches to surrogate regret bounds are developed. The…
Bayesian framework predicts aerodynamic uncertainty from sparse measurements.
problem Calibrating aerodynamic models with sparse and uncertain measurements.
method Bayesian latent Gaussian process for surrogate model calibration.
result Calibrated surrogate model accurately predicts aerodynamic uncertainty.
New methods help calibrate complex ABMs more efficiently.
problem Calibrating parameters in complex ABMs is challenging.
method Integrates different sampling methods and surrogate models.
result Surrogate assisted methods perform better than standard methods.
This paper improves risk bounds and calibration for smart predict-then-optimize method.
problem Improving risk bounds and calibration for smart predict-then-optimize method.
method Develops risk bounds and uniform calibration results for the SPO+ loss relative to the SPO loss.
result Empirical minimizer of the SPO+ loss achieves low excess true risk with high probability.
We study losses for binary classification and class probability estimation and extend the understanding of them from margin losses to general composite losses which are the composition of a proper loss with a link function. We characterise when margin losses can be proper composite losses, explicitly show how to determ…
Study on learning to defer to multiple experts with consistent surrogates and confidence calibration.
problem Addressing the open problems of consistent surrogates, confidence calibration, and ensembling of experts.
method Derive two consistent surrogates (softmax and OvA) and propose a conformal inference technique for choosing experts.
result The OvA-based loss does not cause mis-calibration propagation, while the softmax-based loss does.
EnsLoss combines multiple loss functions to prevent overfitting in classification.
problem Preventing overfitting in classification models.
method EnsLoss is an ensemble method that combines loss functions, ensuring calibration and consistency.
result EnsLoss improves classification accuracy compared to fixed loss methods.
SVI and GP surrogates improve calibration of ABMs in epidemiology.
problem Calibrating stochastic ABMs in epidemiology is computationally expensive.
method Stein Variational Inference (SVI) with Gaussian process (GP) surrogates.
result SVI maintains comparable predictive accuracy and calibration effectiveness to MCMC.
We propose a new framework to improve the calibration of neural networks.
problem Improving the accuracy of model confidence predictions.
method Introducing a differentiable surrogate for expected calibration error (DECE) and a meta-learning framework to optimise model hyper-parameters for validation set calibration.
result Achieved competitive performance with existing calibration approaches.
Novel convex surrogate for non-modular loss functions.
problem Computational tractability for non-modular loss functions.
method Submodular-supermodular decomposition, slack-rescaling, Lov{á}sz hinge.
result First tractable solution for non-modular loss functions.
Novel convex surrogate for submodular losses with tractable computation.
problem Learning with non-modular losses for set prediction.
method Proposed Lovász hinge loss function for submodular losses.
result First tractable convex surrogates for submodular losses.
Study shows uncertainty calibration improves BO performance, but not as much as model type.
problem Effect of model uncertainties on Bayesian optimization performance.
method Extensive study comparing different surrogate models and their uncertainty calibration.
result Gaussian Processes outperform other models in BO, and uncertainty calibration does not significantly improve regret.
The paper proposes methods to directly optimize complex classification metrics.
problem Handling class-imbalanced cases with non-decomposable metrics.
method Calibrated surrogate maximization of linear-fractional utility.
result Calibrated surrogate maximization can coincide with true utility maximization under certain conditions.
We establish linear regret bounds for convex smooth losses using Fenchel-Young losses.
problem Establishing linear regret bounds for convex smooth losses.
method Constructing a convex smooth surrogate loss using Fenchel-Young losses generated by the convolutional negentropy.
result We derive a smooth loss with a linear surrogate regret bound.
Study on H-consistency bounds for machine learning surrogates.
problem Estimating target loss error relative to surrogate loss error in machine learning.
method Developed H-consistency bounds for various surrogates and loss functions. result Stronger guarantees than existing methods, offering distribution-dependent and -independent bounds.
We carefully study how well minimizing convex surrogate loss functions, corresponds to minimizing the misclassification error rate for the problem of binary classification with linear predictors. In particular, we show that amongst all convex surrogate losses, the hinge loss gives essentially the best possible bound, o…
Establishes a condition for multiclass classification-calibration of Gamma-Phi losses.
problem Ensuring classification-calibration of multiclass Gamma-Phi losses.
method Develops a general sufficient condition for classification-calibration of Gamma-Phi losses.
result Proves the first family of nonconvex multiclass surrogate losses for which classification-calibration has been fully justified.
Study improves Bayesian calibration of mechanical properties using active learning and MCMC.
problem Inference of spatially varying material parameters in computational mechanics.
method Comprehensive comparative study of surrogate models and MCMC algorithms.
result Active learning strategy outperforms a priori trained models in posterior estimation.
This study improves hyperparameter optimization for categorical and non-normal data.
problem Bayesian hyperparameter optimization struggles with categorical hyperparameters and non-normal data.
method Integrates conformalized quantile regression to address estimation weaknesses and provides robust calibration guarantees.
result Quantile surrogate architectures and acquisition functions yield superior performance compared to existing methods.
CJE calibrates cheap LLM judges against an oracle, achieving high accuracy at a fraction of the cost.
problem Inexpensive LLM judges can produce biased rankings, leading to unreliable outcomes.
method CJE uses a small oracle to calibrate cheap scores, then evaluates at scale with valid uncertainty.
result CJE achieves 99% pairwise ranking accuracy at 14x lower cost compared to a 16x oracle/judge cost ratio.
Paper extends SMM to weakly convex and multi-convex surrogates for non-convex optimization.
problem Non-convex optimization with weakly convex or multi-convex surrogates.
method Stochastic majorization-minimization with proximal regularization or block-minimization.
result Convergence rates for empirical and expected losses under non-i.i.d. data.
Proposes a robust framework for multiclass classification.
problem General multiclass classification with adversarial robustness.
method Dual formulation as convex optimization with adversarial surrogate loss.
result Competitive performance in multiclass classification problems.
Study of estimation errors in surrogate loss minimizers, providing stronger guarantees than existing methods.
problem Estimation errors in surrogate loss minimizers for various hypothesis sets.
method Detailed study of H-consistency estimation error bounds, proving general theorems for distribution-dependent and independent settings. result Explicit bounds for zero-one and adversarial losses, showing enhancements under distributional assumptions.
Space mapping calibrates financial models, shown feasible for Heston model.
problem Calibrating financial models with few observable parameters and non-linear constraints.
method Space mapping approach using a coarse surrogate model and fine model calibration.
result Space mapping approach feasible for Heston model calibration.
This work improves interpretability and calibration of complex-valued neural networks using Newton-Puiseux analysis.
problem Insufficient interpretability and probability calibration of complex-valued neural networks.
method Newton-Puiseux framework to examine local decision geometry, fitting a polynomial surrogate and factorizing it using Newton-Puiseux expansions.
result Enhanced Expected Calibration Error in ECG and wireless modulation datasets compared to uncalibrated softmax and standard post-hoc baselines.
A new autoregressive SPO method improves decision-making for dependent data.
problem Improving decision-making for dependent data in stochastic optimization.
method An autoregressive Smart Predict-then-Optimize (SPO) method for time series data.
result Generalization bounds and uniform calibration results for the SPO loss in autoregressive models.
A new L2D system produces calibrated probabilities of expert correctness without sacrificing accuracy.
problem Calibration of learning to defer systems for safety.
method One-vs-all classifiers with a consistent surrogate loss function.
result Proposes a calibrated L2D system that outperforms existing methods in accuracy and calibration.
Adversarial consistency depends on the uniqueness of adversarial Bayes classifiers.
problem Consistency of adversarial surrogate losses is not guaranteed.
method Connected consistency of adversarial surrogate losses to the uniqueness of adversarial Bayes classifiers.
result A convex surrogate loss is statistically consistent for adversarial learning if and only if the adversarial Bayes classifier is unique.
The paper proposes a framework for structured prediction using projection oracles.
problem Structured prediction with improved loss functions.
method A general framework for deriving loss functions using convex sets and projection oracles.
result Projections onto the marginal polytope can make the loss smaller and are computationally efficient.
Study inverse problems with measure samples, improving estimator calibration and recovery.
problem Inverse problems with unknown potentials observed through measure samples.
method Introduced convex empirical objectives and sharpened Fenchel--Young losses for finite-dimensional potential classes.
result High-probability parameter recovery bounds for inverse entropic unbalanced optimal transport and inverse JKO learning.
Paper proposes equivalent Lipschitz surrogates for zero-norm and rank optimization problems.
problem Optimization problems involving zero-norm and rank functions.
method Reformulate as MPECs, use global exact penalty, eliminate dual variable to get surrogates.
result Obtained equivalent Lipschitz surrogates for zero-norm and rank optimization problems.
The paper explores trading off consistency and dimensionality in convex surrogates for multiclass classification.
problem Designing consistent surrogate losses for multiclass classification with high-dimensional outcomes.
method Investigates embedding outcomes into convex polytopes and examining consistency under low-noise assumptions.
result Consistency can be achieved with less than n−1 dimensions, but hallucination occurs for some distributions. We develop a framework for consistent polyhedral surrogates in classification and prediction.
problem Designing consistent polyhedral surrogates for classification and prediction problems.
method Formalizing and studying embeddings of predictions as points in R^d, assigning original loss values, and convexifying to create surrogates.
result Established a strong connection between embeddings and polyhedral surrogates, providing constructions and proofs of consistency or inconsistency.
Optimizes hard-to-optimize metrics using adaptive surrogates.
problem Training models with black-box and hard-to-optimize metrics.
method Expresses metric as a function of surrogates, solves optimization problem over relaxed surrogate space.
result Approach performs on par with known methods and adds value when metric form is unknown.
XGB-Chiarella model generates realistic intra-day financial price data using agent-based models.
problem Generating accurate intra-day financial price data for research and risk management.
method Agent-based financial market simulation with XGBoost machine learning calibration.
result XGB-Chiarella model accurately reflects real market behaviours and generates realistic price time series.
In statistical learning theory, convex surrogates of the 0-1 loss are highly preferred because of the computational and theoretical virtues that convexity brings in. This is of more importance if we consider smooth surrogates as witnessed by the fact that the smoothness is further beneficial both computationally- by at…
A robust multiclass logistic regression method using Tsallis divergence.
problem Noise robustness in multiclass logistic regression.
method Two-temperature logistic regression with Tsallis divergence.
result Significant robustness to outliers and noise.