Optimizing proper loss yields calibrated models under specific conditions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Pre-season prediction of crop production outcomes such as grain yields and N losses can provide insights to stakeholders when making decisions. Simulation models can assist in scenario planning, but their use is limited because of data requirements and long run times. Thus, there is a need for more computationally expe…
Boosted decision trees typically yield good accuracy, precision, and ROC area. However, because the outputs from boosting are not well calibrated posterior probabilities, boosting yields poor squared error and cross-entropy. We empirically demonstrate why AdaBoost predicts distorted probabilities and examine three cali…
Study assesses sugar beet yields under EU's neonicotinoids ban and climate change.
We present a new machine learning approach to estimate personalized treatment effects in the classical potential outcomes framework with binary outcomes. To overcome the problem that both treatment and control outcomes for the same unit are required for supervised learning, we propose surrogate loss functions that inco…
Improved diffusion bridge sampling with rKL-LD loss.
Unhinged loss minimization fails to improve classifier accuracy for simple data.
Loss minimization leads to multicalibration for neural networks.
This paper studies Fenchel-Young losses, a generic way to construct convex loss functions from a regularization function. We analyze their properties in depth, showing that they unify many well-known loss functions and allow to create useful new ones easily. Fenchel-Young losses constructed from a generalized entropy, …
We use surrogate losses to obtain several new regret bounds and new algorithms for contextual bandit learning. Using the ramp loss, we derive new margin-based regret bounds in terms of standard sequential complexity measures of a benchmark class of real-valued regression functions. Using the hinge loss, we derive an ef…
We consider the problem of learning a loss function which, when minimized over a training dataset, yields a model that approximately minimizes a validation error metric. Though learning an optimal loss function is NP-hard, we present an anytime algorithm that is asymptotically optimal in the worst case, and is provably…
Closed-form flow matching yields similar performance to stochastic version, improving model performance.
Study predicts doubling of U.S. maize insurance claims due to climate change.
This work uses PAC-Bayes for structured prediction with ILE, yielding insights and algorithms.
Mini-batch stochastic gradient descent (SGD) and variants thereof approximate the objective function's gradient with a small number of training examples, aka the batch size. Small batch sizes require little computation for each model update but can yield high-variance gradient estimates, which poses some challenges for…
We present a generalization of the Cauchy/Lorentzian, Geman-McClure, Welsch/Leclerc, generalized Charbonnier, Charbonnier/pseudo-Huber/L1-L2, and L2 loss functions. By introducing robustness as a continuous parameter, our loss function allows algorithms built around robust loss minimization to be generalized, which imp…
We propose a max-pooling based loss function for training Long Short-Term Memory (LSTM) networks for small-footprint keyword spotting (KWS), with low CPU, memory, and latency requirements. The max-pooling loss training can be further guided by initializing with a cross-entropy loss trained network. A posterior smoothin…
Optimal convex loss function improves regression coefficient estimation.
LoRA-Curve connects independent LoRA optima through continuous low-loss valleys, improving Bayesian model averaging.
New PG losses improve decision optimization in misspecified models.
Adaptive loss function formulation is an active area of research and has gained a great deal of popularity in recent years, following the success of deep learning. However, existing frameworks of adaptive loss functions often suffer from slow convergence and poor choice of weights for the loss components. Traditionally…
Motivated by the developments in cyber risk treatment in the finance industry, we propose a general framework of cyber bond, whose main purpose is to insure (compensate) losses of a cyber attack. Based on a database of publicly available cyber events, we determine cyber loss distribution parameters and use them to nume…
Malware detection is a popular application of Machine Learning for Information Security (ML-Sec), in which an ML classifier is trained to predict whether a given file is malware or benignware. Parameters of this classifier are typically optimized such that outputs from the model over a set of input samples most closely…
Locus scores predictions for risk, reducing large-loss events.
Introduces Fitzpatrick losses, tighter than Fenchel-Young losses.
Learning with non-modular losses is an important problem when sets of predictions are made simultaneously. The main tools for constructing convex surrogate loss functions for set prediction are margin rescaling and slack rescaling. In this work, we show that these strategies lead to tight convex surrogates iff the unde…
Unified framework for fair regression under demographic parity.
Optimistic algorithm reduces regret and constraint violations in online convex optimization with adversarial constraints.
Proposes new loss functions for GANs to improve estimation accuracy and robustness.
Symmetric losses improve classifier robustness from corrupted labels.
The paper decomposes probabilistic scores into reliability, uncertainty, and information loss.
Analyzes adversarial training's impact on loss landscape, proposing PAS to improve model performance.
Loss minimisation fails to capture epistemic uncertainty in second-order predictors.
Paper improves learning rates for GSC loss functions using iterated Tikhonov regularization.
Focal loss improves classification but not class-posterior probability estimation.
Fisher loss improves deep domain adaptation by learning discriminative within-class compact and between-class separable representations.
Structured entropy improves classification performance on structured targets.
Optimizes bond portfolios to avoid worst-case losses.
This work analyzes two methods for combining multiple binary labels in bipartite ranking.
In this paper, we present and illustrate some new tools for rigorously analyzing training data selection methods. These tools focus on the information theoretic losses that occur when sampling data. We use this framework to prove that two methods, Facility Location Selection and Transductive Experimental Design, reduce…
Paper proposes self-supervised method for accurate speaker diarization.
In this work, we study data preconditioning, a well-known and long-existing technique, for boosting the convergence of first-order methods for regularized loss minimization. It is well understood that the condition number of the problem, i.e., the ratio of the Lipschitz constant to the strong convexity modulus, has a h…
Neural network model improves robustness of mortgage bond yield curve estimation.
New expressive losses improve adversarial robustness without sacrificing accuracy.
Replacing MSE with f-divergence in diffusion models improves robustness under data contamination.
It has been argued in the past that high-dimensional neural networks do not exhibit local minima capable of trapping an optimisation algorithm. However, the relationship between loss surface modality and the neural architecture parameters, such as the number of hidden neurons per layer and the number of hidden layers, …
Bayesian neural networks are shown to be minimax and admissible under certain conditions.
The stability of the financial system is associated with systemic risk factors such as the concurrent default of numerous small obligors. Hence it is of utmost importance to study the mutual dependence of losses for different creditors in the case of large, overlapping credit portfolios. We analytically calculate the m…