This paper improves operational risk modeling by selecting better loss severity distributions.
problem Inconsistent regulatory capital calculations due to changing loss severity distribution families.
method Presented truncation probability estimates and a consistent quantile scoring function for selection criteria. Also, recommended collecting loss frequencies below the minimum reporting threshold.
result More stable regulatory capital calculations through better selection of loss severity distributions.
The study explores loss functions for learning distributions, finding the log loss and others are sufficient under certain conditions.
problem Understanding loss functions for distribution learning and density estimation.
method An axiomatic approach to design loss functions, proposing criteria and showing that no single loss function satisfies all criteria.
result No loss function satisfies all criteria, but the log loss and others do under the condition of candidate distributions being calibrated.
Paper introduces a new robust loss function for RL.
problem Heuristic selection of threshold parameters in quantile Huber loss.
method Derived from Wasserstein distance, captures noise in quantile values.
result Enhances robustness against outliers and enables parameter adjustment.
Wrapped loss function improves convergence and accuracy in multi-output models.
problem Nonconforming residual distributions in multi-output models.
method Proposes a 'Wrapped Loss Function' to regularize nonconforming residual distributions.
result Advanced properties of faster convergence, better accuracy, and improved handling of imbalanced data.
We study proper losses for discrete generative models without knowing the target distribution.
problem Evaluating generative models in the discrete setting without direct access to the target distribution.
method Define and construct black-box proper losses using statistical estimation theory.
result Black-box proper losses must be of polynomial form and involve more samples than the polynomial degree.
The impact of a stress scenario of default events on the loss distribution of a credit portfolio can be assessed by determining the loss distribution conditional on these events. While it is conceptually easy to estimate loss distributions conditional on default events by means of Monte Carlo simulation, it becomes imp…
Locus scores predictions for risk, reducing large-loss events.
problem Deployment cost from inaccurate predictions, especially large losses.
method Distribution-free loss-scale reliability score using any predictive distribution.
result Reduces large-loss frequency compared to standard heuristics.
The study analyzes a model for aggregate losses with dependent and overdispersed inter-losses times.
problem Analyzing aggregate loss models with dependent and overdispersed inter-losses times.
method The study uses a two-state Markovian arrival process (MAP2) and a Markov renewal process to model the inter-losses times. Severities are modeled using a heavy-tailed, double-Pareto Lognormal distribution. The model is estimated via direct maximization of the likelihood function.
result The model with dependence and overdispersion in inter-losses times leads to higher capital charges compared to a Poisson process.
Proposes a new loss function for distributional learning.
problem Learning sparse and singular distributions.
method Entropy-regularized optimal transport and Fenchel duality.
result Geometric loss results in unconstrained convex objective functions.
A new loss function improves neural networks' out-of-distribution detection without side effects.
problem Neural networks struggle with out-of-distribution detection due to SoftMax loss issues.
method Proposes IsoMax loss replacing SoftMax loss, maintaining high entropy and fast inferences.
result Significantly improves neural networks' out-of-distribution detection performance.
Estimation of the operational risk capital under the Loss Distribution Approach requires evaluation of aggregate (compound) loss distributions which is one of the classic problems in risk theory. Closed-form solutions are not available for the distributions typically used in operational risk. However with modern comput…
Generative models learn to capture target distribution support with extreme value loss.
problem Mode collapse in generative models for non-trivial target distributions.
method Optimizing against the minimal value of the loss function, rather than the mean.
result Models trained with extreme value loss learn to capture the support of the target distribution.
Study risk bounds for distributed ERM with general loss functions and hypothesis spaces.
problem Limited theoretical analysis for distributed ERM with general loss functions and hypothesis spaces.
method Derive tight risk bounds under assumptions on hypothesis space and loss function.
result Developed more general risk bound for distributed ERM without strong convexity restriction.
We consider distributed convex optimization problems originated from sample average approximation of stochastic optimization, or empirical risk minimization in machine learning. We assume that each machine in the distributed computing system has access to a local empirical loss function, constructed with i.i.d. data sa…
Factor copula models simplify joint default probabilities and loss distributions.
problem Modeling joint default probabilities and loss distributions for high-dimensional portfolios.
method Factor copula models that nest standard models and efficiently compute loss distributions.
result Exact and efficient computation of loss distributions for contingent claims.
New risk class defined based on loss location and deviation.
problem Risk assessment in loss distributions.
method Wrapper around smooth loss functions, M-estimators, stochastic gradient methods.
result Finite-sample stationarity guarantees for stochastic gradient methods.
Model predicts credit portfolio losses with contagion effects.
problem Predicting credit portfolio losses with contagion effects.
method Introduced a model with a recursive algorithm and flexible distributions.
result Good fit for synthetic CDO tranches of the iTraxx index.
Paper analyzes statistical properties of log-cosh loss function.
problem No statistical analysis of log-cosh loss function in literature.
method Presented statistical properties of log-cosh loss function, compared to Cauchy distribution, and examined various statistical procedures.
result Characterized statistical properties of log-cosh loss function, including distribution, likelihood function, and Fisher information.
This paper explains why distributional reinforcement learning is better than vanilla RL using small-loss bounds.
problem Understanding when and why distributional reinforcement learning (DistRL) is superior to vanilla reinforcement learning (RL).
method The paper uses small-loss bounds to explain the benefits of DistRL, proposing algorithms and proving bounds for different RL settings.
result Distributional reinforcement learning (DistRL) outperforms vanilla RL when optimal costs are small, as shown by small-loss bounds.
Proposes an alternative probabilistic interpretation of Huber loss.
problem Lack of intuitive understanding of Huber loss transition point.
method Relates Huber loss to Kullback-Leibler divergence between Laplace distributions.
result Identifies optimal transition point intuitively based on data noise.
Paper compares VaR from aggregated and single loss distributions in credit risk.
problem Estimating VaR in credit risk portfolios with varying severities.
method Uses Monte Carlo simulation with Gamma and truncated exponential distributions.
result Truncated exponential distribution yields VaR closer to aggregated loss approach.
Study of estimation errors in surrogate loss minimizers, providing stronger guarantees than existing methods.
problem Estimation errors in surrogate loss minimizers for various hypothesis sets.
method Detailed study of H-consistency estimation error bounds, proving general theorems for distribution-dependent and independent settings. result Explicit bounds for zero-one and adversarial losses, showing enhancements under distributional assumptions.
Under the Basel II standards, the Operational Risk (OpRisk) advanced measurement approach is not prescriptive regarding the class of statistical model utilised to undertake capital estimation. It has however become well accepted to utlise a Loss Distributional Approach (LDA) paradigm to model the individual OpRisk loss…
Feature noise causes loss discrepancies across groups even with equal data.
problem Loss discrepancies observed in learning procedures across different groups.
method Characterized the effect of feature noise on loss discrepancy in linear regression.
result Feature noise leads to loss discrepancy even when groups have equal data.
We study cross-country GDP losses due to financial crises in terms of frequency (number of loss events per period) and severity (loss per occurrence). We perform the Loss Distribution Approach (LDA) to estimate a multi-country aggregate GDP loss probability density function and the percentiles associated to extreme eve…
Agents prefer non-diversification in markets with extreme losses.
problem Optimal risk allocation and equilibria in markets with extremely heavy-tailed losses.
method Analysis of super-Pareto loss distributions and stochastic dominance.
result Non-diversification is preferred in markets with super-Pareto losses.
Flexible framework for bounding high-loss predictions using quantiles.
problem Need for rigorous guarantees in risk-sensitive applications.
method Order statistics of loss values, flexible quantile-based metrics.
result Ability to rigorously control loss quantiles on real-world datasets.
This work improves regression performance by using distributional losses, finding better optimization leads to improved generalization.
problem Improving regression performance in reinforcement learning.
method Introduced a novel distributional regression loss and investigated its effects on optimization and generalization.
result The novel distributional regression loss leads to improved prediction accuracy and better optimization.
This paper examines the Histogram Loss for regression, revealing its effectiveness without needing complex tuning.
problem Improving regression models by learning the entire distribution.
method Investigates Histogram Loss, a method that minimizes cross-entropy between a target distribution and a histogram prediction.
result The performance gain in regression models using Histogram Loss comes from optimization improvements, not extra modeling.
The paper explores how different loss functions impact reinforcement learning algorithms.
problem Improving reinforcement learning algorithms by optimizing loss functions.
method Comprehensive survey on loss functions in reinforcement learning, proving the benefits of specific loss functions.
result Binary cross-entropy loss leads to first-order bounds and is more efficient than squared loss.
The study finds a trade-off between model size, test loss, and training loss for linear predictors.
problem Finding the optimal balance between model size, test loss, and training loss for linear predictors.
method Established an algorithm and distribution-independent trade-off using non-asymptotic analysis.
result Models with low test loss are either classical (close to noise level training loss) or modern (large number of parameters).
Paper studies Fenchel-Young losses for classifier construction.
problem Creating effective loss functions for classifiers.
method Analyzes Fenchel-Young losses from generalized entropies, formulates conditions for separation margins and sparse support.
result Fenchel-Young losses can induce predictive distributions with separation margins and sparse support.
Study large deviations in life insurance portfolios without identical distributions.
problem Large deviations in life insurance portfolios with bounded losses and variances.
method Upper bound from standard large deviations, counterexample for full large deviation principle.
result Exponential bound for average loss exceeding a threshold.
Proposes PER loss to regularize neural network activations to normal distribution.
problem Improving neural network generalization and training speed.
method Regularizes activations to standard normal distribution via projected error function and Wasserstein distance.
result Minimizes Wasserstein distance between activation distribution and standard normal.
A new method optimizes anomaly scoring from score distribution to improve AD performance.
problem Vulnerability to anomaly contamination and lack of adaptability in existing AD methods.
method Optimizes anomaly scoring function from score distribution perspective, using Overlap loss.
result Overlap loss-based AD models significantly outperform state-of-the-art methods.
Proposes a robust framework for multiclass classification.
problem General multiclass classification with adversarial robustness.
method Dual formulation as convex optimization with adversarial surrogate loss.
result Competitive performance in multiclass classification problems.
Paper proposes SinkhornDRL for distributional RL using Sinkhorn divergence and regularized Wasserstein loss.
problem Improving distributional reinforcement learning by minimizing Bellman return distribution differences.
method Introduces SinkhornDRL, a distributional RL algorithm using Sinkhorn divergence and regularized Wasserstein loss.
result SinkhornDRL consistently outperforms or matches existing algorithms on Atari games, especially in multi-dimensional reward settings.
New methods connect low-loss points on neural network surfaces.
problem Connecting low-loss points on neural network loss surfaces.
method Macroscopic distributional assumptions and global connection models.
result Accuracy correlates with complexity and sensitivity.
Study on loss probabilities for diversified financial systems with light-tailed claims.
problem Analyzing risks in systems of diversified financial agents with light-tailed claims.
method Assuming exponentially distributed claims, we derive conditional loss distributions and compare with heavy-tailed claims.
result Conditional loss distributions reveal different risk profiles for agents and systems compared to heavy-tailed claims.
Over-parameterized models reduce Out-of-Distribution (OOD) generalization loss.
problem Understanding how over-parameterized models handle non-trivial distributional shifts.
method Investigating random feature models and examining non-trivial natural distributional shifts.
result Increasing model parameterization reduces OOD loss.
Optimizes hybrid insurance contracts for heavy-tailed losses.
problem Providing insurance against heavy-tailed losses with finite expected loss.
method Combines traditional and parametric insurance, using a Pareto-type criterion for optimization.
result The hybrid contract outperforms traditional contracts in simulations and real data.
Proposes new loss functions for better handling bimodal predictive uncertainty.
problem Bimodal predictive uncertainty in machine learning models.
method Family of distribution-aware loss functions integrating normalized RMSE with Wasserstein and Cramér distances.
result Proposed loss functions reduce predictive uncertainty estimation error by 45% on complex bimodal datasets.
Gradient descent struggles to achieve zero loss in deep learning models due to non-generic data distributions.
problem Achieving zero loss minimizers in deep learning networks.
method Analysis of gradient descent algorithm in deep learning, focusing on underparametrized networks.
result Zero loss minimization cannot be achieved generically in deep learning networks.
Paper develops a generative model using Wasserstein-2 loss.
problem Creating realistic data samples from limited data.
method Uses a distribution-dependent ODE with a gradient flow for W2 loss.
result The method converges to the true data distribution exponentially.
This paper extends stock trading results to include stop-loss orders.
problem Generalizing stock trading results with stop-loss orders.
method Geometric Brownian motion model, affine feedback controller, closed-form expression for cumulative distribution function.
result Affine feedback controller with stop-loss order generalizes results without stop-loss orders.
We prove a law of large numbers for the loss from default and use it for approximating the distribution of the loss from default in large, potentially heterogenous portfolios. The density of the limiting measure is shown to solve a non-linear SPDE, and the moments of the limiting measure are shown to satisfy an infinit…
LLMs learn peaked distributions slowly due to power-law losses.
problem Slow convergence of loss in training large language models.
method Systematic analysis of toy models and empirical evaluation of LLMs.
result Power-law time scaling with an exponent of 1/3 for learning peaked distributions.
Paper tackles distributed high-dimensional regression with quantile loss, overcoming heavy-tailed noise challenges.
problem High-dimensional linear regression with heavy-tailed noise.
method Adopting quantile regression loss, transforming response variable, and using gradient information for distributed estimation.
result Proposed distributed estimator achieves near-oracle convergence rate and supports recovery without machine number restrictions.