Loss-calibrated EP improves Bayesian decision-making by focusing on utility-sensitive posterior approximations.
problem Bayesian decision-making under asymmetric utility functions.
method Loss-calibrated expectation propagation (Loss-EP) that tilts the posterior towards higher utility decisions.
result Loss-EP can capture useful information for decision-making under asymmetric penalties.
Paper introduces arctan pinball loss for XGBoost quantile regression.
problem Efficiently predicting multiple quantiles with XGBoost.
method Smooth approximation of pinball loss for XGBoost, using arctan pinball loss.
result Arctan pinball loss reduces quantile crossings and improves efficiency.
SIFT reduces training time by selecting samples with approximate losses.
problem Reducing training time by selecting samples with large approximate losses.
method Developed SIFT which uses early exiting to obtain approximate losses with intermediate layer representations for sample selection.
result SIFT achieves significant gains in training time and number of backpropagation steps without optimized implementation.
New expressive losses improve adversarial robustness without sacrificing accuracy.
problem Training networks for robustness at the expense of accuracy.
method Formalizing expressivity, using convex combinations of adversarial attacks and IBP bounds.
result Trivial expressive losses yield state-of-the-art results in various settings.
Unified framework approximates gradient descent's implicit bias in high dimensions.
problem Understanding gradient descent's behavior in overparameterized settings with convex losses.
method Unified framework for convex losses, including sensitivity analysis.
result Approximation of minimum-norm interpolation in high dimensions.
A new method for exponentially weighted moving models using approximations.
problem Efficiently updating moving averages for time series data.
method Approximates EWMM using a fixed window and quadratic term, solving non-growing problems.
result Approximation produces estimates similar to exact EWMM.
The impact of a stress scenario of default events on the loss distribution of a credit portfolio can be assessed by determining the loss distribution conditional on these events. While it is conceptually easy to estimate loss distributions conditional on default events by means of Monte Carlo simulation, it becomes imp…
We prove a law of large numbers for the loss from default and use it for approximating the distribution of the loss from default in large, potentially heterogenous portfolios. The density of the limiting measure is shown to solve a non-linear SPDE, and the moments of the limiting measure are shown to satisfy an infinit…
This paper concerns a method of selecting a subset of features for a sequential logit model. Tanaka and Nakagawa (2014) proposed a mixed integer quadratic optimization formulation for solving the problem based on a quadratic approximation of the logistic loss function. However, since there is a significant gap between …
This work tackles Bayesian neural networks by addressing loss landscape symmetries.
problem Understanding and optimizing the loss landscape of Bayesian neural networks.
method The approach involves extending marginalized loss barrier formalism to BNNs, proposing a matching algorithm to search for linearly connected solutions using permutation matrices and combinatorial optimization.
result Nearly zero marginalized loss barriers for linearly connected solutions were found.
A new method approximates expected empirical loss for stochastic deep learning tasks.
problem Determining optimal step sizes for stochastic gradient descent in deep learning.
method Applying one-dimensional function fitting to noisy losses of vertical cross sections to approximate expected empirical loss.
result The method leads to a robust and straightforward optimization method that performs well across datasets and architectures.
Improved online learning algorithms using ADP for adversarial environments.
problem Minimizing regret in adversarial online learning with vector-valued losses.
method Approximate dynamic programming to characterize lower Pareto frontier of expected losses.
result Improved performance bounds compared to existing online learning algorithms.
We analyze the fluctuation of the loss from default around its large portfolio limit in a class of reduced-form models of correlated firm-by-firm default timing. We prove a weak convergence result for the fluctuation process and use it for developing a conditionally Gaussian approximation to the loss distribution. Nume…
This paper proposes a new methodology to compute Value at Risk (VaR) for quantifying losses in credit portfolios. We approximate the cumulative distribution of the loss function by a finite combination of Haar wavelets basis functions and calculate the coefficients of the approximation by inverting its Laplace transfor…
New method calibrates Bayesian neural network approximations for better task-specific predictions.
problem Inaccurate approximations of Bayesian neural networks without task-specific knowledge.
method Introduces a loss-calibrated evidence lower bound informed by Bayesian decision theory.
result Achieves higher utility for applications with asymmetric utility functions.
This work provides efficient approximations for linear classifiers' performance.
problem Improving the efficiency and accuracy of linear classifiers.
method Developed smooth functions approximating the expected error and ranking loss of linear classifiers, derived from data moments.
result The proposed approximations and optimization algorithms achieve similar or better performance than state-of-the-art methods, significantly faster.
Proposes sigmoidF1 loss for multilabel classification, improving performance metrics.
problem Lack of smooth, tractable loss functions for multilabel classification.
method Introduces sigmoidF1, a smooth F1 score surrogate loss function.
result sigmoidF1 outperforms other loss functions on various datasets and metrics.
New PG losses improve decision optimization in misspecified models.
problem Improving decision optimization in models that are not perfectly specified.
method Introducing Perturbation Gradient (PG) losses to connect decision loss with directional derivatives and optimizing using gradient techniques.
result PG losses yield best-in-class policies asymptotically, even in misspecified settings.
A new method uses local sensitivity to improve importance sampling for approximating complex loss functions.
problem Approximating complex loss functions using subsampling with strong theoretical guarantees.
method Introducing local sensitivity to measure data point importance and using leverage scores for efficient estimation.
result Local sensitivity sampling can be efficiently estimated and used to approximate complex loss functions with strong guarantees.
Fine-grained analysis of gradient descent with momentum provides modified loss equations.
problem Understanding the dynamics of gradient descent with momentum.
method Fine-grained analysis and derivation of modified loss equations.
result Global approximation bounds and continuous modified equations for HB.
Study visualizes actor-critic loss landscapes for inventory optimization.
problem Difficulties in solving multi-store dynamic inventory control problems.
method Low-dimensional visualizations of actor loss function.
result Loss landscapes favor optimal policies in reinforcement learning.
The paper corrects Bayesian neural network approximations to improve decision quality.
problem Inaccurate posterior approximations in Bayesian neural networks lead to suboptimal decisions.
method Develops methods to calibrate approximate posterior predictive distributions for better decision making.
result Empirically produces higher quality decisions compared to previous methods.
Simplified screening tests for data points in optimization.
problem Discarding irrelevant data points in empirical risk minimization.
method Designing loss functions and regularizing convex losses to induce sparsity, using ellipsoidal approximations.
result Automatic discarding of data samples without losing optimization guarantees.
We set the context for capital approximation within the framework of the Basel II / III regulatory capital accords. This is particularly topical as the Basel III accord is shortly due to take effect. In this regard, we provide a summary of the role of capital adequacy in the new accord, highlighting along the way the s…
Positive results for agnostic regression with various losses.
problem Agnostic regression with bounded sample compression.
method Generic and efficient sample compression schemes for real-valued functions.
result Exact and approximate compression schemes for specific losses.
Gradient descent struggles to achieve zero loss in deep learning models due to non-generic data distributions.
problem Achieving zero loss minimizers in deep learning networks.
method Analysis of gradient descent algorithm in deep learning, focusing on underparametrized networks.
result Zero loss minimization cannot be achieved generically in deep learning networks.
This research introduces a line search method for deep learning that uses parabolic approximations.
problem Finding optimal step sizes for deep learning optimization.
method Parabolic approximation line search approach.
result The batch loss over lines in negative gradient direction is mostly convex locally and suitable for parabolic approximations.
Paper introduces a new robust loss function for RL.
problem Heuristic selection of threshold parameters in quantile Huber loss.
method Derived from Wasserstein distance, captures noise in quantile values.
result Enhances robustness against outliers and enables parameter adjustment.
Paper proves LCVB method's consistency in Bayesian posteriors and decision rules.
problem Approximating Bayesian posteriors and decision rules.
method Loss-calibrated variational Bayes (LCVB) method.
result LCVB method's consistency in both approximate posterior and decision rules.
A new method approximates loss functions asymmetrically to prevent catastrophic forgetting.
problem Catastrophic forgetting in deep neural networks.
method Approximating a true loss function using an asymmetric quadratic function with one side overestimated.
result Achieves state-of-the-art accuracy close to upper-bound performance on benchmark datasets.
New method approximates curvature from symmetries in deep networks.
problem Hard to approximate curvature in large deep networks.
method Analytically averaging over group actions that leave the loss invariant to construct structured Hessian approximations.
result Structured Hessian approximations from single gradients can be estimated, stored, and inverted.
Proposes a method for inference in high-dimensional classification with non-differentiable surrogate losses.
problem Lack of inference procedures for identifying driving factors in high-dimensional classification with non-differentiable surrogate losses.
method Kernel-smoothed decorrelated score and cross-fitted version for hypothesis tests and interval estimators.
result Valid and superior inference methods for high-dimensional classification with non-differentiable surrogate losses.
Proposes variational Gaussian approximations for solving the Kushner equation.
problem Solving the Kushner equation for state estimation with observations.
method Tractable variational Gaussian approximations of proximal losses based on Wasserstein and Fisher metrics.
result The proposed method leads to a Gaussian flow consistent with Kalman-Bucy and Riccati flows.
New approximations for Value at Risk and Expected Shortfall accounting for kurtosis.
problem Approximating Value at Risk and Expected Shortfall with positive skewness and kurtosis.
method Extensions of the Normal Power Approximation incorporating skewness and kurtosis.
result Improved precision for various loss distributions.
Exact minimization of saturated loss functions for robust regression and subspace estimation.
problem Minimizing saturated loss functions for robust regression and subspace estimation.
method Developed an exact algorithm with polynomial time-complexity for robust regression and subspace estimation, relating the problems to linear model approximation.
result Exact minimization of saturated loss functions for robust regression and subspace estimation is possible with polynomial time-complexity.
A fast, accurate method for estimating extreme quantiles in insurance and operational risk models.
problem Estimating extreme quantiles of compound loss distributions in insurance and operational risk models.
method Interpolated Single Loss Approximation (ISLA) and modified ISLA (MISLA).
result MISLA is comparable in speed and accuracy to the best competing method (PE2) and is easier to implement.
The Nyström method improves learning efficiency for convex losses.
problem Improving computational efficiency in empirical risk minimization.
method Using random subspaces to approximate hypothesis spaces in convex loss functions.
result Computational gains can be achieved without sacrificing learning performance for general convex Lipschitz losses.
New algorithms find conditions and linear rules with high probability and loss.
problem Finding conditions and rules with high probability and loss in conditional sparse regression.
method Efficient algorithms for identifying conditions and rules with optimal probability and loss.
result Achieved algorithms that nearly match the probability of the ideal condition and improve the approximation to the target loss.
A new federated learning algorithm improves on existing methods by exploiting data smoothness.
problem Federated learning optimization with smooth loss functions.
method Federated Low Rank Gradient Descent (FedLRGD) algorithm.
result FedLRGD outperforms Federated Averaging (FedAve) in federated oracle complexity under certain conditions.
Paper revisits weighted likelihood bootstrap and extends it to loss-likelihood bootstrap.
problem Generating samples from approximate Bayesian posterior of a parametric model.
method Bayesian nonparametric model with minimising expected negative log-likelihood.
result Loss-likelihood bootstrap method for posterior sampling.
Gradient-based methods find saddle points, not critical points, in neural networks.
problem Gradient-based optimization methods converge to saddle points rather than critical points in deep neural networks.
method Critical point-finding methods used to analyze neural network losses.
result Gradient-based methods often converge to or pass through gradient-flat regions, where gradient norm has a stationary point.
Deep neural networks' loss surfaces contain every low-dimensional pattern.
problem Finding arbitrary low-dimensional patterns in neural network loss surfaces.
method Empirical and theoretical analysis of loss landscapes of deep neural networks.
result Deep universal approximators exhibit a property where arbitrary smooth patterns exist in their loss surfaces.
Efficiently learns Single-Index Models with constant factor approximation.
problem Learning Single-Index Models under L22 loss with unknown link functions. method An efficient algorithm using alignment sharpness for optimization.
result Achieves constant factor approximation to optimal loss for various distributions and link functions.
New protocols for locally private learning of linear models with reduced data leakage.
problem Locally private learning of linear models with minimal data leakage.
method Noninteractive LDP convex optimization protocols for generalized linear losses and Euclidean median problems.
result First algorithms with sub-exponential dependence on dimensionality for nonsmooth losses.
Uniform diffusion approximation for SGD in non-convex settings.
problem Finite-time diffusion approximation for SGD.
method Establishing uniform-in-time diffusion approximation with strong convexity and mild conditions.
result Uniform-in-time diffusion approximation of SGD without convexity of each loss function.
Wasserstein GANs fail to approximate Wasserstein distance, leading to their success.
problem Approximating Wasserstein distance in deep generative models.
method Analysis of differences between theoretical setup and training reality.
result Wasserstein GANs' success is due to their failure to approximate Wasserstein distance.
New method achieves small-loss regret bounds in random-order model.
problem Online learning with adversarial loss functions in random order.
method Extending batch-to-online transformation, using average sensitivity and stability.
result Small-loss regret bounds of order ildeO(φ⋆(OPTT)). Method sanitizes IFM in CNN layers to control privacy loss.
problem Controlling privacy loss in CNNs using input feature maps.
method Sample-and-hold approximation scheme to sanitize IFM, unfolding tensors for independence from CNN configuration.
result Control the privacy loss by adjusting the sanitization degree.