A new method approximates expected empirical loss for stochastic deep learning tasks.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The goal of regression and classification methods in supervised learning is to minimize the empirical risk, that is, the expectation of some loss function quantifying the prediction error under the empirical distribution. When facing scarce training data, overfitting is typically mitigated by adding regularization term…
We present -loss, , a tunable loss function for binary classification that bridges log-loss () and - loss (). We prove that -loss has an equivalent margin-based form and is classification-calibrated, two desirable properties for a good surrogate loss function for the ideal y…
The paper studies risk-sensitive learning schemes and provides learning bounds for empirical OCE minimizers.
Estimates MLP expected output without sampling, using fewer FLOPs.
Given two networks with the same training loss on a dataset, when would they have drastically different test losses and errors? Better understanding of this question of generalization may improve practical applications of deep networks. In this paper we show that with cross-entropy loss it is surprisingly simple to ind…
We present GLASSES: Global optimisation with Look-Ahead through Stochastic Simulation and Expected-loss Search. The majority of global optimisation approaches in use are myopic, in only considering the impact of the next function value; the non-myopic approaches that do exist are able to consider only a handful of futu…
Multi-output is essential in machine learning that it might suffer from nonconforming residual distributions, i.e., the multi-output residual distributions are not conforming to the expected distribution. In this paper, we propose "Wrapped Loss Function" to wrap the original loss function to alleviate the problem. This…
Study asset pricing with reference-dependent preferences, finding matching equity premia.
In this paper we study the differentially private Empirical Risk Minimization (ERM) problem in different settings. For smooth (strongly) convex loss function with or without (non)-smooth regularization, we give algorithms that achieve either optimal or near optimal utility bounds with less gradient complexity compared …
Submodularity is studied for convex risk measures, including Expected Shortfall.
Examines optimal risk sharing with realistic risk attitudes, finding risk seeking in certain subdomains.
We study the regret of optimal strategies for online convex optimization games. Using von Neumann's minimax theorem, we show that the optimal regret in this adversarial setting is closely related to the behavior of the empirical minimization algorithm in a stochastic process setting: it is equal to the maximum, over jo…
Optimizes hybrid insurance contracts for heavy-tailed losses.
Uniform deviation bounds limit the difference between a model's expected loss and its loss on an empirical sample uniformly for all models in a learning problem. As such, they are a critical component to empirical risk minimization. In this paper, we provide a novel framework to obtain uniform deviation bounds for loss…
Develops uniform convergence guarantees for a broad class of risk functionals in supervised learning.
Study analyzes landscape complexity of empirical loss functions with correlated data.
New risk control method for non-monotonic losses in complex parameters.
Bayesian approach uses Gaussian process for reinforcement learning.
We consider regression with square loss and general classes of functions without the boundedness assumption. We introduce a notion of offset Rademacher complexity that provides a transparent way to study localization both in expectation and in high probability. For any (possibly non-convex) class, the excess loss of a …
We propose a robust inferential procedure for assessing uncertainties of parameter estimation in high-dimensional linear models, where the dimension can grow exponentially fast with the sample size . Our method combines the de-biasing technique with the composite quantile function to construct an estimator that …
The Alternating Direction Method of Multipliers (ADMM) has been studied for years. The traditional ADMM algorithm needs to compute, at each iteration, an (empirical) expected loss function on all training examples, resulting in a computational complexity proportional to the number of training examples. To reduce the ti…
New loss function equivalence reveals PER's uniform sampling can be improved.
Theory integrates loss aversion into expected utility for monetary returns.
Develops a framework for consistent loss functions with variable transformations.
The Basel II internal ratings-based (IRB) approach to capital adequacy for credit risk plays an important role in protecting the Australian banking sector against insolvency. We outline the mathematical foundations of regulatory capital for credit risk, and extend the model specification of the IRB approach to a more g…
Loss-calibrated EP improves Bayesian decision-making by focusing on utility-sensitive posterior approximations.
We empirically investigate the (negative) expected accuracy as an alternative loss function to cross entropy (negative log likelihood) for classification tasks. Coupled with softmax activation, it has small derivatives over most of its domain, and is therefore hard to optimize. A modified, leaky version is evaluated on…
A main puzzle of deep neural networks (DNNs) revolves around the apparent absence of "overfitting", defined in this paper as follows: the expected error does not get worse when increasing the number of neurons or of iterations of gradient descent. This is surprising because of the large capacity demonstrated by DNNs to…
Study shows AMM liquidity providers lose more than they earn, with varying profitability across pairs.
This paper argues against using calibration metrics for assessing posterior probabilities and proposes expected proper scoring rules instead.
Investigates conditions for risk or utility functionals to be sensitive to large losses.
Curriculum Learning - the idea of teaching by gradually exposing the learner to examples in a meaningful order, from easy to hard, has been investigated in the context of machine learning long ago. Although methods based on this concept have been empirically shown to improve performance of several learning algorithms, …
MRCs minimize worst-case expected 0-1 loss and provide performance guarantees.
In structural credit risk models, default events and the ensuing losses are both derived from the asset values at maturity. Hence it is of utmost importance to choose a distribution for these asset values which is in accordance with empirical data. At the same time, it is desirable to still preserve some analytical tra…
This work investigates square loss in overparametrized neural networks, revealing its advantages in robustness and calibration.
The predictive quality of machine learning models is typically measured in terms of their (approximate) expected prediction accuracy or the so-called Area Under the Curve (AUC). Minimizing the reciprocals of these measures are the goals of supervised learning. However, when the models are constructed by the means of em…
Domain adaptation is the supervised learning setting in which the training and test data are sampled from different distributions: training data is sampled from a source domain, whilst test data is sampled from a target domain. This paper proposes and studies an approach, called feature-level domain adaptation (FLDA), …
The paper improves sparse Gaussian processes by optimizing predictive loss.
This paper rethinks confidence calibration under covariate shifts.
Minimizing a convex risk function is the main step in many basic learning algorithms. We study protocols for convex optimization which provably leak very little about the individual data points that constitute the loss function. Specifically, we consider differentially private algorithms that operate in the local model…
Maximum drawdown, the largest cumulative loss from peak to trough, is one of the most widely used indicators of risk in the fund management industry, but one of the least developed in the context of measures of risk. We formalize drawdown risk as Conditional Expected Drawdown (CED), which is the tail mean of maximum dr…
We analyze the generalization and robustness of the batched weighted average algorithm for V-geometrically ergodic Markov data. This algorithm is a good alternative to the empirical risk minimization algorithm when the latter suffers from overfitting or when optimizing the empirical risk is hard. For the generalization…
Given a task of predicting from , a loss function , and a set of probability distributions on , what is the optimal decision rule minimizing the worst-case expected loss over ? In this paper, we address this question by introducing a generalization of the principle of maximum entropy. Applying t…
This paper provides a PAC-Bayesian bound for CVaR in machine learning.
We introduce an equilibrium asset pricing model, which we build on the relationship between a novel risk measure, the Expected Downside Risk (EDR) and the expected return. On the one hand, our proposed risk measure uses a nonparametric approach that allows us to get rid of any assumption on the distribution of returns.…
Gaptron algorithm reduces mistakes in online multiclass classification.
The paper introduces neural INGARCH models for time series of counts.