In this work we study loss functions for learning and evaluating probability distributions over large discrete domains. Unlike classification or regression where a wide variety of loss functions are used, in the distribution learning and density estimation literature, very few losses outside the dominant ar…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This work explores alternative learning criteria beyond traditional risk.
Paper introduces a new topological loss for better convergence.
Estimation of importance sampling weights for off-policy evaluation of contextual bandits often results in imbalance - a mismatch between the desired and the actual distribution of state-action pairs after weighting. In this work we present balanced off-policy evaluation (B-OPE), a generic method for estimating weights…
Proposes a method to improve classification robustness against label noise.
Generative models learn complex data from low-dimensional manifolds.
In structural credit risk models, default events and the ensuing losses are both derived from the asset values at maturity. Hence it is of utmost importance to choose a distribution for these asset values which is in accordance with empirical data. At the same time, it is desirable to still preserve some analytical tra…
Operational risk models commonly employ maximum likelihood estimation (MLE) to fit loss data to heavy-tailed distributions. Yet several desirable properties of MLE (e.g. asymptotic normality) are generally valid only for large sample-sizes, a situation rarely encountered in operational risk. In this paper, we study how…
Significant advances have been made recently on training neural networks, where the main challenge is in solving an optimization problem with abundant critical points. However, existing approaches to address this issue crucially rely on a restrictive assumption: the training data is drawn from a Gaussian distribution. …
Generative models often misrepresent class frequencies; this paper calibrates them.
Neural Processes (NPs) are a class of models that learn a mapping from a context set of input-output pairs to a distribution over functions. They are traditionally trained using maximum likelihood with a KL divergence regularization term. We show that there are desirable classes of problems where NPs, with this loss, f…
The paper proposes a new method to measure risk with fine-grained tail sensitivity.
Wasserstein GANs fail to approximate Wasserstein distance, leading to their success.
Estimates conditional distribution function using neural networks for censored and uncensored data.
This paper examines how different decoding algorithms for LLMs align with various goals.
Generative models learn distributions, new method finds inputs matching desired conditional distributions.
Fewer degrees of freedom can train deep networks, showing a sharp phase transition.
Proposes a new Huber loss combining absolute and quadratic properties.
Proposes a new loss function for robust learning.
In the UK betting market, bookmakers often offer a free coupon to new customers. These free coupons allow the customer to place extra bets, at lower risk, in combination with the usual betting odds. We are interested in whether a customer can exploit these free coupons in order to make a sure gain, and if so, how the c…
We present -loss, , a tunable loss function for binary classification that bridges log-loss () and - loss (). We prove that -loss has an equivalent margin-based form and is classification-calibrated, two desirable properties for a good surrogate loss function for the ideal y…
UAMM uses external market prices to improve AMM efficiency and reduce liquidity provider risk.
Hi-fi priors enhance BNNs by learning flexible activations.
Paper introduces privacy-preserving inventory policy learning for feature-based newsvendor with unknown demand.
Paper extends FFT-based differential privacy method to heterogeneous compositions.
Paper proposes a method to efficiently cluster stretched mixtures.
Sharpe et al. proposed the idea of having an expected utility maximizer choose a probability distribution for future wealth as an input to her investment problem instead of a utility function. They developed a computer program, called The Distribution Builder, as one way to elicit such a distribution. In a single-perio…
Max-margin learning is a powerful approach to building classifiers and structured output predictors. Recent work on max-margin supervised topic models has successfully integrated it with Bayesian topic models to discover discriminative latent semantic structures and make accurate predictions for unseen testing data. Ho…
Selective removal of data subsets can efficiently unlearn unwanted distributions.
Heavy Lasso improves robustness in high-dimensional linear regression with heavy-tailed errors.
The gain-loss ratio is known to enjoy very good properties from a normative point of view. As a confirmation, we show that the best market gain-loss ratio in the presence of a random endowment is an acceptability index and we provide its dual representation for general probability spaces. However, the gain-loss ratio w…
ANGLE tackles circular data regression, improving predictive performance.
Model calculates capital requirements for multi-line insurance companies.
Proposes a deep ordinal regression framework using optimal transport loss and unimodal output probabilities.
Cross-entropy loss together with softmax is arguably one of the most common used supervision components in convolutional neural networks (CNNs). Despite its simplicity, popularity and excellent performance, the component does not explicitly encourage discriminative learning of features. In this paper, we propose a gene…
In high-dimensional classification settings, we wish to seek a balance between high power and ensuring control over a desired loss function. In many settings, the points most likely to be misclassified are those who lie near the decision boundary of the given classification method. Often, these uninformative points sho…
Support vector machines (SVMs) are special kernel based methods and belong to the most successful learning methods since more than a decade. SVMs can informally be described as a kind of regularized M-estimators for functions and have demonstrated their usefulness in many complicated real-life problems. During the last…
Cyclical MCMC tackles high-dimensional multimodal distributions, showing convergence under certain conditions.
While optimizing convex objective (loss) functions has been a powerhouse for machine learning for at least two decades, non-convex loss functions have attracted fast growing interests recently, due to many desirable properties such as superior robustness and classification accuracy, compared with their convex counterpa…
SEMF predicts prediction intervals for ML models using latent variables.
This paper extends performative prediction to nonlinear cases.
Paper proposes new loss functions for training energy networks.
Stochastic GD converges linearly for CV@R learning under certain conditions.
Paper proposes a method to design molecules with specific properties.
Data containing human or social attributes may over- or under-represent groups with respect to salient social attributes such as gender or race, which can lead to biases in downstream applications. This paper presents an algorithmic framework that can be used as a data preprocessing method towards mitigating such bias.…
We propose a robust inferential procedure for assessing uncertainties of parameter estimation in high-dimensional linear models, where the dimension can grow exponentially fast with the sample size . Our method combines the de-biasing technique with the composite quantile function to construct an estimator that …
Study improves generalization bounds for linear regression across tasks.
We present a fully-supervized method for learning to segment data structured by an adjacency graph. We introduce the graph-structured contrastive loss, a loss function structured by a ground truth segmentation. It promotes learning vertex embeddings which are homogeneous within desired segments, and have high contrast …