Interactive book integrates probability models with big data analytics.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper improves MMD estimation for analytical mean embeddings.
Due to the success of deep learning to solving a variety of challenging machine learning tasks, there is a rising interest in understanding loss functions for training neural networks from a theoretical aspect. Particularly, the properties of critical points and the landscape around them are of importance to determine …
Analytical method finds deeper optima in two-layer ReLU networks.
In this paper we model the loss function of high-dimensional optimization problems by a Gaussian random field, or equivalently a Gaussian process. Our aim is to study gradient descent in such loss functions or energy landscapes and compare it to results obtained from real high-dimensional optimization problems such as …
Paper establishes a formula linking model performance to insurance loss ratio.
The stability of the financial system is associated with systemic risk factors such as the concurrent default of numerous small obligors. Hence it is of utmost importance to study the mutual dependence of losses for different creditors in the case of large, overlapping credit portfolios. We analytically calculate the m…
Edgeworth Accountant calculates privacy loss under differential privacy compositions efficiently.
Analyzes double descent in binary classification models with different losses.
Under the Basel II standards, the Operational Risk (OpRisk) advanced measurement approach is not prescriptive regarding the class of statistical model utilised to undertake capital estimation. It has however become well accepted to utlise a Loss Distributional Approach (LDA) paradigm to model the individual OpRisk loss…
This work analyzes impermanent loss in decentralized markets and provides a hedging strategy.
The impact of a stress scenario of default events on the loss distribution of a credit portfolio can be assessed by determining the loss distribution conditional on these events. While it is conceptually easy to estimate loss distributions conditional on default events by means of Monte Carlo simulation, it becomes imp…
Study evaluates interpretability of time series foundation models' latent spaces.
Most of the banks' operational risk internal models are based on loss pooling in risk and business line categories. The parameters and outputs of operational risk models are sensitive to the pooling of the data and the choice of the risk classification. In a simple model, we establish the link between the number of ris…
New method recovers signals from saturated data using linear loss and nonconvex penalties.
Study how data structure impacts classification performance in overparametrized models.
Proposes a model to update industrial data predictions based on temporal changes.
New loss functions reveal layer roles in deep neural networks.
Max-margin learning is a powerful approach to building classifiers and structured output predictors. Recent work on max-margin supervised topic models has successfully integrated it with Bayesian topic models to discover discriminative latent semantic structures and make accurate predictions for unseen testing data. Ho…
The Soft Nearest Neighbor Loss improves representation quality and generalization.
Proposes efficient Bayesian logistic regression for large sparse datasets.
Data symmetries in neural networks can generate conserved quantities.
In this paper, we address the aggregation of dependent stop loss reinsurance risks where the dependence among the ceding insurer(s) risks is governed by the Sarmanov distribution and each individual risk belongs to the class of Erlang mixtures. We investigate the effects of the ceding insurer(s) risk dependencies on th…
Supervised learning is an active research area, with numerous applications in diverse fields such as data analytics, computer vision, speech and audio processing, and image understanding. In most cases, the loss functions used in machine learning assume symmetric noise models, and seek to estimate the unknown function …
This work considers the problem of binary classification: given training data from a certain population, together with associated labels , determine the best label for an element not among the training data. More specifically, this work considers a variant o…
The article provides formulas to hedge impermanent loss in decentralized markets.
OccamNet finds interpretable symbolic fits to data efficiently.
In structural credit risk models, default events and the ensuing losses are both derived from the asset values at maturity. Hence it is of utmost importance to choose a distribution for these asset values which is in accordance with empirical data. At the same time, it is desirable to still preserve some analytical tra…
Unified framework for binary responses using AUC loss and low-rank constraint.
In this paper we prove local analytic hypoellipticity for a degenerate sum of squares of complex vector fields generalizing those of Kohn in "Hypoellipticity and Loss of Derivatives". Kohn's article is to appear in the Annals of Mathematics with an appendix by Derridj and Tartakoff proving local analyticity in that cas…
Analyzes impermanent loss in decentralized exchanges and provides a replication formula.
In this paper we develop a statistical arbitrage trading strategy with two key elements in hi-frequency trading: stop-loss and leverage. We consider, as in Bertram (2009), a mean-reverting process for the security price with proportional transaction costs; we show how to introduce stop-loss and leverage in an optimal t…
This paper analyzes the landscape of supervised contrastive loss in over-parameterized networks.
Enhances deep learning robustness to noise without sacrificing clean data accuracy.
Bayesian PINNs optimize loss weights for PDEs and data.
Domain adaptation is the supervised learning setting in which the training and test data are sampled from different distributions: training data is sampled from a source domain, whilst test data is sampled from a target domain. This paper proposes and studies an approach, called feature-level domain adaptation (FLDA), …
We address the problem of detecting changes in multivariate datastreams, and we investigate the intrinsic difficulty that change-detection methods have to face when the data dimension scales. In particular, we consider a general approach where changes are detected by comparing the distribution of the log-likelihood of …
New loss function helps models avoid noisy labels, improving robustness.
The paper examines how loss aversion impacts multi-armed bandit decisions over long periods.
Improved neural network predicts spectral functions more accurately than traditional methods.
DiMS sampler explores neural network loss minima via dissipative dynamics.
This paper finds ReLU restores symmetry in SCL under class imbalances.
Under the Basel II standards, the Operational Risk (OpRisk) advanced measurement approach allows a provision for reduction of capital as a result of insurance mitigation of up to 20%. This paper studies the behaviour of different insurance policies in the context of capital reduction for a range of possible extreme los…
New mechanisms improve differential privacy for scalar queries.
A new method for efficient label retrieval in large output spaces.
A framework connects VAEs to GLMs for better model initialization and performance.
The paper examines statistical properties of IL and LVR in automated market makers.
Motivated by the industry practice of pairs trading, we study the optimal timing strategies for trading a mean-reverting price spread. An optimal double stopping problem is formulated to analyze the timing to start and subsequently liquidate the position subject to transaction costs. Modeling the price spread by an Orn…