Obtaining accurate and well calibrated probability estimates from classifiers is useful in many applications, for example, when minimising the expected cost of classifications. Existing methods of calibrating probability estimates are applied globally, ignoring the potential for improvements by applying a more fine-gra…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Midicoth compresses online probability estimates by correcting prior smoothing biases.
The article applies Occam's Razor to non-parametric model building, minimizing the number of bits for data encoding.
The estimate of a Multiperiod probability of default applied to residential mortgages can be obtained using the mean of the observed default, so called the Mean of ratios estimator, or aggregating the default and the issued mortgages and computing the ratio of their sum, that is the Ratio of means. This work studies th…
We propose a betting strategy based on Bayesian logistic regression modeling for the probability forecasting game in the framework of game-theoretic probability by Shafer and Vovk (2001). We prove some results concerning the strong law of large numbers in the probability forecasting game with side information based on …
We study the problem of supervised learning for both binary and multiclass classification from a unified geometric perspective. In particular, we propose a geometric regularization technique to find the submanifold corresponding to a robust estimator of the class probability . The regularization term meas…
The path probability of a particle undergoing stochastic motion is studied by the use of functional technique, and the general formula is derived for the path probability distribution functional. The probability of finding paths inside a tube/band, the center of which is stipulated by a given path, is analytically eval…
We apply the formalism of the continuous time random walk to the study of financial data. The entire distribution of prices can be obtained once two auxiliary densities are known. These are the probability densities for the pausing time between successive jumps and the corresponding probability density for the magnitud…
Paper finds robust -quantiles equal to extremal distributions.
Implied posterior probability of a given model (say, Support Vector Machines (SVM)) at a point is an estimate of the class posterior probability pertaining to the class of functions of the model applied to a given dataset. It can be regarded as a score (or estimate) for the true posterior probability, which ca…
OPAA estimates probability densities using functional analysis.
This paper applies quantum probability theory to model asset returns, avoiding assumptions about quantum effects.
Paper uses stats to predict treatment choice based on illness probability.
The paper develops approximations for Pearson's chi-square statistic and applies them to confidence intervals.
A method for diffusion on probability simplex for generative models.
The concordance probability or C-index is a popular measure to capture the discriminatory ability of a regression model. In this article, the definition of this measure is adapted to the specific needs of the frequency and severity model, typically used during the technical pricing of a non-life insurance product. Due …
Categorical d-separation criterion simplifies probability graph analysis.
Predicting potential credit default accounts in advance is challenging. Traditional statistical techniques typically cannot handle large amounts of data and the dynamic nature of fraud and humans. To tackle this problem, recent research has focused on artificial and computational intelligence based approaches. In this …
Paper connects probability density cuts to graph theory eigenfunctions.
NNLMs optimize poorly for word probabilities due to embedding space structure.
Within the framework of the cumulative prospective theory of Kahneman and Tversky, this paper considers a continuous-time behavioral portfolio selection problem whose model includes both running and terminal terms in the objective functional. Despite the existence of S-shaped utility functions and probability distortio…
A key prerequisite to optimal reasoning under uncertainty in intelligent systems is to start with good class probability estimates. This paper improves on the current best probability estimation trees (Bagged-PETs) and also presents a new ensemble-based algorithm (MOB-ESP). Comparisons are made using several benchmark …
We extend Bayes' theorem for upper probabilities considering likelihood uncertainty.
While Gaussian probability densities are omnipresent in applied mathematics, Gaussian cumulative probabilities are hard to calculate in any but the univariate case. We study the utility of Expectation Propagation (EP) as an approximate integration method for this problem. For rectangular integration regions, the approx…
We analyze the data of the Italian and U.S. futures on the stock markets and we test the validity of the Continuous Time Random Walk assumption for the survival probability of the returns time series via a renewal aging experiment. We also study the survival probability of returns sign and apply a coarse graining proce…
Bayesian approach approximates probability functions of Gaussian mixtures.
Link invariants fail to detect most links with high probability.
We apply multiple testing procedures to the validation of estimated default probabilities in credit rating systems. The goal is to identify rating classes for which the probability of default is estimated inaccurately, while still maintaining a predefined level of committing type I errors as measured by the familywise …
The framework of this paper is that of risk measuring under uncertainty, which is when no reference probability measure is given. To every regular convex risk measure on , we associate a unique equivalence class of probability measures on Borel sets, characterizing the riskless non positive elements of $…
New formulae identify discrete probability laws without needing normalization constants.
We develop a general framework for applying the Kelly criterion to stock markets. By supplying an arbitrary probability distribution modeling the future price movement of a set of stocks, the Kelly fraction for investing each stock can be calculated by inverting a matrix involving only first and second moments. The fra…
MPT improves CNN and energy-based models' OOD detection and generalization.
We study two-layer belief networks of binary random variables in which the conditional probabilities Pr[childlparents] depend monotonically on weighted sums of the parents. In large networks where exact probabilistic inference is intractable, we show how to compute upper and lower bounds on many probabilities of intere…
Understanding how users navigate in a network is of high interest in many applications. We consider a setting where only aggregate node-level traffic is observed and tackle the task of learning edge transition probabilities. We cast it as a preference learning problem, and we study a model where choices follow Luce's a…
This paper uses probability tensors for efficient path planning in complex scenarios.
New method calibrates classifier probabilities with guaranteed coverage.
The authors examine the concept of probability of default for asset-backed loans. In contrast to unsecured loans it is shown that probability of default can be defined as either a measure of the likelihood of the borrower failing to make required payments, or as the likelihood of an insufficiency of collateral value on…
Volatility measures the amplitude of price fluctuations. Despite it is one of the most important quantities in finance, volatility is not directly observable. Here we apply a maximum likelihood method which assumes that price and volatility follow a two-dimensional diffusion process where volatility is the stochastic d…
The paper analyzes how contagion affects the survival probability of investment groups in microfinance.
Proposes a method to reconcile count time series forecasts.
We present a new algorithm for identifying the transition and emission probabilities of a hidden Markov model (HMM) from the emitted data. Expectation-maximization becomes computationally prohibitive for long observation records, which are often required for identification. The new algorithm is particularly suitable fo…
The paper explores statistical and topological properties of sliced probability divergences.
The restricted Boltzmann machine is a network of stochastic units with undirected interactions between pairs of visible and hidden units. This model was popularized as a building block of deep learning architectures and has continued to play an important role in applied and theoretical machine learning. Restricted Bolt…
Dynamic Vocabulary Pruning stabilizes LLM training by removing low-probability tokens.
This paper calculates the exact probability distribution of hypervolume improvement for bi-objective problems.
We describe a method to perform functional operations on probability distributions of random variables. The method uses reproducing kernel Hilbert space representations of probability distributions, and it is applicable to all operations which can be applied to points drawn from the respective distributions. We refer t…
A new machine learning method calculates failure probability efficiently and accurately.
DoSE improves OOD detection by estimating model probability density.