We quantify predictive uncertainty using the posterior predictive variance.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper analyzes the bias-variance tradeoff for Bregman divergences.
Measures three types of noise in LLM evaluations.
The paper introduces a method to decompose variance in twin networks for better treatment effect estimation.
Language models allocate information storage, not collapsing into uniform representations.
The law of total probability may be deployed in binary classification exercises to estimate the unconditional class probabilities if the class proportions in the training set are not representative of the population class proportions. We argue that this is not a conceptually sound approach and suggest an alternative ba…
Study on RL on volatility surfaces, proving no free lunch for law-seeking methods.
Theory explains neural network scaling with dataset and model size.
Scaling laws in linear regression explain model performance improvements with size and data.
We study how the presence of correlations in physical variables contributes to the form of probability distributions. We investigate a process with correlations in the variance generated by (i) a Gaussian or (ii) a truncated Lévy distribution. For both (i) and (ii), we find that due to the correlations in the variance,…
Ensembles improve classifier performance by reducing bias, not variance.
New theory shows how multi-head attention reduces variance and decorrelates outputs.
This paper investigates the use of multiple directions of stratification as a variance reduction technique for Monte Carlo simulations of path-dependent options driven by Gaussian vectors. The precision of the method depends on the choice of the directions of stratification and the allocation rule within each strata. S…
Before training a neural net, a classic rule of thumb is to randomly initialize the weights so the variance of activations is preserved across layers. This is traditionally interpreted using the total variance due to randomness in both weights \emph{and} samples. Alternatively, one can interpret the rule of thumb as pr…
Improves diffusion models by controlling total variance and signal-to-noise-ratio.
For the first time ever, we analyze a unique public procurement database, which includes information about a number of bidders for a contract, a final price, an identification of a winner and an identification of a contracting authority for each of more than 40,000 public procurements in the Czech Republic between 2006…
We consider an ideal closed stock market, in which 100 traders have economic activities. The assets of the traders change through buying and selling stocks. We simulate the assets under conservation of both total currency and total number of stocks. If the traders are identical, then the assets are distributed as a sta…
In this paper we analyze a dynamic recursive extension of the (static) notion of a deviation measure and its properties. We study distribution invariant deviation measures and show that the only dynamic deviation measure which is law invariant and recursive is the variance. We also solve the problem of optimal risk-sha…
New bounds on neural network convergence using information theory.
New algorithm reduces variance in Monte Carlo simulations using deep neural networks and policy gradients.
Large batch sizes reduce gradient variance in DP-SGD, improving privacy.
We propose a general yet simple theorem describing the convergence of SGD under the arbitrary sampling paradigm. Our theorem describes the convergence of an infinite array of variants of SGD, each of which is associated with a specific probability law governing the data selection rule used to form mini-batches. This is…
Study simulates Variance Gamma processes for energy derivatives pricing.
The paper proves the law of one price in a continuous-time setting without friction.
New Monte Carlo method outperforms existing strategy for estimating Sobol' indices.
Efficiently designs experiments without integrating posterior distributions.
This note finds closed-form solutions for mean-risk portfolios using a specific type of mixture distribution.
Taylor's law of temporal fluctuation scaling, variance mean, is ubiquitous in natural and social sciences. We report for the first time convincing evidence of a solid temporal fluctuation scaling law in stock illiquidity by investigating the mean-variance relationship of the high-frequency illiquidity o…
Pareto law, which states that wealth distribution in societies have a power-law tail, has been a subject of intensive investigations in statistical physics community. Several models have been employed to explain this behavior. However, most of the agent based models assume the conservation of number of agents and wealt…
The paper develops estimators for variance in graph structures using fused lasso.
Generalized Lotka-Volterra (GLV) models extending the (70 year old) logistic equation to stochastic systems consisting of a multitude of competing auto-catalytic components lead to power distribution laws of the (100 year old) Pareto-Zipf type. In particular, when applied to economic systems, GLV leads to power laws in…
Ensembles of random-feature models can't outperform a single large model.
Two approaches integrate qualitative views into portfolio optimization, showing aggregation methods outperform robust optimization.
To know the statistical distribution of a variable is an important problem in management of resources. Distributions of the power law type are observed in many real systems. However power law distributions have an infinite variance and thus can not be used as a standard distribution. Normally professionals in the area …
The key idea of this model is that firms are the result of an evolutionary process. Based on demand and supply considerations the evolutionary model presented here derives explicitly Gibrat's law of proportionate effects as the result of the competition between products. Applying a preferential attachment mechanism for…
Learning shrinks hard tail, improving inference performance.
We study the regular conditional law of mixed Gaussian Volterra processes under the influence of model disturbances. More precisely, we study prediction of Gaussian Volterra processes driven by a Brownian motion in a case where the Brownian motion is not observable, but only a noisy version is observed. As an applicati…
Using an exhaustive list of Japanese bankruptcy in 1997, we discover a Zipf law for the distribution of total liabilities of bankrupted firms in high debt range. The life-time of these bankrupted firms has exponential distribution in correlation with entry rate of new firms. We also show that the debt and size are high…
Study tests UK FTSE-listed companies' financial data for Benford's Law conformity.
We undertake a systematic comparison between implied volatility, as represented by VIX (new methodology) and VXO (old methodology), and realized volatility. We compare visually and statistically distributions of realized and implied variance (volatility squared) and study the distribution of their ratio. We find that t…
The study compares parametric and nonparametric models for estimating mean-variance mixtures and finds that nonparametric models perform better.
Analyzes how diffusion models learn, revealing a spectral bias in structure mastery.
Employing profits data of Japanese firms in 2003--2005, we kinematically exhibit the static log-normal distribution in the middle scale region. In the derivation, a Non-Gibrat's law under the detailed balance is adopted together with following two approximations. Firstly, the probability density function of profits gro…
We study the probability distribution of stock returns at mesoscopic time lags (return horizons) ranging from about an hour to about a month. While at shorter microscopic time lags the distribution has power-law tails, for mesoscopic times the bulk of the distribution (more than 99% of the probability) follows an expon…
Proposes ENVAR for causal discovery in structural VAR models with equal noise variance.
The ARCH process (R. F. Engle, 1982) constitutes a paradigmatic generator of stochastic time series with time-dependent variance like it appears on a wide broad of systems besides economics in which ARCH was born. Although the ARCH process captures the so-called "volatility clustering" and the asymptotic power-law prob…
In our simplified description `wealth' is money (). A kinetic theory of gas like model of money is investigated where two agents interact (trade) selectively and exchange some amount of money between them so that sum of their money is unchanged and thus total money of all the agents remains conserved. The probabilit…
Study shows how anisotropic data affects learning dynamics in phase retrieval.