Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

6481,2961,9442,592 · Jun 202019922001200920172026
48 results for law of total variance

We quantify predictive uncertainty using the posterior predictive variance.

problem Quantifying uncertainty in predictive models.
method Using the law of total variance, we generate expansions for the posterior predictive variance.
result Identify the main contributors to prediction intervals and quantify term-wise uncertainty.

The paper analyzes the bias-variance tradeoff for Bregman divergences.

problem Understanding the bias-variance tradeoff for Bregman divergences.
method Analyzes the bias-variance tradeoff through operations in dual space.
result Derives several results including a generalized law of total variance and ensembling operations.

The paper introduces a method to decompose variance in twin networks for better treatment effect estimation.

problem Accurate treatment effect estimation requires reliable uncertainty measures to locate model failures.
method Layer-wise variance decomposition using Monte Carlo Dropout in twin networks.
result The encoder component dominates under distributional shift, providing a practical diagnostic for data collection.

Language models allocate information storage, not collapsing into uniform representations.

problem Incomplete neural collapse in language model representations.
method Analyzing variance and information sharing across 14 models, proving an information floor.
result Within-class variance is allocated information storage, not collapsed into uniform representations.

The law of total probability may be deployed in binary classification exercises to estimate the unconditional class probabilities if the class proportions in the training set are not representative of the population class proportions. We argue that this is not a conceptually sound approach and suggest an alternative ba…

2013-12-02abs ↗pdf ↗

Study on RL on volatility surfaces, proving no free lunch for law-seeking methods.

problem Aligning RL agents with no-arbitrage laws in volatile markets.
method Built a law manifold, defined penalties, and used a Goodhart decomposition.
result No free lunch theorem: Law-seeking RL cannot outperform baselines.

Scaling laws in linear regression explain model performance improvements with size and data.

problem Disagreement between empirical neural scaling laws and conventional wisdom on variance error.
method Infinite dimensional linear regression setup, one-pass SGD, Gaussian prior, power-law spectrum.
result Variance error is dominated by other errors, disappearing from the bound due to SGD's implicit regularization.

Ensembles improve classifier performance by reducing bias, not variance.

problem Improving classifier performance through ensemble methods.
method Extended bias-variance decomposition for classification tasks, introducing dual reparameterization.
result Ensembling reduces bias in classifiers, contrary to the traditional view.

New theory shows how multi-head attention reduces variance and decorrelates outputs.

problem Understanding and optimizing multi-head attention in neural networks.
method Developed a statistical theory linking multi-head attention to ensemble Nadaraya-Watson estimators.
result MHA variance reduction depends on head decorrelation, not just head count.

This paper investigates the use of multiple directions of stratification as a variance reduction technique for Monte Carlo simulations of path-dependent options driven by Gaussian vectors. The precision of the method depends on the choice of the directions of stratification and the allocation rule within each strata. S…

2010-04-28abs ↗pdf ↗

Before training a neural net, a classic rule of thumb is to randomly initialize the weights so the variance of activations is preserved across layers. This is traditionally interpreted using the total variance due to randomness in both weights \emph{and} samples. Alternatively, one can interpret the rule of thumb as pr…

2019-02-13abs ↗pdf ↗

For the first time ever, we analyze a unique public procurement database, which includes information about a number of bidders for a contract, a final price, an identification of a winner and an identification of a contracting authority for each of more than 40,000 public procurements in the Czech Republic between 2006…

2013-09-01abs ↗pdf ↗

We consider an ideal closed stock market, in which 100 traders have economic activities. The assets of the traders change through buying and selling stocks. We simulate the assets under conservation of both total currency and total number of stocks. If the traders are identical, then the assets are distributed as a sta…

2003-12-22abs ↗pdf ↗

New algorithm reduces variance in Monte Carlo simulations using deep neural networks and policy gradients.

problem Reducing variance in Monte Carlo simulations for estimating function values.
method Optimal correlation search using deep neural networks and policy gradients.
result Optimal correlation function reduces variance by approximating and calibrating policy.

Large batch sizes reduce gradient variance in DP-SGD, improving privacy.

problem Understanding why large batch sizes work in DP-SGD.
method Decomposed total gradient variance into subsampling and noise-induced variances, proving batch size independence in the limit.
result Large batch sizes reduce effective total gradient variance, improving privacy in DP-SGD.

We propose a general yet simple theorem describing the convergence of SGD under the arbitrary sampling paradigm. Our theorem describes the convergence of an infinite array of variants of SGD, each of which is associated with a specific probability law governing the data selection rule used to form mini-batches. This is…

2019-01-27abs ↗pdf ↗

Study simulates Variance Gamma processes for energy derivatives pricing.

problem Simulating Variance Gamma processes for accurate energy derivative pricing.
method Three-step procedure to relate self-decomposability to increments, derived from Qu et al. (2019). Exact simulation of skeleton of Variance Gamma and symmetric Variance Gamma driven Ornstein-Uhlenbeck processes.
result Exact simulation of Variance Gamma and related processes without numerical inversion.

The paper proves the law of one price in a continuous-time setting without friction.

problem Identifying conditions under which the law of one price holds in a continuous-time setting without frictions.
method Formulating a new mechanism for LOP failure and proving a novel variant of the uniform boundedness principle.
result Establishes the equivalence of the economic concept of LOP with the probabilistic property of the existence of a local $\scr{E}$-martingale state price density.

Efficiently designs experiments without integrating posterior distributions.

problem Computational inefficiency in Bayesian experimental design for PDE-based models.
method Likelihood-free approach using ANN to approximate conditional expectation.
result Significant reduction in observation model evaluations.

This note finds closed-form solutions for mean-risk portfolios using a specific type of mixture distribution.

problem Finding optimal portfolios under mean-risk criteria for general distributions.
method Using normal mean-variance mixture (NMVM) distributions, the paper derives closed-form expressions for mean-risk frontiers by optimizing a Markowitz model with adjusted return vectors.
result Closed-form solutions for mean-risk portfolios are found for return vectors following NMVM distributions.

Taylor's law of temporal fluctuation scaling, variance \sim a(a(mean)b)^b, is ubiquitous in natural and social sciences. We report for the first time convincing evidence of a solid temporal fluctuation scaling law in stock illiquidity by investigating the mean-variance relationship of the high-frequency illiquidity o…

2016-10-04abs ↗pdf ↗

The paper develops estimators for variance in graph structures using fused lasso.

problem Variance estimation in graph-structured problems.
method Developed linear time estimator for homoscedastic case and total variation regularization estimator for heteroscedastic case.
result Minimax rates and consistency for variance estimation in various graph structures.

Generalized Lotka-Volterra (GLV) models extending the (70 year old) logistic equation to stochastic systems consisting of a multitude of competing auto-catalytic components lead to power distribution laws of the (100 year old) Pareto-Zipf type. In particular, when applied to economic systems, GLV leads to power laws in…

2000-12-27abs ↗pdf ↗

Ensembles of random-feature models can't outperform a single large model.

problem Finding the optimal balance between model size and ensemble size.
method Deterministic equivalent risk estimates and scaling laws analysis.
result Ensembles of random-feature models achieve near-optimal performance only under specific conditions.

Two approaches integrate qualitative views into portfolio optimization, showing aggregation methods outperform robust optimization.

problem Incorporating qualitative views into portfolio optimization models.
method Robust optimization and order aggregation methods.
result Aggregation methods outperform robust optimization in portfolio performance analysis.

The key idea of this model is that firms are the result of an evolutionary process. Based on demand and supply considerations the evolutionary model presented here derives explicitly Gibrat's law of proportionate effects as the result of the competition between products. Applying a preferential attachment mechanism for…

2012-08-06abs ↗pdf ↗

Learning shrinks hard tail, improving inference performance.

problem Improving inference performance in neural networks.
method Latent Instance Difficulty (LID) model analyzing fine-tuning of neural networks.
result Training-dependent inference scaling, with βexteffβ_ ext{eff} growing with sample size before saturating.

We study the regular conditional law of mixed Gaussian Volterra processes under the influence of model disturbances. More precisely, we study prediction of Gaussian Volterra processes driven by a Brownian motion in a case where the Brownian motion is not observable, but only a noisy version is observed. As an applicati…

2019-04-22abs ↗pdf ↗

Using an exhaustive list of Japanese bankruptcy in 1997, we discover a Zipf law for the distribution of total liabilities of bankrupted firms in high debt range. The life-time of these bankrupted firms has exponential distribution in correlation with entry rate of new firms. We also show that the debt and size are high…

2003-10-03abs ↗pdf ↗

Study tests UK FTSE-listed companies' financial data for Benford's Law conformity.

problem Ensuring the fairness of public revenue collection and reducing tax avoidance risks.
method Utilised pre-tax income and total assets data from 567 FTSE companies, tested for Benford's Laws conformity using χ2\chi^2 and MAD tests.
result MAD test rejects Benford's Laws conformity, suggesting potential issues with reported financial data.

The study compares parametric and nonparametric models for estimating mean-variance mixtures and finds that nonparametric models perform better.

problem Estimating the distribution of a normal mean-variance mixture under uncertainty.
method Comparison of six parametric mixing laws with a grid nonparametric maximum likelihood estimator, using a paired block bootstrap for score comparison.
result Nonparametric models outperform parametric models in estimating the distribution of a normal mean-variance mixture.

Analyzes how diffusion models learn, revealing a spectral bias in structure mastery.

problem Understanding the learning dynamics and bias in diffusion models.
method Developed an analytical framework using a Gaussian-equivalence principle to solve gradient-flow dynamics and integrate probability-flow ODEs.
result Exposes a universal inverse-variance spectral law: high-variance structure is mastered faster than low-variance detail.

Proposes ENVAR for causal discovery in structural VAR models with equal noise variance.

problem Challenges in causal discovery from multivariate time series with contemporaneous effects.
method Introduces observational equivalence and the observational alignment discrepancy for structural VAR models with equal noise variance.
result Shows that multiple structural VAR parameterizations can induce the same stationary observed process law.

The ARCH process (R. F. Engle, 1982) constitutes a paradigmatic generator of stochastic time series with time-dependent variance like it appears on a wide broad of systems besides economics in which ARCH was born. Although the ARCH process captures the so-called "volatility clustering" and the asymptotic power-law prob…

2007-05-23abs ↗pdf ↗

In our simplified description `wealth' is money (mm). A kinetic theory of gas like model of money is investigated where two agents interact (trade) selectively and exchange some amount of money between them so that sum of their money is unchanged and thus total money of all the agents remains conserved. The probabilit…

2005-09-21abs ↗pdf ↗

Study shows how anisotropic data affects learning dynamics in phase retrieval.

problem Understanding learning dynamics in phase retrieval with anisotropic Gaussian inputs.
method Developed a tractable reduction to reveal a three-phase trajectory and derived scaling laws.
result Found that anisotropy leads to a three-phase trajectory: fast escape, slow convergence, and spectral-tail learning.