Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

103206308411 · Jun 202019922001200920172026
48 results for empirical law

This paper reformulates systemic risk measures and finds new properties and estimators.

problem Understanding and measuring systemic risk in financial networks.
method Representation of systemic risk measures in terms of univariate risk measures and quantiles determined by copulas. Empirical properties and estimators derived.
result MES is not suitable for measuring extreme risks. ES-based measures are more sensitive to power-law tails and large losses.

We study the growth dynamics of the size of manufacturing firms considering competition and normal distribution of competency. We start with the fact that all components of the system struggle with each other for growth as happened in real competitive bussiness world. The detailed quantitative agreement of the theory w…

2002-01-14abs ↗pdf ↗

Employing profits data of Japanese companies in 2002 and 2003, we identify the non-Gibrat's law which holds in the middle profits region. From the law of detailed balance in all regions, Gibrat's law in the high region and the non-Gibrat's law in the middle region, we kinematically derive the profits distribution funct…

2005-08-24abs ↗pdf ↗

By employing exhaustive lists of large firms in European countries, we show that the upper-tail of the distribution of firm size can be fitted with a power-law (Pareto-Zipf law), and that in this region the growth rate of each firm is independent of the firm's size (Gibrat's law of proportionate effect). We also find t…

2003-10-03abs ↗pdf ↗

Scaling laws in linear regression explain model performance improvements with size and data.

problem Disagreement between empirical neural scaling laws and conventional wisdom on variance error.
method Infinite dimensional linear regression setup, one-pass SGD, Gaussian prior, power-law spectrum.
result Variance error is dominated by other errors, disappearing from the bound due to SGD's implicit regularization.

This paper tightens the law of the iterated logarithm for empirical KL_inf, applicable to unbounded data.

problem Developing nonasymptotic concentration bounds for empirical KL_inf with optimal constants and rates.
method Presenting a tight law of the iterated logarithm for empirical KL_inf, applicable to unbounded data.
result A tight law of the iterated logarithm for empirical KL_inf, applicable to unbounded data.

Data pruning algorithms struggle in high compression regimes, as shown by theoretical and empirical studies.

problem Limitations of score-based data pruning algorithms in high compression regimes.
method Theoretical and empirical analysis of score-based data pruning algorithms.
result Score-based data pruning algorithms fail in high compression regimes due to 'No Free Lunch' theorems.

A new scaling law predicts optimal batch size for training models.

problem Finding the optimal batch size for training models efficiently.
method Proposed a three-term scaling law that considers model size, training data, training steps, and batch size.
result The three-term law accurately recovers the optimal batch size and can be robustly fit with fewer training runs.

Study uncovers scaling laws and spectral properties of shallow neural networks.

problem Understanding scaling laws and spectral properties of shallow neural networks.
method Leveraging connections with matrix compressed sensing and LASSO, derived a phase diagram for excess risk.
result Uncovered crossovers between scaling regimes and plateau behaviors, validated empirical observations.

We explain theoretically a curious empirical phenomenon: "Approximating a matrix by deterministically selecting a subset of its columns with the corresponding largest leverage scores results in a good low-rank matrix surrogate". To obtain provable guarantees, previous work requires randomized sampling of the columns wi…

2014-04-06abs ↗pdf ↗

The study explains transformer scaling laws using statistical and approximation theories.

problem Understanding why transformer scaling laws exist for large models trained on low-dimensional data.
method Established statistical estimation and mathematical approximation theories for transformers on low-dimensional manifolds.
result Predicted a power law between generalization error and model and data sizes, with power depending on intrinsic data dimension.

Study uses OT to simulate markets, revealing power-law returns are driven by informational effect.

problem Reproduce power-law returns in financial markets using realistic simulations.
method Constructed artificial markets, used optimal transport (OT) to measure similarity, incrementally introduced behavioral components.
result Informational effect of prices is dominant in reproducing power-law returns, and multiple components interact synergistically.

A power-law fit to the empirical inference-compute frontier in LOB prediction suggests a scaling-law-style frontier.

problem Limit order book prediction
method Using a suite of models ranging from small decision trees to neural LOB architectures
result A power-law fit to the low- and mid-compute non-MLPLOB frontier extrapolates across multiple orders of magnitude and attains R2=0.941R^2=0.941 on the excluded high-compute MLPLOB target frontier.

Based on empirical financial time-series, we show that the "silence-breaking" probability follows a super-universal power law: the probability of observing a large movement is inversely proportional to the length of the on-going low-variability period. Such a scaling law has been previously predicted theoretically [R. …

2008-12-24abs ↗pdf ↗

Empirical study finds IT project costs follow a power-law distribution, exposing risk underestimation.

problem IT project cost overruns are underestimated due to normal distribution assumptions.
method Analyzed 5,392 IT projects to examine cost overruns following a power-law distribution.
result IT project cost overruns follow a power-law distribution with a fat tail of extreme overruns.

We respond to the issues discussed by Farmer and Lillo (FL) related to our proposed approach to understanding the origin of power-law distributions in stock price fluctuations. First, we extend our previous analysis to 1000 US stocks and perform a new estimation of market impact that accounts for splitting of large ord…

2004-03-02abs ↗pdf ↗

We provide an empirical investigation aimed at uncovering the statistical properties of intricate stock trading networks based on the order flow data of a highly liquid stock (Shenzhen Development Bank) listed on Shenzhen Stock Exchange during the whole year of 2003. By reconstructing the limit order book, we can extra…

2010-03-12abs ↗pdf ↗

Controlled interacting particle systems such as the ensemble Kalman filter (EnKF) and the feedback particle filter (FPF) are numerical algorithms to approximate the solution of the nonlinear filtering problem in continuous time. The distinguishing feature of these algorithms is that the Bayesian update step is implemen…

2019-10-05abs ↗pdf ↗

We analyze the European transition economies and show that time series for most of major indices exhibit (i) power-law correlations in their values, power-law correlations in their magnitudes, and (iii) asymmetric probability distribution. We propose a stochastic model that can generate time series with all the previou…

2006-08-02abs ↗pdf ↗

Study on RL on volatility surfaces, proving no free lunch for law-seeking methods.

problem Aligning RL agents with no-arbitrage laws in volatile markets.
method Built a law manifold, defined penalties, and used a Goodhart decomposition.
result No free lunch theorem: Law-seeking RL cannot outperform baselines.

We use data on wealth of the richest persons taken from the "rich lists" provided by business magazines like Forbes to verify if upper tails of wealth distributions follow, as often claimed, a power-law behaviour. The data sets used cover the world's richest persons over 1996-2012, the richest Americans over 1988-2012,…

2013-03-31abs ↗pdf ↗

The paper examines the short-time implied volatility of additive processes and finds key parameters.

problem Characterizing the short-time implied volatility of equity markets.
method Examined pure jump exponential additive processes with power-law scaling parameters.
result The implied volatility is consistent with equity market characteristics if and only if β=1 and δ=-1/2.

This paper addresses the statistical properties of time series driven by rational bubbles a la Blanchard and Watson (1982), corresponding to multiplicative maps, whose study has recently be revived recently in physics as a mechanism of intermittent dynamics generating power law distributions. Using insights on the beha…

1999-10-08abs ↗pdf ↗

In this paper we present an analysis of power law statistics on land markets. There have been no other studies that have analyzed power law statistics on land markets up to now. We analyzed a database of the assessed value of land, which is officially monitored and made available to the public by the Ministry of Land, …

2003-02-24abs ↗pdf ↗

The paper studies neural networks with wide layers and finds a deformed semicircle law.

problem Investigating spectral distributions of neural networks in the ultra-wide regime.
method Analyzes empirical kernel matrices, proves deformed semicircle law, provides nonlinear Hanson-Wright inequality.
result Emergence of a deformed semicircle law in the ultra-wide neural network regime.
Physics of Personal Incomecond-mat.stat-mech

We report empirical studies on the personal income distribution, and clarify that the distribution pattern of the lognormal with power law tail is the universal structure. We analyze the temporal change of Pareto index and Gibrat index to investigate the change of the inequality of the income distribution. In addition …

2002-02-22abs ↗pdf ↗

Analyzes how BPE tokenisation affects corpus statistics and model entropy in transformer models.

problem Understanding how natural language properties relate to tokenisation schemes in transformer models.
method Analyzes Shannon entropy of corpora under Zipfian distribution, investigates BPE transformations, trains language models, and uses attention diagnostics.
result Transformer models trained on BPE-tokenised corpora increasingly agree with Zipfian predictions as BPE depth increases, indicating reduced local token dependencies.

New method integrates real and synthetic data to improve machine learning models.

problem Expensive or impractical collection of high-quality data limits machine learning.
method Weighted empirical risk minimization approach for integrating surrogate data.
result Integrating surrogate data can significantly reduce test error on the original distribution.

This work analyzes neural scaling laws using power-law data spectra and derives analytical expressions for generalization error.

problem Understanding how neural network performance scales with key factors like data size and model complexity.
method Statistical mechanics techniques applied to one-pass stochastic gradient descent in a student-teacher framework.
result Derivation of analytical expressions for generalization error under power-law data spectra and identification of conditions for power-law scaling.

It is now well established empirically that financial price changes are distributed according to a power law, with cubic exponent. This is a fascinating regularity, as it holds for various classes of securities, on various markets, and on various time scales. The universality of this law suggests that there must be som…

2016-12-27abs ↗pdf ↗