Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

142283425566 · Jun 202019922001200920172026
48 results for statistical limit laws

New neural scaling law found for simple quadratic function.

problem Neural scaling laws and their predictions for model performance.
method Analysis of neural networks, lottery ticket ensembling, statistical interpretation.
result Found a new scaling law (α=1α=1) for a simple quadratic function, contradicting previous theories.

This paper tightens the law of the iterated logarithm for empirical KL_inf, applicable to unbounded data.

problem Developing nonasymptotic concentration bounds for empirical KL_inf with optimal constants and rates.
method Presenting a tight law of the iterated logarithm for empirical KL_inf, applicable to unbounded data.
result A tight law of the iterated logarithm for empirical KL_inf, applicable to unbounded data.

Testing independence is of significant interest in many important areas of large-scale inference. Using extreme-value form statistics to test against sparse alternatives and using quadratic form statistics to test against dense alternatives are two important testing procedures for high-dimensional independence. However…

2015-12-30abs ↗pdf ↗

Study compares exponential and power-law kernels in modeling high-frequency trading data.

problem Modeling high-frequency trading data with specific kernel types.
method Proposes and analyzes two bivariate Hawkes processes with exponential and power-law kernels.
result Identifies strengths and limitations of exponential and power-law kernels for high-frequency trading data.

This work analyzes neural scaling laws using power-law data spectra and derives analytical expressions for generalization error.

problem Understanding how neural network performance scales with key factors like data size and model complexity.
method Statistical mechanics techniques applied to one-pass stochastic gradient descent in a student-teacher framework.
result Derivation of analytical expressions for generalization error under power-law data spectra and identification of conditions for power-law scaling.

Simplified kernel ridge regression with a conservation law.

problem Understanding the test risk and generalization of kernel ridge regression.
method Identification of a conservation law that limits KRR's learning ability, leading to simplified expressions for test risk.
result Transparency in test risk expressions through the conserved quantity in the kernel eigenbasis.

Signals consisting of a sequence of pulses show that inherent origin of the 1/f noise is a Brownian fluctuation of the average interevent time between subsequent pulses of the pulse sequence. In this paper we generalize the model of interevent time to reproduce a variety of self-affine time series exhibiting power spec…

2003-03-05abs ↗pdf ↗

The relevance of data quantifies learning efficiency.

problem Understanding the statistical nature of high-dimensional, sparse data.
method Defining relevance as information content, and using it to define ideal limits of samples and learning machines.
result Maximally informative samples and optimal learning machines exhibit critical features like power-law frequency distributions and anomalously large susceptibility.

Deviation inequalities and limit laws for random walks on metric spaces.

problem Understanding random walks on metric spaces with contracting isometries.
method Adapting Gouëzel's pivotal time construction to establish deviation inequalities.
result Exponential bounds and limit laws for random walks on mapping class groups and CAT(0) spaces.

Unified model explains market dynamics, linking order flow, volatility, and impact.

problem Understanding the dynamics of order flow, market impact, and volatility in financial markets.
method Proposes a microstructural model using Hawkes processes to distinguish core orders and reaction flow, and analyzes their scaling limits.
result Estimates the persistence parameter H0H_0 and finds it consistent with market impact and volatility properties.

NN-Turb generates turbulent velocity statistics using neural networks.

problem Creating a 1D field with turbulent velocity statistics.
method Fully-convolutional neural network (NN-Turb) to generate the field.
result NN-Turb generates a 1D field that satisfies Kolmogorov's 2/3 and 4/5 laws, exhibiting intermittency.

The study explains transformer scaling laws using statistical and approximation theories.

problem Understanding why transformer scaling laws exist for large models trained on low-dimensional data.
method Established statistical estimation and mathematical approximation theories for transformers on low-dimensional manifolds.
result Predicted a power law between generalization error and model and data sizes, with power depending on intrinsic data dimension.

New insights into how depth and width affect in-context learning in deep models.

problem Understanding how various resources impact in-context learning in deep models.
method Analyzed linear regression in a deep linear self-attention model, varying resources like depth, width, context length, and training steps.
result Increasing depth improves in-context learning even at infinite context length, contrary to previous findings.

We provide an empirical investigation aimed at uncovering the statistical properties of intricate stock trading networks based on the order flow data of a highly liquid stock (Shenzhen Development Bank) listed on Shenzhen Stock Exchange during the whole year of 2003. By reconstructing the limit order book, we can extra…

2010-03-12abs ↗pdf ↗

Following the work of Okuyama, Takayasu and Takayasu [Okuyama, Takayasu and Takayasu 1999] we analyze huge databases of Japanese companies' financial figures and confirm that the Zipf's law, a power law distribution with the exponent -1, has been maintained over 30 years in the income distribution of Japanese companies…

2003-08-19abs ↗pdf ↗

This study reveals statistical patterns in ERC20 token transactions on Ethereum blockchain.

problem Understanding transactional dynamics in decentralized systems.
method Examined over 44 million ERC20 token transfers, categorized by address type (EOA or SC), and analyzed using scaling laws.
result EOA-driven transactions exhibit consistent statistical behavior, while SC-driven activity displays sublinear scaling and bursty activity.

The paper studies scaling laws for associative memory mechanisms.

problem Understanding and optimizing learning and memorization processes.
method High-dimensional matrices of outer products of embeddings, relating to transformer models. Derived scaling laws with sample and parameter sizes. Extensive numerical experiments.
result Precise scaling laws and statistical efficiency of estimators.

The paper proves local laws for non-separable sample covariance matrices.

problem Analyzing non-separable sample covariance matrices with dependent or nonlinearly transformed data.
method Tensor network framework for analyzing fluctuation averaging in the presence of higher-order cumulant structure.
result Optimal averaged local law and full anisotropic local law for non-separable sample covariance matrices.

We uncover scaling laws and statistical structure in complex datasets.

problem Understanding universal traits in complex datasets.
method Analogizing data to physical systems, using statistical physics and RMT.
result Real-world datasets and Gaussian data with long-range correlations share the same RMT universality class.

Ridge regression reveals surprising high-dimensional behaviors via random matrix theory.

problem Understanding power-law scalings in high-dimensional regression models.
method Random matrix theory and free probability.
result Analytic formulas for training and generalization errors derived from SS-transform.

New formulae identify discrete probability laws without needing normalization constants.

problem Characterizing non-normalized discrete probability distributions.
method Derive explicit formulae for mass functions using Stein's method.
result Developed tools for solving statistical problems without normalization constants.

We study the volume distribution of nodal domains of random band-limited functions on generic manifolds, and find that in the high energy limit a typical instance obeys a deterministic universal law, independent of the manifold. Some of the basic qualitative properties of this law, such as its support, monotonicity and…

2016-06-18abs ↗pdf ↗

New method uses sufficient statistics to infer causal relationships from observational data.

problem Inferring causal relationships from observational data with hidden variables.
method Information Bottleneck method applied to find functional sufficient statistics.
result New causal rules not obtainable from standard methods, validated on simulated and real data.

This paper improves the robustness of risk estimation for financial positions.

problem Ensuring robustness of risk measures in the presence of data noise.
method Proposes a quantitative approach using the Fortet-Mourier metric to quantify the variation of true probability measures.
result Derives explicit error bounds for discrepancies between laws of estimators based on true and perturbed data.

In this paper, we quantitatively investigate the statistical properties of a statistical ensemble of stock prices. We selected 1200 stocks traded on the Tokyo Stock Exchange, and formed a statistical ensemble of daily stock prices for each trading day in the 3-year period from January 4, 1999 to December 28, 2001, corr…

2006-03-17abs ↗pdf ↗

Paper studies deep learning for solving elliptic PDEs, proving optimal bounds and neural scaling laws.

problem Solving elliptic PDEs from random samples using machine learning.
method Deep Ritz Method and Physics-Informed Neural Networks (PINNs) for the Schrödinger equation.
result Proves minimax optimal bounds and neural scaling laws for deep PDE solvers.

Study shows how anisotropic data affects learning dynamics in phase retrieval.

problem Understanding learning dynamics in phase retrieval with anisotropic Gaussian inputs.
method Developed a tractable reduction to reveal a three-phase trajectory and derived scaling laws.
result Found that anisotropy leads to a three-phase trajectory: fast escape, slow convergence, and spectral-tail learning.

In the present work we demonstrate the application of different physical methods to high-frequency or tick-by-tick financial time series data. In particular, we calculate the Hurst exponent and inverse statistics for the price time series taken from a range of futures indices. Additionally, we show that in a limit orde…

2007-12-18abs ↗pdf ↗