Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

74148222296 · Jun 202019922001200920172026
48 results for scaling law crossover

New principles needed for scaling large language models, challenging traditional regularization methods.

problem The shift from generalization to scaling in machine learning requires new guiding principles.
method Examining the effectiveness of traditional regularization methods in the scaling-centric era.
result Traditional principles of regularization may not generalize to larger scales, highlighting new phenomena like scaling law crossover.

Study uncovers scaling laws and spectral properties of shallow neural networks.

problem Understanding scaling laws and spectral properties of shallow neural networks.
method Leveraging connections with matrix compressed sensing and LASSO, derived a phase diagram for excess risk.
result Uncovered crossovers between scaling regimes and plateau behaviors, validated empirical observations.

We propose a simple stochastic volatility model which is analytically tractable, very easy to simulate and which captures some relevant stylized facts of financial assets, including scaling properties. In particular, the model displays a crossover in the log-return distribution from power-law tails (small time) to a Ga…

2010-06-01abs ↗pdf ↗

We investigate the waiting-time distribution of the absolute return in the Korean stock-market index KOSPI. We define the waiting time as a time interval during which the normalized absolute return remains continuously below a threshold rcr_c. Through an exponential bin plot, we observe that the waiting-time distributi…

2005-08-30abs ↗pdf ↗

Based on the tick-by-tick price changes of the companies from the U.S. and from the German stock markets over the period 1998-99 we reanalyse several characteristics established by the Boston Group for the U.S. market in the period 1994-95, which serves to verify their space and time-translational invariance. By increa…

2002-08-12abs ↗pdf ↗

Using a large database of 8 million institutional trades executed in the U.S. equity market, we establish a clear crossover between a linear market impact regime and a square-root regime as a function of the volume of the order. Our empirical results are remarkably well explained by a recently proposed dynamical theory…

2018-11-13abs ↗pdf ↗

The detrending moving average (DMA) algorithm is one of the best performing methods to quantify the long-term correlations in nonstationary time series. Many long-term correlated time series in real systems contain various trends. We investigate the effects of polynomial trends on the scaling behaviors and the performa…

2015-04-28abs ↗pdf ↗

Unified thermodynamic approach to Transformer attention dynamics.

problem Understanding the statistical mechanics of Transformer attention.
method Constructing a Lagrangian on the information manifold to analyze attention dynamics.
result Establishes a formal correspondence between scaled dot-product attention and canonical ensemble statistics.

We analyze the sequence of time intervals between consecutive stock trades of thirty companies representing eight sectors of the U. S. economy over a period of four years. For all companies we find that: (i) the probability density function of intertrade times may be fit by a Weibull distribution; (ii) when appropriate…

2004-03-27abs ↗pdf ↗

Dropout schedules can be optimized to significantly reduce model test loss.

problem Improving model performance in neural networks.
method Developed a mean-field theory of dropout at the edge of chaos, proposing front-loaded dropout schedules.
result Front-loaded dropout schedules reduce test loss by 18-35% over constant dropout.

The herd behavior of returns is investigated in Korean futures exchange market. It is obtained that the probability distribution of returns for three types of herding parameter scales as a power law RβR^{-β} with the exponents β=3.6 β=3.6(KTB203) and 2.9(KTB209) in two kinds of Korean treasury bond. For our case since the…

2003-04-07abs ↗pdf ↗

We study the statistical properties of volatility---a measure of how much the market is likely to fluctuate. We estimate the volatility by the local average of the absolute price changes. We analyze (a) the S&P 500 stock index for the 13-year period Jan 1984 to Dec 1996 and (b) the market capitalizations of the largest…

1999-03-24abs ↗pdf ↗

An exclusion particle model is considered as a highly simplified model of a limit order market. Its price behavior reproduces the well known crossover from over-diffusion (Hurst exponent H>1/2) to diffusion (H=1/2) when the time horizon is increased, provided that orders are allowed to be canceled. For early times a ma…

2002-06-24abs ↗pdf ↗

Optimal income crossover found using particle swarm optimization.

problem Determining the crossover point between two income distributions.
method Particle swarm optimization for finding crossover income, temperature, and Pareto index.
result Optimization method finds boundaries of two income distributions.

Several models of stock trading [P. Bak et al, Physica A {\bf 246}, 430 (1997)] are analyzed in analogy with one-dimensional, two-species reaction-diffusion-branching processes. Using heuristic and scaling arguments, we show that the short-time market price variation is subdiffusive with a Hurst exponent H=1/4H=1/4. Biase…

1998-11-09abs ↗pdf ↗

We study the tick dynamical behavior of the yen-dollar exchange rate using the rescaled range analysis in financial market. It is found that the multifractal Hurst exponents with the short and long-run memory effects can be obtained from the yen-dollar exchange rate. This exists one crossover for the Hurst exponents at…

2004-05-09abs ↗pdf ↗

The intraday pattern, long memory, and multifractal nature of the intertrade durations, which are defined as the waiting times between two consecutive transactions, are investigated based upon the limit order book data and order flows of 23 liquid Chinese stocks listed on the Shenzhen Stock Exchange in 2003. An inverse…

2008-06-15abs ↗pdf ↗

The study examines Kernel Ridge Regression error rates across noiseless and noisy conditions.

problem Characterizing Kernel Ridge Regression error rates in different noise levels.
method Unified analysis of Kernel Ridge Regression under various noise and regularization conditions.
result A crossover from noiseless to noisy error rates is observed as sample complexity increases.

The paper analyzes return distribution of Chinese stock market indices over various time scales.

problem Understanding return distribution properties of Chinese stock markets.
method Systematic analysis of 1-min to 4000-min composite index datasets from 2005-2021.
result Return distribution properties are similar to mature markets, with distinct behavior at different time scales.

We investigate the herd behavior of returns for the yen-dollar exchange rate in the Japanese financial market. It is obtained that the probability distribution P(R)P(R) of returns RR satisfies the power-law behavior P(R)RβP(R) \simeq R^{-β} with the exponents β=3.11 β=3.11(the time interval τ=τ= one minute) and 3.36(τ=τ= one da…

2004-05-09abs ↗pdf ↗

The correlation function of a financial index of the New York stock exchange, the S&P 500, is analyzed at 1 min intervals over the 13-year period, Jan 84 -- Dec 96. We quantify the correlations of the absolute values of the index increment. We find that these correlations can be described by two different power laws wi…

1997-06-03abs ↗pdf ↗

A new scaling law predicts optimal batch size for training models.

problem Finding the optimal batch size for training models efficiently.
method Proposed a three-term scaling law that considers model size, training data, training steps, and batch size.
result The three-term law accurately recovers the optimal batch size and can be robustly fit with fewer training runs.

This work extends the scaling law to multiple and kernel regression, challenging traditional machine learning principles.

problem Challenging traditional machine learning wisdom with scaling law in large practical models.
method Demonstrates the scaling law in multiple and kernel regression settings.
result The scaling law extends to multiple and kernel regression, providing deeper insights into LLMs.

Study reveals neural scaling laws in random graphs and natural language models.

problem Understanding the origin of neural scaling laws in complex systems.
method Examined scaling laws in transformers trained on random walks and simplified natural language models.
result Neural scaling laws emerge in the absence of power law structure in data correlations.

This work analyzes neural scaling laws using power-law data spectra and derives analytical expressions for generalization error.

problem Understanding how neural network performance scales with key factors like data size and model complexity.
method Statistical mechanics techniques applied to one-pass stochastic gradient descent in a student-teacher framework.
result Derivation of analytical expressions for generalization error under power-law data spectra and identification of conditions for power-law scaling.

Unified theory for neural scaling laws in hierarchically compositional data.

problem Understanding neural scaling laws in hierarchically compositional data.
method Probabilistic context-free grammars and power-law distributed production rules.
result Unified learning curve behavior for classification and next-token prediction tasks.

Scaling laws in linear regression explain model performance improvements with size and data.

problem Disagreement between empirical neural scaling laws and conventional wisdom on variance error.
method Infinite dimensional linear regression setup, one-pass SGD, Gaussian prior, power-law spectrum.
result Variance error is dominated by other errors, disappearing from the bound due to SGD's implicit regularization.

Employing profits data of Japanese companies in 2002 and 2003, we confirm that Pareto's law and the Pareto index are derived from the law of detailed balance and Gibrat's law. The last two laws are observed beyond the region where Pareto's law holds. By classifying companies into job categories, we find that companies …

2005-06-08abs ↗pdf ↗

The herd behaviors of returns for the won-dollar exchange rate and the KOSPI are analyzed in Korean financial markets. It is shown that the probability distribution P(R)P(R) of price returns RR for three values of the herding parameter tends to a power-law behavior P(R)RβP(R) \simeq R^{-β} with the exponents β=2.2 β=2.2(the wo…

2003-04-21abs ↗pdf ↗

The study explains transformer scaling laws using statistical and approximation theories.

problem Understanding why transformer scaling laws exist for large models trained on low-dimensional data.
method Established statistical estimation and mathematical approximation theories for transformers on low-dimensional manifolds.
result Predicted a power law between generalization error and model and data sizes, with power depending on intrinsic data dimension.

New neural scaling law found for simple quadratic function.

problem Neural scaling laws and their predictions for model performance.
method Analysis of neural networks, lottery ticket ensembling, statistical interpretation.
result Found a new scaling law (α=1α=1) for a simple quadratic function, contradicting previous theories.

Improved scaling laws in linear regression using data reuse.

problem Sustainability of neural scaling laws when running out of new data.
method Data reuse in multi-pass stochastic gradient descent (multi-pass SGD) for MM-dimensional linear models trained on NN data with sketched features.
result Multi-pass SGD achieves a test error of Θ(M1b+L(1b)/a)Θ(M^{1-b} + L^{(1-b)/a}) with L>NL>N, improving scaling laws in data-constrained regimes.