Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

71141212282 · Jun 202019922001200920172026
48 results for scaling exponent

The state of a stochastic process evolving over a time tt is typically assumed to lie on a normal distribution whose width scales like t1/2t^{1/2}. However, processes where the probability distribution is not normal and the scaling exponent differs from 12\frac{1}{2} are known. The search for possible origins of such "a…

2017-04-07abs ↗pdf ↗

Optimizer choice affects neural scaling laws, changing the exponent α\alpha.

problem The exponent α\alpha in neural scaling laws L(N)NαL(N) \propto N^{-\alpha} varies with the optimizer used.
method Controlled random-feature regression experiments with five optimizer variants and six spectral conditions.
result Preconditioned optimizers yield steeper scaling (larger α\alpha), with the α\alpha-shift increasing across most of the tested spectral range.

This paper investigates the scaling dependencies between measures of "activity" and of "size" for companies included in the FTSE 100. The "size" of companies is measured by the total market capitalization. The "activity" is measured with several quantities related to trades (transaction value per trade, transaction val…

2004-07-29abs ↗pdf ↗

We consider strictly stationary heavy tailed time series whose finite-dimensional exponent measures are concentrated on axes, and hence their extremal properties cannot be tackled using classical multivariate regular variation that is suitable for time series with extremal dependence. We recover relevant information ab…

2013-07-05abs ↗pdf ↗

Improved loss scaling for stochastic momentum algorithms in high dimensions.

problem Improving loss scaling for stochastic momentum algorithms in high dimensions.
method Dimension-adapted Nesterov acceleration (DANA) scales momentum hyperparameters based on model size and data complexity.
result DANA improves loss scaling exponents across various data and target complexities.

Model shows feature learning can improve neural scaling laws for hard tasks.

problem Understanding and improving neural network scaling laws for various task difficulties.
method Developed a solvable model of neural scaling laws, identified three scaling regimes, and demonstrated feature learning's impact on scaling exponents.
result Feature learning can improve scaling with training time and compute for hard tasks, nearly doubling the exponent.

Are large scale research programs that include many projects more productive than smaller ones with fewer projects? This problem of economy of scale is particularly relevant for understanding recent mergers in particular in the pharmaceutical industry. We present a quantitative theory based on the characterization of d…

2000-01-31abs ↗pdf ↗

This paper improves traditional Markowitz optimization by considering variance at multiple time scales.

problem Traditional Markowitz optimization limits to a single time scale, ignoring variance across different frequencies.
method Introduces multifrequency optimization allowing specification of target Hurst exponents across multiple time scales.
result Effective risk management strategy that aligns with investor preferences at various time scales.

Taylor's law of temporal fluctuation scaling, variance \sim a(a(mean)b)^b, is ubiquitous in natural and social sciences. We report for the first time convincing evidence of a solid temporal fluctuation scaling law in stock illiquidity by investigating the mean-variance relationship of the high-frequency illiquidity o…

2016-10-04abs ↗pdf ↗

The paper studies entropy calibration in language models and finds that miscalibration improves slowly with scale.

problem The problem is whether language model entropy calibration improves with scale and if it's possible to calibrate without reducing log loss.
method The authors study a simplified theoretical setting to characterize miscalibration scaling behavior and measure it empirically in language models ranging from 0.5B to 70B parameters.
result The observed scaling behavior of miscalibration is similar to theoretical predictions, indicating slow improvement with scale. The authors also prove theoretically that it is possible to reduce entropy while preserving log loss if access to a black box predicting future entropy is available.

We analyze a database comprising quarterly sales of 55624 pharmaceutical products commercialized by 3939 pharmaceutical firms in the period 1992--2001. We study the probability density function (PDF) of growth in firms and product sales and find that the width of the PDF of growth decays with the sales as a power law w…

2005-02-15abs ↗pdf ↗

Study tests rough fractional volatility model across different time scales, revealing new volatility patterns.

problem Testing robustness of rough fractional volatility model over various time scales.
method Used large dataset on FX rates, included smoothing and measurement errors, analyzed log-log plots of realized variance increments.
result Found new stylized facts in volatility patterns, including convexity and nonlinear behavior.

Noise-resilient method improves Hurst exponent estimation accuracy in noisy data.

problem Noise degrades accuracy of Hurst exponent estimation methods.
method Noise-Controlled ALPHEE (NC-ALPHEE) using wavelet multi-scale analysis and neural network combination.
result NC-ALPHEE consistently outperforms existing techniques in noisy conditions.

Detection of power-law behavior and studies of scaling exponents uncover the characteristics of complexity in many real world phenomena. The complexity of financial markets has always presented challenging issues and provided interesting findings, such as the inverse cubic law in the tails of stock price fluctuation di…

2018-03-22abs ↗pdf ↗

We empirically analyze the most volatile component of the electricity price time series from two North-American wholesale electricity markets. We show that these time series exhibit fluctuations which are not described by a Brownian Motion, as they show multi-scaling, high Hurst exponents and sharp price movements. We …

2015-07-21abs ↗pdf ↗

Neural networks' performance scales with data size, explained by data manifold dimensionality.

problem Understanding the scaling of neural network performance with the number of parameters.
method Explained by the intrinsic dimension of the data manifold, confirmed through teacher/student framework and various datasets.
result The scaling exponent α is approximately 4 divided by the intrinsic dimension d of the data manifold.

A method for estimating the cross-correlation Cxy(τ)C_{xy}(τ) of long-range correlated series x(t)x(t) and y(t)y(t), at varying lags ττ and scales nn, is proposed. For fractional Brownian motions with Hurst exponents H1H_1 and H2H_2, the asymptotic expression of Cxy(τ)C_{xy}(τ) depends only on the lag ττ (wide-sense stationarit…

2008-04-13abs ↗pdf ↗

Different investment strategies are adopted in short-term and long-term depending on the time scales, even though time scales are adhoc in nature. Empirical mode decomposition based Hurst exponent analysis and variance technique have been applied to identify the time scales for short-term and long-term investment from …

2019-06-13abs ↗pdf ↗

The multifractal behavior for tick data of prices is investigated in Korean financial market. Using the rescaled range analysis(R/S analysis), we show the multifractal nature of returns for the won-dollar exchange rate and the KOSPI. We also estimate the Hurst exponent and the generalized qqth-order Hurst exponent in …

2003-05-13abs ↗pdf ↗

Unified theory for neural scaling laws in hierarchically compositional data.

problem Understanding neural scaling laws in hierarchically compositional data.
method Probabilistic context-free grammars and power-law distributed production rules.
result Unified learning curve behavior for classification and next-token prediction tasks.

High-dimensional shrinkage risk depends on the default prior for the common scale.

problem Choosing the default prior for the common scale in high-dimensional shrinkage.
method Using radial-power benchmark to compare variance-flat and standard deviation-flat priors.
result The standard deviation-flat prior has a one-unit asymptotic risk advantage near the origin.

Study on heat content for domains with fractal boundaries.

problem Analyzing short-time asymptotics of heat content for domains with fractal boundaries.
method Developing mathematical analysis on de Gennes' hypothesis and exploring fractal curvatures.
result Fractal curvatures and their scaling exponents may emerge in the short-time heat content asymptotics of domains with fractal boundaries.

Algorithm identifies fractal system's scaling exponents in high dimensions.

problem Statistical identification of Hurst distribution in high-dimensional fractal systems.
method Wavelet random matrices, modified spectral clustering, model selection.
result Algorithm consistently estimates Hurst distribution in moderately high dimensions.

The conventional formal tool to detect effects of the financial persistence is in terms of the Hurst exponent. A typical corresponding result is that its value comes out close to 0.5, as characteristic for geometric Brownian motion, with at most small departures from this value in either direction depending on the mark…

2005-04-22abs ↗pdf ↗

Estimates roughness of stochastic processes without assuming specific models.

problem Estimating roughness of stochastic processes without assuming specific models.
method Using Faber-Schauder coefficients and martingales, we provide a method to estimate the roughness exponent of stochastic processes.
result The roughness exponent can be estimated without assuming specific models, providing a strong consistency result for the Gladyshev estimators.

We propose a network description of large market investments, where both stocks and shareholders are represented as vertices connected by weighted links corresponding to shareholdings. In this framework, the in-degree (kink_{in}) and the sum of incoming link weights (vv) of an investor correspond to the number of asset…

2003-10-21abs ↗pdf ↗

We study the tick dynamical behavior of the bond futures in Korean Futures Exchange(KOFEX) market. Since the survival probability in the continuous-time random walk theory is applied to the bond futures transaction, the form of the decay function in our bond futures model is discussed from two kinds of Korean Treasury …

2002-12-17abs ↗pdf ↗

We consider the structure functions S^(q)(T), i.e. the moments of order q of the increments X(t+T)-X(t) of the Foreign Exchange rate X(t) which give clear evidence of scaling (S^(q)(T)~T^z(q)). We demonstrate that the nonlinearity of the observed scaling exponent z(q) is incompatible with monofractal additive stochasti…

2001-02-21abs ↗pdf ↗

A new concept, called balanced estimator of diffusion entropy, is proposed to detect scalings in short time series. The effectiveness of the method is verified by means of a large number of artificial fractional Brownian motions. It is used also to detect scaling properties and structural breaks in stock price series o…

2012-11-13abs ↗pdf ↗

Deep learning (DL) creates impactful advances following a virtuous recipe: model architecture search, creating large training data sets, and scaling computation. It is widely believed that growing training sets and models should improve accuracy and result in better products. As DL application domains grow, we would li…

2017-12-01abs ↗pdf ↗

Optimal learning rates decay to zero in easy tasks and maintain a warmup phase in hard tasks.

problem Optimizing learning rates under functional scaling laws for model training.
method Deriving optimal learning-rate schedules based on exponents ss and ββ.
result Sharp phase transition between easy and hard tasks, with different decay behaviors.

Model shows loss curve with two distinct exponents due to sparse activations.

problem Sparse activations impact neural network scaling laws.
method Introduced a model for neural scaling laws under sparse activations, derived asymptotic population loss, and analyzed gradient-descent dynamics.
result Loss curve exhibits double-descent peak near interpolation threshold with two distinct scaling exponents.