Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

103206309412 · May 202619922001200920172026
48 results for scaling regimes

Three training regimes found for scale-invariant neural networks on the sphere.

problem Training scale-invariant neural networks on the sphere with varying effective learning rate.
method Investigated three regimes of training: convergence, chaotic equilibrium, and divergence.
result Discovered three distinct training regimes with unique characteristics.

This paper proposes a multi-scale Markov-Switching GARCH model for EUR/USD volatility.

problem Non-stationary financial volatility requires models that capture changing market conditions across multiple timescales.
method Triple-timeframe Markov-Switching GARCH (MS-GARCH) framework with AR(1)-MS-GARCH models and TVTP for short horizons.
result The proposed model produces statistically distinct regimes and superior volatility forecasting performance.

Study uncovers scaling laws and spectral properties of shallow neural networks.

problem Understanding scaling laws and spectral properties of shallow neural networks.
method Leveraging connections with matrix compressed sensing and LASSO, derived a phase diagram for excess risk.
result Uncovered crossovers between scaling regimes and plateau behaviors, validated empirical observations.

We investigate multifractality in the Korean stock-market index KOSPI. The generalized qqth order height-height correlation function shows multiscaling properties. There are two scaling regimes with a crossover time around tc=40t_c =40 min. We consider the original data sets and the modified data sets obtained by removin…

2004-12-15abs ↗pdf ↗

We develop a second-order model for limit order books in a single scaling regime.

problem Modeling price and volume dynamics in a limit order book with market and limit orders at a common time scale.
method Established a first- and second-order approximation for an infinite dimensional limit order book model.
result Proved the existence and uniqueness of a solution for the second-order approximation.

New scaling framework for MoE architectures ensures stability and optimal performance at scale.

problem Lack of principled understanding of how hyperparameters should scale in MoE architectures.
method Developed a novel Dynamical Mean Field Theory (DMFT) for three scaling regimes of MoE architectures.
result Derived Maximally Scale-Stable Parameterization (MSSP) for SGD and Adam, providing robust learning rate transfer and monotonic improvement with scale.

Large learning rates work surprisingly well in standard parameterization, contrary to theory.

problem Theoretical limits of large learning rates do not match practical network behavior.
method Fine-grained analysis of learning rates and network behavior under cross-entropy loss.
result There are two distinct sub-regimes of unstable learning rates, with a controlled divergence regime where features continue to evolve.

Data pruning algorithms struggle in high compression regimes, as shown by theoretical and empirical studies.

problem Limitations of score-based data pruning algorithms in high compression regimes.
method Theoretical and empirical analysis of score-based data pruning algorithms.
result Score-based data pruning algorithms fail in high compression regimes due to 'No Free Lunch' theorems.

Study shows optimal model performance at critical level of feature learning.

problem Catastrophic forgetting in neural networks, especially in non-stationary environments.
method Systematic study on model scale and feature learning, using dynamical mean field theory.
result Optimal performance achieved at a critical level of feature learning, dependent on task non-stationarity and model scale.

The paper analyzes SGD in high-dimensional networks, revealing new scaling limits.

problem Understanding SGD dynamics in high-dimensional networks.
method Analyzing the effective dynamics of SGD using recent work on the subject.
result A new correction term emerges at the critical scaling regime, changing the phase diagram.

Model shows feature learning can improve neural scaling laws for hard tasks.

problem Understanding and improving neural network scaling laws for various task difficulties.
method Developed a solvable model of neural scaling laws, identified three scaling regimes, and demonstrated feature learning's impact on scaling exponents.
result Feature learning can improve scaling with training time and compute for hard tasks, nearly doubling the exponent.

Framework models multiscale dynamics with Bayesian learning for regime changes.

problem Analyzing complex interactions between fast and slow processes.
method Hierarchical state-space modeling with Sequential Monte Carlo.
result Bayesian approach accurately tracks state transitions and identifies switching dynamics.

Empirical analysis of financial market trends and reversions across various time scales.

problem Understanding trends and reversions in financial markets over different time scales.
method Analysis of 14 years of futures tick data, 30 years of daily futures prices, 330 years of monthly asset prices, and yearly financial data since medieval times.
result Markets exhibit trending and reversion regimes with different time scales, explaining trends persistence and reversions.

Study how generalization scales with model size and data in quadratic neural networks.

problem Understanding how generalization scales with model size and data in quadratic neural networks.
method Analyzed 2\ell_2-regularized empirical test error minimization in a quadratic two-layer network with finite-sample setting and structured data.
result Revealed a phase diagram with distinct scaling regimes as the number of parameters varies, showing data-dependent power laws controlled by spectral structure of the target.

Two distinct limits for deep learning have been derived as the network width hh\rightarrow \infty, depending on how the weights of the last layer scale with hh. In the Neural Tangent Kernel (NTK) limit, the dynamics becomes linear in the weights and is described by a frozen kernel ΘΘ. By contrast, in the Mean-Field …

2019-06-19abs ↗pdf ↗

Model shows loss curve with two distinct exponents due to sparse activations.

problem Sparse activations impact neural network scaling laws.
method Introduced a model for neural scaling laws under sparse activations, derived asymptotic population loss, and analyzed gradient-descent dynamics.
result Loss curve exhibits double-descent peak near interpolation threshold with two distinct scaling exponents.

Scaling ResNets requires careful consideration of the layer depth and output scaling factors.

problem Avoiding vanishing or exploding gradients in deep ResNets as depth increases.
method Probabilistic analysis and continuous-time limit interpretation of ResNets.
result The optimal scaling factor is αL=1Lα_L = \frac{1}{\sqrt{L}} for standard i.i.d. initializations.

The paper analyzes learning curves for kernel ridge regression with dot-product kernels.

problem Understanding the learning curves for different scaling regimes of data and model.
method Precise formulas for mean test error, bias, and variance in the mom o\infty with m/drm/d^r constant regime.
result A peak in the learning curve at mdr/r!m \approx d^r/r! for any integer rr.

Using a large database of 8 million institutional trades executed in the U.S. equity market, we establish a clear crossover between a linear market impact regime and a square-root regime as a function of the volume of the order. Our empirical results are remarkably well explained by a recently proposed dynamical theory…

2018-11-13abs ↗pdf ↗

Develops a new framework to analyze gradient flow regimes and derive explicit solutions.

problem Analyzing scaling regimes and deriving explicit analytic solutions for gradient flow in large learning problems.
method Formal power series expansion of the loss evolution with coefficients encoded by diagrams.
result Reveals different learning phases and obtains explicit solutions in some cases.

The study provides precise asymptotic theory for in-context learning by Transformers.

problem Understanding the sample complexity, pretraining task diversity, and context length for successful in-context learning.
method An exactly solvable model of linear regression task by linear attention, deriving sharp asymptotics.
result Double-descent learning curve with increasing pretraining examples, phase transition between low and high task diversity regimes.

A new approach is presented to describe the change in the statistics of the log return distribution of financial data as a function of the timescale. To this purpose a measure is introduced, which quantifies the distance of a considered distribution to a reference distribution. The existence of a small timescale regime…

2005-09-30abs ↗pdf ↗

Bayesian neural networks reveal multimodal predictive distributions.

problem Uncertainty quantification and interpretability in neural networks.
method Discretized prior for inner layer weights, Gaussian mixture approximation of posterior predictive distribution.
result Distinct parameter realizations can produce the same training error but different posterior predictive distributions.

New model captures asymmetric rough volatility with Zumbach effect.

problem Capturing asymmetric rough volatility and Zumbach effect.
method Proposes a bivariate QHawkes process to model asymmetric buying and selling actions.
result Derives a super-rough-Heston model preserving the Zumbach effect.

Gaussian equivalence fails for simple polynomial embeddings in quadratic scaling RF models.

problem Failure of Gaussian equivalence in polynomial feature embeddings under quadratic scaling.
method Introduced Conditional Gaussian Equivalent (CGE) model to capture non-Gaussian behavior.
result Correct asymptotics derived for training and test errors in CGE model.

Study on SGD dynamics and scaling laws for training quadratic neural networks in high dimensions.

problem Optimizing and understanding the training dynamics of quadratic neural networks in high-dimensional settings.
method Sharp analysis of SGD dynamics, combining matrix Riccati differential equations and matrix monotonicity arguments.
result Derivation of scaling laws for prediction risk, highlighting power-law dependencies on optimization time, sample size, and model width.

This work bridges two views of feature learning in neural networks.

problem The relationship between kernel scale changes and data-adaptive feature learning in neural networks remains unresolved.
method Using statistical mechanics, the work derives analytical expressions for network output statistics across scaling regimes.
result Kernel adaptation can be reduced to an effective kernel rescaling, but multi-scale adaptive approach provides richer insights.

This paper optimizes importance sampling for rare-event options pricing under the Heston model.

problem Efficiently pricing European call options with short maturity and deep out-of-the-money strikes.
method Asymptotic importance sampling schemes leveraging the large deviation principle and state-dependent change of measure.
result Proposed IS methods achieve logarithmic efficiency in short-maturity and deep OTM regimes, significantly reducing variance.

New dynamics for SGD in small learning rate regime.

problem Improving stochastic gradient descent in small learning rate regime.
method Introducing stochastic modified flows and distribution dependent stochastic modified flows.
result Captures fluctuating dynamics of SGD in small learning rate - infinite width scaling regime.

Optimizes dividend control in a bankruptcy process using a special Levy process.

problem Optimizing dividend payouts in a bankruptcy process.
method Using a non-standard spectrally negative Levy process with endogenous regime switching.
result Optimal dividend control is of the barrier type and the optimal barrier can be identified.

LINTEL improves INTEL's time series prediction by optimizing computation and accuracy.

problem Online prediction of time series with regime switching and outliers.
method Gaussian process-based approach with exact filtering distribution and constant-time updates.
result LINTEL is over five times faster with better quality predictions.

Optimization of neural networks scales with γ, revealing unique loss curves and optimal learning rates.

problem Understanding the impact of feature learning strength on neural network optimization.
method Empirical investigation of neural networks with varying γ, analyzing the γγ-ηη plane, and examining loss curves.
result Optimal learning rate scales non-trivially with γ, with ηγ2η^* \propto γ^2 for small γ and ηγ2/Lη^* \propto γ^{2/L} for large γ.

Develops identifiability theory for multi-lag regime-switching models.

problem Ensuring interpretability of deep latent variable models with multi-lag dependencies.
method Formulates a general theoretical framework for multi-lag Regime-Switching Models (RSMs), proving identifiability of number of regimes and multi-lag transitions.
result Establishes identifiability conditions for multi-lag regime-switching models, including Markov Switching Models and Switching Dynamical Systems.

Ensembles of random-feature models can't outperform a single large model.

problem Finding the optimal balance between model size and ensemble size.
method Deterministic equivalent risk estimates and scaling laws analysis.
result Ensembles of random-feature models achieve near-optimal performance only under specific conditions.

We analyze the Hessian spectra of large models up to 100B parameters.

problem Accurate Hessian spectra of large foundation models are difficult to obtain.
method We use shard-local finite-difference Hessian vector products and stochastic Lanczos quadrature.
result We produce the first large-scale spectral density estimates of foundation models.

New algorithm learns switching dynamics from multiple neural signals.

problem Learning accurate switching dynamical system models from multimodal neural data.
method Unsupervised learning algorithm for multiscale switching dynamical system models.
result Switching multiscale dynamical system models outperform single-scale models in behavior decoding.

Embedded ensembles improve neural network performance efficiently.

problem Improving neural network performance with fewer resources.
method Analyzing the wide network limit of gradient descent dynamics using Neural-Tangent-Kernel.
result Embedded ensembles exhibit two regimes: independent and collective, affecting performance.

Study on deep multi-head self-attention dynamics, proving homogenized limits under specific scalings.

problem Understanding the behavior of deep multi-head self-attention models as depth increases.
method Random model of deep multi-head self-attention, viewing depth as time, and analyzing the residual stream as a particle system.
result Homogenized limit of the dynamics, leading to deterministic or stochastic behavior depending on scaling, with implications for representation collapse.