We prove a scaling limit theorem for the super-replication cost of options in a Cox--Ross--Rubinstein binomial model with transient price impact. The correct scaling turns out to keep the market depth parameter constant while resilience over fixed periods of time grows in inverse proportion with the duration between tr…
Theory explains neural network scaling with dataset and model size.
problem Neural network scaling laws with dataset and model size.
method Identified variance-limited and resolution-limited scaling behaviors.
result Four scaling regimes explained: infinite data, infinite width, resolution-limited, and large width.
The paper studies scaling limits of hedging prices in financial models.
problem Scaling limits of exponential utility indifference prices in financial models.
method Formulated dual problem as stochastic control, solved HJB equation for upper bound, used duality result for lower bound.
result Represented scaling limit in terms of specific relative entropy and constructed asymptotic optimal hedging strategies.
New framework for understanding infinite-width neural networks.
problem Understanding the infinite-width limit behavior of neural networks.
method General framework to study limit behavior of neural models based on hyperparameter scaling.
result Derives scaling for existing mean-field and neural tangent kernel limits and introduces new dynamically stable limits.
The paper analyzes SGD in high-dimensional networks, revealing new scaling limits.
problem Understanding SGD dynamics in high-dimensional networks.
method Analyzing the effective dynamics of SGD using recent work on the subject.
result A new correction term emerges at the critical scaling regime, changing the phase diagram.
In this paper we derive a scaling limit for an infinite dimensional limit order book model driven by Hawkes random measures. The dynamics of the incoming order flow is allowed to depend on the current market price as well as on a volume indicator. With our choice of scaling the dynamics converges to a coupled SDE-ODE s…
Study scaling limits of utility indifference prices in discretized Bachelier model.
problem Analyzing utility indifference prices for path-dependent European options in a discretized Bachelier model.
method Purely probabilistic approach, including duality argument, optimal drift control problem, martingale techniques, and strong invariance principles.
result Obtained a scaling limit for utility indifference prices as the number of trading times increases.
We develop a second-order model for limit order books in a single scaling regime.
problem Modeling price and volume dynamics in a limit order book with market and limit orders at a common time scale.
method Established a first- and second-order approximation for an infinite dimensional limit order book model.
result Proved the existence and uniqueness of a solution for the second-order approximation.
We establish decoupled functional CLTs for two-time-scale stochastic approximation.
problem Understanding the asymptotic behavior of two-time-scale stochastic approximation.
method Martingale problem approach and auxiliary sequence.
result The limiting dynamics of two-time-scale SA are independent of each other.
New model captures asymmetric rough volatility with Zumbach effect.
problem Capturing asymmetric rough volatility and Zumbach effect.
method Proposes a bivariate QHawkes process to model asymmetric buying and selling actions.
result Derives a super-rough-Heston model preserving the Zumbach effect.
Study scaling limits for option pricing in trinomial models.
problem Analyzing exponential hedging in trinomial models converging to Black-Scholes.
method Purely probabilistic approach using duality, martingale, and weak-convergence techniques.
result Derives a scaling limit for exponential certainty-equivalent prices in trinomial models.
Derives scaling limits and fluctuations for SGD in high dimensions.
problem Understanding SGD behavior in high-dimensional settings with varying noise levels.
method Interacting particle system approach, treating SGD iterates as such, with covariance structure considered.
result Precise three-step phase transition observed in SGD behavior: ballistic, diffusive, then random.
The paper studies scaling limits of Wasserstein metrics on Gaussian mixture models.
problem Understanding the scaling limits of Wasserstein metrics on Gaussian mixture models.
method Scaling limit approach on Gaussian mixture models, including inhomogeneous and extended models.
result Existence of the limit of the Wasserstein metric after renormalization for GMMs with zero variance.
New phases identified in neural scaling laws with compute limits.
problem Understanding neural scaling laws under compute constraints.
method Solved neural scaling model with stochastic gradient descent, derived loss curves, analyzed model-parameter-count phases.
result Identified 4 phases (+3 subphases) in data-complexity/target-complexity phase-plane, derived exponents.
Financial markets can be described on several time scales. We use data from the limit order book of the London Stock Exchange (LSE) to compare how the fluctuation dominated microstructure crosses over to a more systematic global behavior.
Study utility indifference pricing with delayed investment information in a Bachelier model.
problem Investment decisions based on delayed information in a Bachelier model.
method Developed discrete-time duality and used techniques from [7] to compute scaling limits.
result Utility indifference prices scaling limit for vanishing delay with quadratic penalty.
High-dimensional SGD limits show surprising dynamics and phase transitions.
problem Understanding SGD in high dimensions and its scaling limits.
method Proving limit theorems for SGD trajectories in high dimensions, choosing summary statistics, initialization, and step-size.
result Critical scaling regime for step-size, new correction term, and complex diffusive limits.
Continuous time random walks impose a random waiting time before each particle jump. Scaling limits of heavy tailed continuous time random walks are governed by fractional evolution equations. Space-fractional derivatives describe heavy tailed jumps, and the time-fractional version codes heavy tailed waiting times. Thi…
New scaling framework for MoE architectures ensures stability and optimal performance at scale.
problem Lack of principled understanding of how hyperparameters should scale in MoE architectures.
method Developed a novel Dynamical Mean Field Theory (DMFT) for three scaling regimes of MoE architectures.
result Derived Maximally Scale-Stable Parameterization (MSSP) for SGD and Adam, providing robust learning rate transfer and monotonic improvement with scale.
New method for Bayesian neural networks with unbounded weights.
problem Posterior inference for Bayesian neural networks with unbounded weights.
method Conditionally Gaussian representation for efficient posterior inference.
result Interpretable and computationally efficient procedure for posterior inference.
Study examines infinite limits of transformer dynamics, identifying key parameterizations.
problem Understanding the training dynamics of transformer models in the feature learning regime.
method Analysis of infinite scaling limits using dynamical mean field theory.
result Identified parameterizations that admit well-defined infinite width and depth limits.
Study of splitting maps in Type I Ricci flows for understanding singular set structure.
problem Understanding the structure of the singular set in non-collapsed Ricci limit spaces.
method Construction and investigation of almost splitting maps on Ricci flows that are almost self-similar.
result Sharp splitting maps remain splitting maps at smaller scales under certain conditions.
Study reveals how initialization scale affects training accuracy in linear networks.
problem Understanding implicit bias in linear classification models.
method Asymptotic analysis of gradient flow trajectories and training loss minimization.
result Implicit bias is more complex at reasonable initialization scales and training accuracies.
Theory predicts neural scaling exponents from language statistics.
problem No existing theory could quantitatively predict neural scaling exponents.
method Isolated two key statistical properties of language.
result Derives a simple formula predicting neural scaling exponents.
New method μP2 improves neural network training by scaling perturbations layerwise.
problem Improving neural network performance as models scale up.
method Layerwise perturbation scaling in the infinite-width limit of neural networks.
result Layerwise perturbation scaling ensures all layers are effectively perturbed in the limit.
CrossAD detects anomalies in time series data by considering cross-scale associations and cross-window modeling.
problem Anomaly detection in time series data is challenging due to varying patterns at different scales and fixed window sizes.
method CrossAD incorporates cross-scale reconstruction and a query library to capture dynamic cross-scale associations and comprehensive context.
result CrossAD achieves state-of-the-art performance in anomaly detection across multiple real-world datasets.
The paper studies Hawkes processes under mean-field limits and criticality conditions.
problem Analyzing nearly unstable Hawkes processes in a mean-field regime.
method Extending the method by Jaisson and Rosenbaum, establishing scaling limits and propagation of chaos.
result Scaling limits of Hawkes processes are stochastic Volterra diffusions of affine type, with three distinct limiting regimes.
We prove a relation between the scaling hβ of the elastic energies of shrinking non-Euclidean bodies Sh of thickness h→0, and the curvature along their mid-surface S. This extends and generalizes similar results for plates [BLS16, LRR] to any dimension and co-dimension. In particular, it proves that the na…
We investigate the behavior of limit order books on the meso-scale motivated by order execution scheduling algorithms. To do so we carry out empirical analysis of the order flows from market and limit order submissions, aggregated from tick-by-tick data via volume-based bucketing, as well as various LOB depth and shape…
Sharp theory of neural network scaling laws for hierarchical targets.
problem Learning hierarchical multi-index models in neural networks.
method Sharp information-theoretic scaling laws derived for two-layer neural networks.
result Optimal rates achieved by a simple spectral estimator.
We consider Ricci flow of complete Riemannian manifolds which have bounded non-negative curvature operator, non-zero asymptotic volume ratio and no boundary. We prove scale invariant estimates for these solutions. Using these estimates, we show that there is a limit solution, obtained by scaling down this solution at a…
Large learning rates work surprisingly well in standard parameterization, contrary to theory.
problem Theoretical limits of large learning rates do not match practical network behavior.
method Fine-grained analysis of learning rates and network behavior under cross-entropy loss.
result There are two distinct sub-regimes of unstable learning rates, with a controlled divergence regime where features continue to evolve.
We prove a central limit theorem for the components of the largest eigenvectors of the adjacency matrix of a finite-dimensional random dot product graph whose true latent positions are unknown. In particular, we follow the methodology outlined in \citet{sussman2012universally} to construct consistent estimates for the …
New neural scaling law found for simple quadratic function.
problem Neural scaling laws and their predictions for model performance.
method Analysis of neural networks, lottery ticket ensembling, statistical interpretation.
result Found a new scaling law (α=1) for a simple quadratic function, contradicting previous theories. SGD with constant stepsize converges to a non-Gaussian limit near flat minima.
problem Behavior of SGD near flat minima with convex objectives.
method Analyzes SGD with Markovian noise and contractive driving chain.
result Invariant law concentrates on scale α1/m and converges weakly to a non-Gaussian stationary distribution. Study on deep multi-head self-attention dynamics, proving homogenized limits under specific scalings.
problem Understanding the behavior of deep multi-head self-attention models as depth increases.
method Random model of deep multi-head self-attention, viewing depth as time, and analyzing the residual stream as a particle system.
result Homogenized limit of the dynamics, leading to deterministic or stochastic behavior depending on scaling, with implications for representation collapse.
This paper explores how neural network width and depth behave as they approach infinity.
problem Understanding the behavior of neural functions as width and depth go to infinity.
method Formal definition of commutativity framework, study of neural covariance kernel, novel proof techniques.
result Taking width and depth to infinity in a deep neural network with skip connections results in the same covariance structure, regardless of the order of taking limits.
We study super--replication of European contingent claims in an illiquid market with insider information. Illiquidity is captured by quadratic transaction costs and insider information is modeled by an investor who can peek into the future. Our main result describes the scaling limit of the super--replication prices wh…
Uniform scaling limits in AdamW-trained transformers converge to ODEs.
problem Understanding the dynamics of large-depth transformers trained with AdamW.
method Modeling transformer dynamics as an interacting particle system coupled through attention, proving convergence to ODEs.
result The joint dynamics of hidden states and backpropagated variables converge uniformly to an ODE system.
Paper analyzes infinite-width attention layers using Tensor Programs.
problem Capturing the infinite-width limit of attention layers.
method Tensor Programs framework to rigorously identify the limit distribution.
result Derives exact form of infinite-width limit distribution without Gaussian approximations.
ResNets converge to a limit model with improved error rates.
problem Understanding convergence of ResNets in the large-scale limit.
method Combining cavity method and propagation of chaos arguments on skeleton maps.
result Convergence rate of O(1/sqrt(D)) for ResNets in the large-scale limit.
Combines machine learning and convex limiting for accurate subgrid flux modeling in shallow-water equations.
problem Accurate subgrid flux modeling in shallow-water equations.
method Machine learning and flux limiting for property-preserving subgrid scale modeling.
result The proposed method produces meaningful closures even in untrained scenarios.
Paper proposes efficient GCN learning method for limited data.
problem Learning GCNs from data with extremely limited annotations.
method Adaptive sampling strategy and model compression.
result Cut down annotation requirement by 90% and compress parameters 6x.
SIMPGEN improves SWOT SSH data interpretation by removing noise and preserving fine-scale features.
problem Noisy data and limited fine-scale observations in oceanic processes.
method Simulation-Informed Metric and Prior for Generative Ensemble Networks (SIMPGEN) combining real SWOT observations with simulated reference data.
result SIMPGEN effectively removes noise, preserving fine-scale features better than existing neural methods.
We introduce a solvable model of randomly growing systems consisting of many independent subunits. Scaling relations and growth rate distributions in the limit of infinite subunits are analysed theoretically. Various types of scaling properties and distributions reported for growth rates of complex systems in a variety…
With the increasingly widespread deployment of generative models, there is a mounting need for a deeper understanding of their behaviors and limitations. In this paper, we expose the limitations of Variational Autoencoders (VAEs), which consistently fail to learn marginal distributions in both latent and visible spaces…
Efficiently scales continuous kernels with sparse Fourier domain learning.
problem High computational and memory demands, spectral bias in continuous kernels.
method Sparse learning in the Fourier domain.
result Efficient scaling of continuous kernels, reduced computational and memory requirements, mitigated spectral bias.
The order submission and cancelation processes are two crucial aspects in the price formation of stocks traded in order-driven markets. We investigate the dynamics of order cancelation by studying the statistical properties of inter-cancelation durations defined as the waiting times between consecutive order cancelatio…