Counterexamples show failure of uniform laws of large numbers for subdifferentials.
problem Failure of uniform laws of large numbers for subdifferentials under natural assumptions.
method Univariate and bivariate random Lipschitz and convex functions with smooth pieces.
result Counterexamples demonstrate failure of uniform laws of large numbers for subdifferentials.
Study laws of large numbers in online classification, determining optimal regret bounds.
problem Understanding how sequential sampling affects online learning and classification.
method Characterized online learnable classes and determined optimal regret bounds using Littlestone's dimension.
result Optimal regret bounds in online learning are determined, resolving open questions.
Large models follow power laws in performance with dataset size or parameters.
problem Understanding neural scaling laws in large language models.
method Joint generative data model and random feature model.
result Modeling and solving the dual limit reveals insights into scaling laws.
This note presents a kind of the strong law of large numbers for an insurance risk caused by a single catastrophic event rather than by an accumulation of independent and identically distributed risks. We derive this result by a large diversification effect resulting from optimal allocation of the risk to many reinsure…
The study explores generalized divergences and exponential families with a focus on sufficient conditions and laws of large numbers.
problem Generalization of Kullback-Leibler divergence and exponential families.
method Investigation of (h,τ)-divergence and (h,τ)-exponential families, definition of (h,τ)-dependence, proof of law of large numbers. result Sufficient condition for (h,τ)-divergence to induce Hessian structure on (h,τ)-exponential family, proof of law of large numbers. Study on Volterra Cox-Ingersoll-Ross process, proving asymptotic independence and ergodicity.
problem Analyzing the Volterra Cox-Ingersoll-Ross process and its properties.
method Fine asymptotic analysis of Volterra Riccati equation, affine transformation formula.
result Proves asymptotic independence and ergodicity of the process.
The paper analyzes Bayesian neural networks trained with VI, proving a law of large numbers for different schemes.
problem Training Bayesian neural networks with variational inference.
method Analyzes three training schemes: exact estimation, Bayes by Backprop, and Minimal VI.
result All training schemes converge to the same mean-field limit.
Study uses VIX for zero-coupon Treasury rates, proving long-term stability and returns.
problem Modeling zero-coupon Treasury rates with VIX for volatility.
method Multivariate autoregressive stochastic volatility model, proving stability and Law of Large Numbers.
result VIX accurately models zero-coupon Treasury rates and returns.
Law derived for neural networks with sparse connections.
problem Understanding the behavior of neural networks with sparse connections.
method Law of large numbers for empirical distribution of parameters derived.
result Law for neural networks with sparse connections derived.
Logistic regression gets a new, simpler uniform bound.
problem Finding a uniform bound for logistic regression's empirical risk.
method PAC-Bayes approach with second-order expansion and Rademacher-complexity bounds.
result Provides a dimension-free uniform concentration bound.
In this paper, we briefly discuss a mathematical concept that can be used in economics.
Develops a simple model to understand learning curves for arbitrary power laws.
problem Lack of theoretical understanding of scaling laws in machine learning.
method Analyzes a toy model to determine if learning curves are universal or depend on data distribution.
result Determines that learning curves can exhibit n−β for arbitrary power β>0. Empirical study finds IT project costs follow a power-law distribution, exposing risk underestimation.
problem IT project cost overruns are underestimated due to normal distribution assumptions.
method Analyzed 5,392 IT projects to examine cost overruns following a power-law distribution.
result IT project cost overruns follow a power-law distribution with a fat tail of extreme overruns.
New neural scaling law found for simple quadratic function.
problem Neural scaling laws and their predictions for model performance.
method Analysis of neural networks, lottery ticket ensembling, statistical interpretation.
result Found a new scaling law (α=1) for a simple quadratic function, contradicting previous theories. We prove a law of large numbers for the volumes of families of random hyperbolic mapping tori and Heegaard splittings providing a sharp answer to a conjecture of Dunfield and Thurston.
Survey on random walks on mapping class groups and their properties.
problem Understanding random walks on mapping class groups.
method Analyzing actions on Teichmüller spaces and curve complexes.
result Laws of large numbers and central limit theorems for random walks.
A simple model explains inference scaling in neural models.
problem Understanding how model performance improves with repeated inference attempts.
method A statistical ansatz based on memorization to study inference scaling laws.
result Inference loss exhibits a power law decay with increasing trials.
This work proves that large models can be compressed significantly without losing performance.
problem Achieving comparable performance with smaller models and less data.
method Developed a universal compression theory for neural networks and datasets.
result Proved that a generic permutation-invariant function can be compressed into a function of polylogarithmic size with vanishing error.
The yearly aggregated tax income data of all, more than 8000, Italian municipalities are analyzed for a period of five years, from 2007 to 2011, to search for conformity or not with Benford's law, a counter-intuitive phenomenon observed in large tabulated data where the occurrence of numbers having smaller initial digi…
Randomly glued tetrahedra form connected 3-manifolds with a single boundary.
problem Understanding the properties of random three-manifolds formed by truncated tetrahedra.
method Asymptotic analysis of random glued manifolds, proving laws of large numbers, and bounding various topological and geometric properties.
result The random manifolds are connected, have a single boundary component, and admit a unique hyperbolic metric with a uniform spectral gap.
This work investigates power laws in deep neural network ensembles and predicts their performance.
problem Understanding the performance of deep neural network ensembles and their optimal structure.
method Investigated the behavior of negative log-likelihood (CNLL) of a deep ensemble as a function of ensemble size and member network size, identifying power law dependencies.
result One large network may perform worse than an ensemble of several medium-size networks, known as a memory split.
This paper studies a limit order book (LOB) model, in which the order dynamics depend on both, the current best available prices and the current volume density functions. For the joint dynamics of the best bid price, the best ask price, and the standing volume densities on both sides of the LOB we derive a weak law of …
ReD improves LLM inference efficiency at fixed budget, reducing attempts and cost.
problem Improving LLM inference efficiency at a fixed budget.
method Reset-and-Discard (ReD) query method.
result ReD increases coverage@cost for a given budget, reducing attempts and cost.
Statistical properties of an order book and the effect they have on price dynamics were studied using the high-frequency NASDAQ Level II data. It was observed that the size distribution of marketable orders (transaction sizes) has power law tails with an exponent 1+mu_{market}=2.4 \pm 0.1. The distribution of limit ord…
New framework reveals thermodynamic principles for LLM training.
problem Understanding the training dynamics of large language models.
method Introducing Neural Thermodynamic Laws (NTL) under river-valley loss landscape assumptions.
result Key thermodynamic quantities and principles naturally emerge in LLM training.
Study on Haantjes tensors for superintegrable systems, focusing on vanishing properties.
problem Understanding the vanishing of Haantjes tensors in superintegrable systems.
method Investigating Killing tensor fields associated with second-order superintegrable systems.
result Characterization of Haantjes-zero Killing tensor fields.
Improved scaling laws in linear regression using data reuse.
problem Sustainability of neural scaling laws when running out of new data.
method Data reuse in multi-pass stochastic gradient descent (multi-pass SGD) for M-dimensional linear models trained on N data with sketched features. result Multi-pass SGD achieves a test error of Θ(M1−b+L(1−b)/a) with L>N, improving scaling laws in data-constrained regimes. Theory explains neural network scaling with dataset and model size.
problem Neural network scaling laws with dataset and model size.
method Identified variance-limited and resolution-limited scaling behaviors.
result Four scaling regimes explained: infinite data, infinite width, resolution-limited, and large width.
This work extends the scaling law to multiple and kernel regression, challenging traditional machine learning principles.
problem Challenging traditional machine learning wisdom with scaling law in large practical models.
method Demonstrates the scaling law in multiple and kernel regression settings.
result The scaling law extends to multiple and kernel regression, providing deeper insights into LLMs.
Space exploration technology advances exponentially, consistent with Moore's and Wright's laws.
problem Predicting the advancement of space exploration technology.
method Analysis of Moore's and Wright's laws applied to space exploration technology.
result Spacecraft technology advances exponentially, consistent with Moore's and Wright's laws.
Empirical law predicts accuracy of Google Translate's translation chains.
problem Predicting accuracy in machine translation with multiple hops.
method Empirical testing of Google Translate's sequential translation.
result Accuracy decreases with the number of translating hops, following a power law.
A simple text model shows word lengths follow Zipf's law.
problem Understanding word statistics in large language models.
method A non-linguistic model of text with independent symbol draws.
result Word lengths follow a geometric distribution and Zipf's law.
Scaling laws govern predictive uncertainties in deep learning models.
problem Understanding predictive uncertainties in deep learning models.
method Empirical analysis and approximate Bayesian inference on vision and language tasks.
result Scaling laws exist for various measures of predictive uncertainty in deep learning models.
Ensembles of random-feature models can't outperform a single large model.
problem Finding the optimal balance between model size and ensemble size.
method Deterministic equivalent risk estimates and scaling laws analysis.
result Ensembles of random-feature models achieve near-optimal performance only under specific conditions.
This note presents an operational measure of fat-tailedness for univariate probability distributions, in [0,1] where 0 is maximally thin-tailed (Gaussian) and 1 is maximally fat-tailed. Among others,1) it helps assess the sample size needed to establish a comparative n needed for statistical significance, 2) allows…
We derive scaling laws for optimizing neural networks in hardware.
problem Optimizing the large parameter space of neural networks in hardware.
method Analytical derivation of scaling laws for Coordinate Descent optimization.
result Convergence is exponential and scales linearly with the number of neurons.
The relaxation dynamics of aftershocks after large volatility shocks are investigated based on two high-frequency data sets of the Shanghai Stock Exchange Composite (SSEC) index. Compared with previous relevant work, we have defined main financial shocks based on large volatilities rather than large crashes. We find th…
Let S=Γ\H be a hyperbolic surface of finite topological type, such that the Fuchsian group Γ≤PSL2(R) is non-elementary, and consider any generating set S of Γ. When sampling by an n-step random walk in π1(S)≅Γ with each step given by an element…
Employing profits data of Japanese companies in 2002 and 2003, we confirm that Pareto's law and the Pareto index are derived from the law of detailed balance and Gibrat's law. The last two laws are observed beyond the region where Pareto's law holds. By classifying companies into job categories, we find that companies …
Unified theory for neural scaling laws in hierarchically compositional data.
problem Understanding neural scaling laws in hierarchically compositional data.
method Probabilistic context-free grammars and power-law distributed production rules.
result Unified learning curve behavior for classification and next-token prediction tasks.
Following the work of Okuyama, Takayasu and Takayasu [Okuyama, Takayasu and Takayasu 1999] we analyze huge databases of Japanese companies' financial figures and confirm that the Zipf's law, a power law distribution with the exponent -1, has been maintained over 30 years in the income distribution of Japanese companies…
We study the relaxation dynamics of a financial market just after the occurrence of a crash by investigating the number of times the absolute value of an index return is exceeding a given threshold value. We show that the empirical observation of a power law evolution of the number of events exceeding the selected thre…
We prove the existence of limiting distributions for a large class of Markov chains on a general state space in a random environment. We assume suitable versions of the standard drift and minorization conditions. In particular, the system dynamics should be contractive on the average with respect to the Lyapunov functi…
We develop a dynamic point process model of correlated default timing in a portfolio of firms, and analyze typical default profiles in the limit as the size of the pool grows. In our model, a firm defaults at a stochastic intensity that is influenced by an idiosyncratic risk process, a systematic risk process common to…
Based on empirical financial time-series, we show that the "silence-breaking" probability follows a super-universal power law: the probability of observing a large movement is inversely proportional to the length of the on-going low-variability period. Such a scaling law has been previously predicted theoretically [R. …
The paper examines the unexpected losses and risk ratios for co-monotonic alternatives in large portfolios.
problem Understanding the unexpected losses and risk ratios for large portfolios with co-monotonic alternatives.
method Analyzes the asymptotic behavior of unexpected losses and risk ratios for co-monotonic alternatives using monotone cash-additive risk measures and Choquet insurance premia.
result Unexpected losses of large weighted portfolios are of order o(nλn), where λn is the average weight. We suggest an analytical approach for Pareto-Zipf law, where we assume random multiplicative noise and fragmentation processes for the growth of the number of citizens of each city and the number of the cities, respectively.
The study improves bounds on the number of closed geodesics and logarithmic improvements in the Weyl law.
problem Estimating the number of closed geodesics and improving logarithmic bounds in the Weyl law.
method Study of non-degeneracy properties of nearly closed orbits for predominant sets of metrics.
result Logarithmic improvements in the Weyl law and exponential bounds on the number of closed geodesics.