Counterexamples show failure of uniform laws of large numbers for subdifferentials.
problem Failure of uniform laws of large numbers for subdifferentials under natural assumptions.
method Univariate and bivariate random Lipschitz and convex functions with smooth pieces.
result Counterexamples demonstrate failure of uniform laws of large numbers for subdifferentials.
Logistic regression gets a new, simpler uniform bound.
problem Finding a uniform bound for logistic regression's empirical risk.
method PAC-Bayes approach with second-order expansion and Rademacher-complexity bounds.
result Provides a dimension-free uniform concentration bound.
Study laws of large numbers in online classification, determining optimal regret bounds.
problem Understanding how sequential sampling affects online learning and classification.
method Characterized online learnable classes and determined optimal regret bounds using Littlestone's dimension.
result Optimal regret bounds in online learning are determined, resolving open questions.
We show that the sets in a family with finite VC dimension can be uniformly approximated within a given error by a finite partition. Immediate corollaries include the fact that VC classes have finite bracketing numbers, satisfy uniform laws of averages under strong dependence, and exhibit uniform mixing. Our results ar…
For any family of measurable sets in a probability space, we show that either (i) the family has infinite Vapnik-Chervonenkis (VC) dimension or (ii) for every epsilon > 0 there is a finite partition pi such the pi-boundary of each set has measure at most epsilon. Immediate corollaries include the fact that a family wit…
Study on order book dynamics with uniform catastrophes, explaining volatility and trends.
problem Understanding volatility and trends in financial markets with different types of liquidity.
method Stochastic models and population processes with uniform catastrophes.
result Law of large numbers, central limit theorem, and large deviations proved for the model.
Sharp bounds on uniform generalization errors in binary linear classification.
problem Understanding the uniform generalization errors in binary linear classification.
method Isoperimetric arguments, Poincaré and log-Sobolev inequalities for joint distributions.
result Sharp concentration bounds on uniform generalization errors, almost sure convergence in broad settings.
Randomly glued tetrahedra form connected 3-manifolds with a single boundary.
problem Understanding the properties of random three-manifolds formed by truncated tetrahedra.
method Asymptotic analysis of random glued manifolds, proving laws of large numbers, and bounding various topological and geometric properties.
result The random manifolds are connected, have a single boundary component, and admit a unique hyperbolic metric with a uniform spectral gap.
Novel groups exhibit contradictory behaviors with respect to Burnside laws.
problem Understanding probabilistic behaviors of groups under Burnside laws.
method Geometric analysis of relations, information-theoretic coding, combinatorial and probabilistic methods.
result Groups can satisfy Burnside laws with probability 1 for some generating sets and 0 for others.
Large models follow power laws in performance with dataset size or parameters.
problem Understanding neural scaling laws in large language models.
method Joint generative data model and random feature model.
result Modeling and solving the dual limit reveals insights into scaling laws.
Study uniform convergence of random walk Laplacians to diffusion Laplacian on smooth manifolds.
problem Uniform convergence of random walk Laplacians to diffusion Laplacian on smooth manifolds.
method Analysis of random walks on geometric and directed kNN graphs, using concentration tools and differential geometry.
result Uniform convergence of kNN Laplacians to diffusion Laplacian, without continuity of transition kernel. This note presents a kind of the strong law of large numbers for an insurance risk caused by a single catastrophic event rather than by an accumulation of independent and identically distributed risks. We derive this result by a large diversification effect resulting from optimal allocation of the risk to many reinsure…
The paper examines the unexpected losses and risk ratios for co-monotonic alternatives in large portfolios.
problem Understanding the unexpected losses and risk ratios for large portfolios with co-monotonic alternatives.
method Analyzes the asymptotic behavior of unexpected losses and risk ratios for co-monotonic alternatives using monotone cash-additive risk measures and Choquet insurance premia.
result Unexpected losses of large weighted portfolios are of order o(nλn), where λn is the average weight. The study explores generalized divergences and exponential families with a focus on sufficient conditions and laws of large numbers.
problem Generalization of Kullback-Leibler divergence and exponential families.
method Investigation of (h,τ)-divergence and (h,τ)-exponential families, definition of (h,τ)-dependence, proof of law of large numbers. result Sufficient condition for (h,τ)-divergence to induce Hessian structure on (h,τ)-exponential family, proof of law of large numbers. We study a resource utilization scenario characterized by intrinsic fitness. To describe the growth and organization of different cities, we consider a model for resource utilization where many restaurants compete, as in a game, to attract customers using an iterative learning process. Results for the case of restauran…
Study on Volterra Cox-Ingersoll-Ross process, proving asymptotic independence and ergodicity.
problem Analyzing the Volterra Cox-Ingersoll-Ross process and its properties.
method Fine asymptotic analysis of Volterra Riccati equation, affine transformation formula.
result Proves asymptotic independence and ergodicity of the process.
The paper analyzes Bayesian neural networks trained with VI, proving a law of large numbers for different schemes.
problem Training Bayesian neural networks with variational inference.
method Analyzes three training schemes: exact estimation, Bayes by Backprop, and Minimal VI.
result All training schemes converge to the same mean-field limit.
Uniform heat kernel and diffusion bridge asymptotics for sub-Riemannian geometry.
problem Analyzing sub-Riemannian heat kernels and their derivatives on incomplete manifolds.
method Localized asymptotic analysis, focusing on minimizing geodesics and the non-abnormal cut locus.
result Uniform bounds and expansions for heat kernels and their derivatives on compacts, including the diffusion bridge measure.
Study uses VIX for zero-coupon Treasury rates, proving long-term stability and returns.
problem Modeling zero-coupon Treasury rates with VIX for volatility.
method Multivariate autoregressive stochastic volatility model, proving stability and Law of Large Numbers.
result VIX accurately models zero-coupon Treasury rates and returns.
Algorithms based on spectral graph cut objectives such as normalized cuts, ratio cuts and ratio association have become popular in recent years because they are widely applicable and simple to implement via standard eigenvector computations. Despite strong performance for a number of clustering tasks, spectral graph cu…
Law derived for neural networks with sparse connections.
problem Understanding the behavior of neural networks with sparse connections.
method Law of large numbers for empirical distribution of parameters derived.
result Law for neural networks with sparse connections derived.
Let F be a family of Borel measurable functions on a complete separable metric space. The gap (or fat-shattering) dimension of F is a combinatorial quantity that measures the extent to which functions f in F can separate finite sets of points at a predefined resolution gamma > 0. We establish a connection between the g…
Interpolating label noise makes models vulnerable to adversarial attacks.
problem Adversarial vulnerability of models trained on noisy labels.
method Theoretical analysis of label noise and adversarial risk relationship.
result Uniform label noise induces adversarial risk similar to worst-case poisoning.
In this paper, we briefly discuss a mathematical concept that can be used in economics.
New method resolves causal heterogeneity by defining a resolution profile.
problem Causal subgroup analyses often oversimplify heterogeneity into a small number of groups.
method Introduces a resolution profile as a functional of the causal feature law, using Bayesian-bootstrap inference.
result Shows that the resolution profile is a continuous path with discontinuities at knots, providing integer-valued subgroup numbers.
Improved neural network training for speech recognition using power-law nonlinearity and uniform distribution criterion.
problem Stability and uniformity of feature distribution in neural network training.
method Power-function based and histogram-based Maximum Uniformity of Distribution (MUD) algorithms.
result Power-function based MUD outperforms conventional MFCCs in speech recognition systems.
A universal learner achieves best rates for all distributions.
problem Improving learning algorithm rates under various settings.
method Simple extension of Levin's universal search.
result Achieves best-possible rates for all distributions.
Study spectral distribution of twisted Laplacian on high genus hyperbolic surfaces.
problem Estimating spectral distribution of twisted Laplacian on hyperbolic surfaces.
method Estimate spectral distribution by supremum norm of harmonic form; show small supremum norm for high genus surfaces; prove uniform Weyl law.
result Prove uniform Weyl law for real parts of spectrum on high genus hyperbolic surfaces.
We study the rigidity of polyhedral surfaces using variational principle. The action functionals are derived from the cosine laws. The main focus of this paper is on the cosine law for a non-triangular region bounded by three possibly disjoint geodesics. Several of these cosine laws were first discovered and used by Fe…
Develops a simple model to understand learning curves for arbitrary power laws.
problem Lack of theoretical understanding of scaling laws in machine learning.
method Analyzes a toy model to determine if learning curves are universal or depend on data distribution.
result Determines that learning curves can exhibit n−β for arbitrary power β>0. PCA whitening weighted by Zipfian word frequencies improves task performance.
problem Skewed word embedding spaces in neural models.
method PCA whitening weighted by empirical word frequencies following Zipf's law.
result Significantly improves task performance, surpassing baselines.
Empirical study finds IT project costs follow a power-law distribution, exposing risk underestimation.
problem IT project cost overruns are underestimated due to normal distribution assumptions.
method Analyzed 5,392 IT projects to examine cost overruns following a power-law distribution.
result IT project cost overruns follow a power-law distribution with a fat tail of extreme overruns.
New neural scaling law found for simple quadratic function.
problem Neural scaling laws and their predictions for model performance.
method Analysis of neural networks, lottery ticket ensembling, statistical interpretation.
result Found a new scaling law (α=1) for a simple quadratic function, contradicting previous theories. We prove a law of large numbers for the volumes of families of random hyperbolic mapping tori and Heegaard splittings providing a sharp answer to a conjecture of Dunfield and Thurston.
Survey on random walks on mapping class groups and their properties.
problem Understanding random walks on mapping class groups.
method Analyzing actions on Teichmüller spaces and curve complexes.
result Laws of large numbers and central limit theorems for random walks.
A simple model explains inference scaling in neural models.
problem Understanding how model performance improves with repeated inference attempts.
method A statistical ansatz based on memorization to study inference scaling laws.
result Inference loss exhibits a power law decay with increasing trials.
New inequalities for unbounded functions improve denoising score matching.
problem Statistical error bounds for denoising score matching with unbounded objective functions.
method Derive new concentration inequalities using McDiarmid's inequality and Rademacher complexity bounds.
result Improved statistical error bounds for denoising score matching.
This work proves that large models can be compressed significantly without losing performance.
problem Achieving comparable performance with smaller models and less data.
method Developed a universal compression theory for neural networks and datasets.
result Proved that a generic permutation-invariant function can be compressed into a function of polylogarithmic size with vanishing error.
Study finds root vertex in large networks with high probability.
problem Finding the root vertex in large growing networks.
method Constructs confidence sets for the root vertex in various random network models.
result Confidence sets of size independent of the number of vertices contain the root vertex with high probability.
Different models of capital exchange among economic agents have been proposed recently trying to explain the emergence of Pareto's wealth power law distribution. One important factor to be considered is the existence of risk aversion. In this paper we study a model where agents posses different levels of risk aversion,…
The yearly aggregated tax income data of all, more than 8000, Italian municipalities are analyzed for a period of five years, from 2007 to 2011, to search for conformity or not with Benford's law, a counter-intuitive phenomenon observed in large tabulated data where the occurrence of numbers having smaller initial digi…
Uniform systole bounds for arithmetic orbifolds and number fields.
problem Bounding systole lengths in arithmetic orbifolds.
method Geometric methods and Mahler measure.
result Uniform lower bounds for systole lengths.
This work investigates power laws in deep neural network ensembles and predicts their performance.
problem Understanding the performance of deep neural network ensembles and their optimal structure.
method Investigated the behavior of negative log-likelihood (CNLL) of a deep ensemble as a function of ensemble size and member network size, identifying power law dependencies.
result One large network may perform worse than an ensemble of several medium-size networks, known as a memory split.
This paper studies a limit order book (LOB) model, in which the order dynamics depend on both, the current best available prices and the current volume density functions. For the joint dynamics of the best bid price, the best ask price, and the standing volume densities on both sides of the LOB we derive a weak law of …
ReD improves LLM inference efficiency at fixed budget, reducing attempts and cost.
problem Improving LLM inference efficiency at a fixed budget.
method Reset-and-Discard (ReD) query method.
result ReD increases coverage@cost for a given budget, reducing attempts and cost.
Statistical properties of an order book and the effect they have on price dynamics were studied using the high-frequency NASDAQ Level II data. It was observed that the size distribution of marketable orders (transaction sizes) has power law tails with an exponent 1+mu_{market}=2.4 \pm 0.1. The distribution of limit ord…
Noise-Aware Conformal Prediction (NACP) calibrates CP for noisy labels.
problem Calibrating Conformal Prediction with noisy labels.
method Estimate conformal threshold from noisy labels using uniform noise coverage guarantee.
result Finite sample coverage guarantee for uniform noise remains effective in high-class tasks.
New framework reveals thermodynamic principles for LLM training.
problem Understanding the training dynamics of large language models.
method Introducing Neural Thermodynamic Laws (NTL) under river-valley loss landscape assumptions.
result Key thermodynamic quantities and principles naturally emerge in LLM training.