This work analyzes neural scaling laws using power-law data spectra and derives analytical expressions for generalization error.
problem Understanding how neural network performance scales with key factors like data size and model complexity.
method Statistical mechanics techniques applied to one-pass stochastic gradient descent in a student-teacher framework.
result Derivation of analytical expressions for generalization error under power-law data spectra and identification of conditions for power-law scaling.
The paper proves geometric and spectral alignment for deep neural networks.
problem Understanding the singular spectra of deep neural network layers.
method Proves deterministic quotient-geometric estimates for singular spectra of Frobenius-normalized layer factors.
result Exact power-law spectra form a trace-normalized Cartan orbit under Frobenius normalization.
Neural networks learn simpler features first, then more complex ones; Fourier analysis reveals this pattern.
problem Understanding the learning dynamics of neural networks, especially with natural image data.
method Fourier analysis of translation-invariant and power-law spectra to study feature learning.
result Simple neural networks first rely on amplitude information, then phase information, and power-law spectra can accelerate learning phase information.
Robust CD method for real-world time series with power-law distributions.
problem Challenges in causal discovery due to noise sensitivity.
method Power-law spectral feature extraction for robust CD.
result Consistently outperforms state-of-the-art alternatives on real-world datasets.
The log-periodic power law (LPPL) is a model of asset prices during endogenous bubbles. A major open issue is to verify the presence of LPPL in price sequences and to estimate the LPPL parameters. Estimation is complicated by the fact that daily LPPL returns are typically orders of magnitude smaller than measured price…
The study examines cryptocurrency market activity, revealing multifractal inter-transaction times and challenging traditional statistical models.
problem Analyzing long-range autocorrelations and multifractality in cryptocurrency market activity.
method Analysis of tick-by-tick data from multiple cryptocurrency trading platforms, focusing on inter-transaction times, transaction volumes, and volatility.
result Inter-transaction times exhibit multifractality, indicating periods of increased market activity are more complex than quiet periods.
Signatures of universality are detected by comparing individual eigenvalue distributions and level spacings from financial covariance matrices to random matrix predictions. A chopping procedure is devised in order to produce a statistical ensemble of asset-price covariances from a single instance of financial data sets…
The study reveals the spectral structure of attention layers and its implications for generalization.
problem Understanding the spectral structure and generalization of trained attention layers.
method Empirical risk minimization in a single-head tied-attention layer, using random matrix theory, spin-glass theory, and approximate message passing.
result Exact high-dimensional characterization of training and test error, interpolation and recovery thresholds, and spectrum of the key and query matrices.
This work explains scaling laws as redundancy laws in deep learning.
problem The mathematical origins of scaling laws in deep learning models remain unclear.
method Kernel regression and analysis of data covariance spectra.
result Scaling laws can be explained as redundancy laws, revealing the learning curve's slope depends on data redundancy.
Random matrix theory is used to assess the significance of weak correlations and is well established for Gaussian statistics. However, many complex systems, with stock markets as a prominent example, exhibit statistics with power-law tails, that can be modelled with Levy stable distributions. We review comprehensively …
Paper examines the structure of stochastic gradients in deep learning.
problem Exploring the structure and heavy tails of stochastic gradients in deep learning.
method Conducted formal statistical tests on stochastic gradients and gradient noise.
result Stochastic gradients and gradient noise do not exhibit power-law heavy tails, but their covariance spectra do.
Stochastic momentum methods trade compute efficiency for serial runtime.
problem Stochastic momentum methods trade compute efficiency for serial runtime.
method Stochastic HB and ASGD for consistent linear regression with Gaussian covariates.
result HB preserves SGD-level CE over a larger batch-size window, allowing larger batches to reduce serial runtime until HB reaches its deterministic accelerated scale.
Signals consisting of a sequence of pulses show that inherent origin of the 1/f noise is a Brownian fluctuation of the average interevent time between subsequent pulses of the pulse sequence. In this paper we generalize the model of interevent time to reproduce a variety of self-affine time series exhibiting power spec…
Study calculates spectra of minimal hypersurfaces in hyperbolic space.
problem Computing Laplacian spectra of minimal hypersurfaces.
method Analyzes hypersurfaces in hyperbolic space with specific asymptotic data.
result Obtains spectra and extremal properties of the bottom of the spectrum.
Analyzes SGD dynamics on multi-class problems with exact expressions.
problem Analyzing SGD dynamics on multi-class problems.
method Developed a framework for analyzing training and learning rate dynamics using exact expressions.
result Exact expressions for risk and overlap with true signal in terms of ODEs.
Power laws detected in financial data, modeled with random multipliers.
problem Detecting power laws in financial data.
method Investigated data from financial instruments, proposed a model based on sums of Maxwell-Boltzmann distributions with random multipliers.
result Detected power laws with various exponents in financial data, proposed a universal model.
Paper resolves decades-old problem about L-spectra.
problem Identifying L-spectra local information with geometric data. method Proved equivalence of L-orientations and characteristic classes. result Levitt-Ranicki's theory equivalent to Brumfiel-Morgan's classes.
Study on KRR with power-law data, showing better sample complexity.
problem High-dimensional kernel ridge regression with anisotropic power-law covariance.
method Explicit characterization of kernel spectrum and asymptotic analysis of excess risk.
result Sample complexity is governed by effective dimension, not ambient dimension.
Cycle-StarNet bridges theory and data by adapting synthetic spectra to observational data.
problem Lack of consistency between theoretical stellar models and observational data.
method Hybrid generative domain adaptation using unsupervised learning on large spectroscopic surveys.
result Improved spectral fitting and reduced gap between synthetic and observational data.
We use data on wealth of the richest persons taken from the "rich lists" provided by business magazines like Forbes to verify if upper tails of wealth distributions follow, as often claimed, a power-law behaviour. The data sets used cover the world's richest persons over 1996-2012, the richest Americans over 1988-2012,…
Superposition accelerates training to a universal power-law exponent.
problem Training dynamics in neural networks.
method Teacher-student framework and analytic theory.
result Superposition leads to a universal power-law exponent of ~1, independent of data and channel statistics.
New ICA method for sources with mixed spectra.
problem Inaccurate separation of sources with temporal autocorrelations and mixed spectra.
method Estimates spectral density functions and line spectra using cubic splines and indicator functions, then maximizes the Whittle likelihood function.
result Outperforms existing ICA methods in simulations and EEG data applications.
We take prior-to-crash market prices (NASDAQ, Dow Jones Industrial Average) as a signal, a function of time, we project these discrete values onto a vertical axis, thus obtaining a Cantordust. We study said cantordust with the tools of multifractal analysis, obtaining spectra by definition and by lagrangian coordinates…
BSD is a Bayesian framework for analyzing neural spectral data.
problem Challenges in statistical analysis and group-level comparisons of neural power spectra.
method Bayesian Spectral Decomposition (BSD) for parametric models of neural spectra.
result BSD outperforms existing methods in model selection and parameter estimation.
This paper describes a novel energy-based probabilistic distribution that represents complex-valued data and explains how to apply it to direct feature extraction from complex-valued spectra. The proposed model, the complex-valued restricted Boltzmann machine (CRBM), is designed to deal with complex-valued visible unit…
Investigates point spectra of vector fields and their properties.
problem Understanding the point spectra of vector fields.
method Define and study point spectra, prove properties under isometries, and analyze compactly supported fields.
result Point spectra are well-behaved under isometries and trivial for compactly supported fields.
Khovanov spectra are shown to be functorial under certain conditions.
problem Understanding functoriality of Khovanov spectra.
method Proving functoriality up to homotopy and sign for Khovanov spectra.
result Khovanov spectra are functorial under specific conditions.
Method maps imperfect simulations to observed stellar spectra using unsupervised domain adaptation.
problem Mapping from large sets of imperfect simulations and observational data.
method Adversarial autoencoders, cycle-consistency constraint, and generative surrogate physics emulator network.
result Reconstructed spectra quality and discovery of new spectral features.
Unified theory for neural scaling laws in hierarchically compositional data.
problem Understanding neural scaling laws in hierarchically compositional data.
method Probabilistic context-free grammars and power-law distributed production rules.
result Unified learning curve behavior for classification and next-token prediction tasks.
Bayesian nonparametric CMS improves frequency estimation for power-law data.
problem Estimating frequencies of low-frequency tokens in power-law data streams.
method Developed a learning-augmented count-min sketch using a normalized inverse Gaussian process prior.
result The approach achieves remarkable performance in estimating low-frequency tokens.
We apply a novel spectral graph technique, that of locally-biased semi-supervised eigenvectors, to study the diversity of galaxies. This technique permits us to characterize empirically the natural variations in observed spectra data, and we illustrate how this approach can be used in an exploratory manner to highlight…
We give a simple sufficient condition for Quinn's "bordism-type spectra" to be weakly equivalent to strictly associative ring spectra. We also show that Poincare bordism and symmetric L-theory are naturally weakly equivalent to monoidal functors. Part of the proof of these statements involves showing that Quinn's funct…
Method calculates spectra of Rarita-Schwinger operator on symmetric spaces.
problem Calculating spectra of the Rarita-Schwinger operator on compact symmetric spaces.
method Using Weitzenböck formulas, Laplace operator, Casimir operator, Freudenthal's formula, and branching rules.
result Obtained spectra on the sphere, complex projective space, and quaternionic projective space.
Background: High-throughput proteomics techniques, such as mass spectrometry (MS)-based approaches, produce very high-dimensional data-sets. In a clinical setting one is often interested in how mass spectra differ between patients of different classes, for example spectra from healthy patients vs. spectra from patients…
Intertrade duration of equities is an important financial measure characterizing the trading activities, which is defined as the waiting time between successive trades of an equity. Using the ultrahigh-frequency data of a liquid Chinese stock and its associated warrant, we perform a comparative investigation of the sta…
In this paper we tackle the problem of estimating the power-law tail exponent of income distributions by using the Hill's estimator. A subsample semi-parametric bootstrap procedure minimising the mean squared error is used to choose the power-law cutoff value optimally. This technique is applied to personal income data…
Proposes a method for training Bayesian neural networks using synthetic data from Raman and CARS spectra.
problem Limited real observations in Raman and CARS spectroscopy.
method Log-Gaussian Gamma Processes and Bayesian Neural Networks.
result Trained Bayesian neural networks provide accurate estimates of Raman and CARS spectra with uncertainty quantification.
Proves spectra equivalence for Riemannian manifolds.
problem Equivalence of Almgren-Pitts and phase-transition half-volume spectra.
method Proof of spectra equivalence for Riemannian manifolds.
result Confirms conjecture about spectra equivalence.
Proposes a new complex Gaussian distribution for better modeling of complex-valued signals.
problem Limited ability of Gaussian distribution to represent diverse amplitude characteristics.
method Introduces a power-weighted noncentral complex Gaussian distribution on the complex plane.
result Consistently outperforms conventional distributions in log-likelihood for speech power spectra.
We compute the bridge spectra of cables of 2-bridge knots. We also give some results about bridge spectra and distance of Montesinos knots.
It is generally recognized that economical systems, and more in general complex systems, are characterized by power law distributions. Sometime, these distributions show a changing of the slope in the tail so that, more appropriately, they show a multi-power law behavior. We present a method to derive analytically a tw…
Develops a simple model to understand learning curves for arbitrary power laws.
problem Lack of theoretical understanding of scaling laws in machine learning.
method Analyzes a toy model to determine if learning curves are universal or depend on data distribution.
result Determines that learning curves can exhibit n−β for arbitrary power β>0. Diffusion models generalize better with hierarchical data structure and regularization.
problem Understanding generalization in diffusion models with finite data.
method Analyzing diffusion models through data covariance spectra and developing a theoretical framework based on linear neural networks.
result Generalization in diffusion models improves with hierarchical data structure and regularization.
Study compares exponential and power-law kernels in modeling high-frequency trading data.
problem Modeling high-frequency trading data with specific kernel types.
method Proposes and analyzes two bivariate Hawkes processes with exponential and power-law kernels.
result Identifies strengths and limitations of exponential and power-law kernels for high-frequency trading data.
We introduce a new statistical tool (the TP-statistic and TE-statistic) designed specifically to compare the behavior of the sample tail of distributions with power-law and exponential tails as a function of the lower threshold u. One important property of these statistics is that they converge to zero for power laws o…
New metrics compare rational spectra using optimal transport.
problem Comparing rational spectra efficiently and accurately.
method Optimal transport and linear-systems theory.
result Established connection to Wasserstein distance.
Study uses OT to simulate markets, revealing power-law returns are driven by informational effect.
problem Reproduce power-law returns in financial markets using realistic simulations.
method Constructed artificial markets, used optimal transport (OT) to measure similarity, incrementally introduced behavioral components.
result Informational effect of prices is dominant in reproducing power-law returns, and multiple components interact synergistically.
Sparse-mode DMD disambiguates local and global modes in spatiotemporal data.
problem Disambiguating local and global modes in spatiotemporal data.
method Sparse-mode DMD with sparsity-promoting regularization.
result Explicitly constructs discrete and continuous spectra.