LLT transforms time series features based on linear laws.
problem Classifying univariate and multivariate time series.
method Time-delay embedding, spectral decomposition, and feature transformation.
result Transformed features improve classification accuracy.
Scaling laws in linear regression explain model performance improvements with size and data.
problem Disagreement between empirical neural scaling laws and conventional wisdom on variance error.
method Infinite dimensional linear regression setup, one-pass SGD, Gaussian prior, power-law spectrum.
result Variance error is dominated by other errors, disappearing from the bound due to SGD's implicit regularization.
Improved scaling laws in linear regression using data reuse.
problem Sustainability of neural scaling laws when running out of new data.
method Data reuse in multi-pass stochastic gradient descent (multi-pass SGD) for M-dimensional linear models trained on N data with sketched features. result Multi-pass SGD achieves a test error of Θ(M1−b+L(1−b)/a) with L>N, improving scaling laws in data-constrained regimes. Paper finds linear laws in Bitcoin price changes, aiding anomaly detection.
problem Detecting anomalies in Bitcoin price changes.
method Time embedding of autocorrelation function, binary series generation, stepped time windows.
result Linear laws became more complex before major market events, suggesting price manipulation.
SignSGD outperforms SGD in linear regression with optimal scaling laws under PLRF model.
problem Improving linear regression performance with signSGD under power-law random features.
method Analysis of signSGD risk under PLRF model, comparison with SGD, identification of unique effects.
result SignSGD can have a steeper compute-optimal slope than SGD in noisy regimes, especially with WSD schedule.
We study higher-order conservation laws of the non-linearizable elliptic Poisson equation ∂z∂zˉ∂2u=−f(u) as elements of the characteristic cohomology of the associated exterior differential system. The theory of characteristic cohomology determines a normal form for diffe…
LLT-ECG classifies ECG signals without backpropagation using linear laws.
problem ECG signal classification efficiency and verifiability.
method Forward linear law identification for ECG signal features.
result State-of-the-art performance on real-world ECG datasets.
This work extends the scaling law to multiple and kernel regression, challenging traditional machine learning principles.
problem Challenging traditional machine learning wisdom with scaling law in large practical models.
method Demonstrates the scaling law in multiple and kernel regression settings.
result The scaling law extends to multiple and kernel regression, providing deeper insights into LLMs.
This work analyzes neural scaling laws using power-law data spectra and derives analytical expressions for generalization error.
problem Understanding how neural network performance scales with key factors like data size and model complexity.
method Statistical mechanics techniques applied to one-pass stochastic gradient descent in a student-teacher framework.
result Derivation of analytical expressions for generalization error under power-law data spectra and identification of conditions for power-law scaling.
This work studies scaling laws for low-precision training in high-dimensional linear regression.
problem Optimizing trade-off between model quality and training costs in high-dimensional linear regression.
method Theoretical study of scaling laws for low-precision training within a high-dimensional sketched linear regression framework, analyzing multiplicative and additive quantization.
result Multiplicative quantization maintains full-precision model size, while additive quantization reduces effective model size.
This paper extends, to a class of systems of semi-linear hyperbolic second order PDEs in three variables, the geometric study of a single nonlinear hyperbolic PDE in the plane as presented in [Anderson I.M., Kamran N., Duke Math. J. 87 (1997), 265-319]. The constrained variational bi-complex is introduced and used to d…
I consider the existence and structure of conservation laws for the general class of evolutionary scalar second-order differential equations with parabolic symbol. First I calculate the linearized characteristic cohomology for such equations. This provides an auxiliary differential equation satisfied by the conservatio…
This work explains scaling laws as redundancy laws in deep learning.
problem The mathematical origins of scaling laws in deep learning models remain unclear.
method Kernel regression and analysis of data covariance spectra.
result Scaling laws can be explained as redundancy laws, revealing the learning curve's slope depends on data redundancy.
SAEs struggle with curved activation manifolds, revealing layer-dependent scaling laws.
problem Sparse autoencoders' reconstruction error varies across layers, not fitting existing scaling laws.
method Cross-layer study of 844 SAE checkpoints, fitting and regressing on manifold geometry.
result Manifold geometry predicts layer-dependent width exponents in SAEs, with transferable coefficients.
Noether's First Theorem yields conservation laws for Lagrangians with a variational symmetry group. The explicit formulae for the laws are well known and the symmetry group is known to act on the linear space generated by the conservation laws. In recent work the authors showed the mathematical structure behind both th…
Advocates a local feedback approach for RL in unknown systems.
problem Finding optimal feedback laws in unknown nonlinear dynamical systems.
method Searches over a local feedback representation consisting of an open-loop sequence and an optimal linear feedback law.
result Results in highly efficient training and superior performance compared to global methods.
New findings on how certain functionals behave in random variable spaces.
problem Understanding when law-invariant convex functionals simplify to the mean.
method Analyzing a broad class of random variable spaces and mild semicontinuity assumptions.
result The expectation functional is the only law-invariant convex functional that collapses to the mean under certain conditions.
Study on KRR with power-law data, showing better sample complexity.
problem High-dimensional kernel ridge regression with anisotropic power-law covariance.
method Explicit characterization of kernel spectrum and asymptotic analysis of excess risk.
result Sample complexity is governed by effective dimension, not ambient dimension.
We propose a geometric correspondence between (a) linearly degenerate systems of conservation laws with rectilinear rarefaction curves and (b) congruences of lines in projective space whose developable surfaces are planar pencils of lines. We prove that in projective 4-space such congruences are necessarily linear. Bas…
The study uncovers universality laws for Gaussian mixtures in generalized linear models.
problem Understanding the asymptotic behavior of estimators in Gaussian mixture models.
method Investigates the asymptotic joint statistics of generalized linear estimators from empirical risk minimization and Gibbs sampling.
result Characterizes conditions under which the joint statistics depend only on means and covariances of class conditional features.
Study on neural scaling laws for solving linear systems in-context.
problem Theoretical guarantees for solving linear systems using a linear transformer architecture.
method Neural scaling laws and task diversity for in-domain and out-of-domain generalization.
result Novel notion of task diversity for necessary and sufficient condition of generalization under task shifts.
Decomposing market impact into diffusive components
problem Market impact scaling
method Decomposing impact into realized and counterfactual returns
result Implication of square-root law in information-neutral regime
Tensor networks help learn complex physical laws from data.
problem Identifying non-linear dynamical laws from complex physical systems.
method Tensor network parameterizations and rank-adaptive optimization.
result Optimal tensor network models can be learned from data.
ALT improves TSC by capturing complex patterns in time series data.
problem Challenges in traditional TSC methods with time series complexity and variability.
method ALT incorporates variable-length shifted time windows to enhance LLT for better feature representation.
result ALT achieves state-of-the-art performance with few hyperparameters.
Auto-regressive conditionally heteroskedastic (ARCH) family models are still used, by practitioners in business and economic policy making, as a conditional volatility forecasting models. Furthermore ARCH models still are attracting an interest of the researchers. In this contribution we consider the well known GARCH(1…
We succeed in writing 2-dimensional conformally invariant non-linear elliptic PDE (harmonic map equation, prescribed mean curvature equations...etc) in divergence form. This divergence free quantities generalize to target manifolds without symmetries the well known conservation laws for harmonic maps into homogeneous s…
Study uncovers scaling laws and spectral properties of shallow neural networks.
problem Understanding scaling laws and spectral properties of shallow neural networks.
method Leveraging connections with matrix compressed sensing and LASSO, derived a phase diagram for excess risk.
result Uncovered crossovers between scaling regimes and plateau behaviors, validated empirical observations.
Improves machine learning models by incorporating physical laws into feature maps.
problem Lack of model interpretability in classical machine learning approaches.
method Physics-informed feature maps constructed from physical laws and dimensional analysis.
result Enhanced model interpretability and potential discovery of new physical equations.
Noether's Theorem yields conservation laws for a Lagrangian with a variational symmetry group. The explicit formulae for the laws are well known and the symmetry group is known to act on the linear space generated by the conservation laws. The aim of this paper is to explain the mathematical structure of both the Euler…
Conservation laws vanishing along characteristic directions of a given system of PDEs are known as characteristic conservation laws, or characteristic integrals. In 2D, they play an important role in the theory of Darboux-integrable equations. In this paper we discuss characteristic integrals in 3D and demonstrate that…
The paper explores neural scaling laws for deep operator networks, offering a theoretical foundation.
problem Understanding neural scaling laws in deep operator networks.
method Theoretical analysis of approximation and generalization errors.
result Established a theoretical framework to quantify neural scaling laws for deep operator networks.
Gradient descent dynamics studied for DEQs in linear and single-index models.
problem Understanding gradient descent dynamics for DEQs.
method Rigorously studied gradient descent dynamics for DEQs in linear and single-index models.
result Gradient descent converges to a global minimizer for linear DEQs and single-index models.
Researchers use quantum chaos and RMT to analyze turbulence, revealing unique scaling laws.
problem Understanding the statistical structure and scaling laws of turbulence.
method Applied tools from quantum chaos and Random Matrix Theory to analyze turbulence datasets.
result Turbulence Gram matrices exhibit power-law scalings distinct from classical chaos and random data.
We present a deep learning framework for quantifying and propagating uncertainty in systems governed by non-linear differential equations using physics-informed neural networks. Specifically, we employ latent variable models to construct probabilistic representations for the system states, and put forth an adversarial …
Deep learning improves image reconstruction, but scaling up training sets doesn't significantly boost performance.
problem Understanding the impact of training set size on deep learning image reconstruction.
method Empirical study and analytical characterization of performance scaling laws.
result Scaling up training set size does not significantly improve reconstruction quality for deep learning.
Consider a Riemannian metric on two-torus. We prove that the question of existence of polynomial first integrals leads naturally to a remarkable system of quasi-linear equations which turns out to be a Rich system of conservation laws. This reduces the question of integrability to the question of existence of smooth (q…
The exterior differential system for constant mean curvature (CMC) surfaces in a 3-dimensional space form is an elliptic Monge-Ampere system defined on the unit tangent bundle. We determine the infinite sequence of higher-order symmetries and conservation laws via an enhanced prolongation modelled on a loop algebra val…
Investigates ways to train larger models with fewer resources, finding that test loss depends only on the actual number of trainable parameters.
problem Training larger models for cheaper under hardware constraints.
method Emulates an increase in effective parameters using frozen random parameters or fast structured transforms.
result Scaling laws cannot be deceived by spurious parameters; test loss depends only on the actual number of trainable parameters.
We apply an asymmetric version of Kirman's herding model to volatile financial markets. In the relation between returns and agent concentration we use the square root law proposed by Zhang. This can be derived by extending the idea of a critical mean field theory suggested by Plerou et al. We show that this model is eq…
Unified analysis of parameter norms in overparameterized linear models, revealing scaling laws and thresholds.
problem Understanding the scaling of parameter norms in overparameterized linear models.
method Simple dual-ray analysis revealing competition between signal spike and bulk of null coordinates.
result Unified closed-form predictions for parameter norm scaling, including elbow and threshold laws.
We study the residual bootstrap (RB) method in the context of high-dimensional linear regression. Specifically, we analyze the distributional approximation of linear contrasts c⊤(β^ρ−β), where β^ρ is a ridge-regression estimator. When regression coefficients are estimated via least squares, classical…
New law predicts first extinction in resampling processes.
problem Intractable extinction times in resampling processes.
method Modeling multinomial updates as independent square-root diffusions.
result Closed-form law for first-extinction time with linear cost.
New framework finds more efficient linear layers over structured matrices.
problem Efficient alternatives for dense linear layers in neural networks.
method Unified framework searching over all linear operators, developing a taxonomy based on computational and algebraic properties.
result BTT-MoE provides substantial compute-efficiency gains over dense layers and standard MoE.
New insights into how depth and width affect in-context learning in deep models.
problem Understanding how various resources impact in-context learning in deep models.
method Analyzed linear regression in a deep linear self-attention model, varying resources like depth, width, context length, and training steps.
result Increasing depth improves in-context learning even at infinite context length, contrary to previous findings.
We report analytical results for the development of the viscous fingering instability in a cylindrical Hele-Shaw cell of radius a and thickness b. We derive a generalized version of Darcy's law in such cylindrical background, and find it recovers the usual Darcy's law for flow in flat, rectangular cells, with correctio…
Among econophysics investigations, studies of religious groups have been of interest. On one hand, the present paper concerns the Antoinist community financial reports, - a community which appeared at the end of the 19-th century in Belgium. Several growth-decay regimes have been previously found over different time sp…
I consider the geometry of the general class of scalar 2nd-order differential equations with parabolic symbol, including non-linear and non-evolutionary parabolic equations. After defining the appropriate G-structure to model parabolic equations, I apply Cartan techniques to determine local geometric invariants (quan…
Dimension reduction is the process of embedding high-dimensional data into a lower dimensional space to facilitate its analysis. In the Euclidean setting, one fundamental technique for dimension reduction is to apply a random linear map to the data. This dimension reduction procedure succeeds when it preserves certain …