Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

94189283377 · Jun 202019922001200920172026
48 results for linear laws

Scaling laws in linear regression explain model performance improvements with size and data.

problem Disagreement between empirical neural scaling laws and conventional wisdom on variance error.
method Infinite dimensional linear regression setup, one-pass SGD, Gaussian prior, power-law spectrum.
result Variance error is dominated by other errors, disappearing from the bound due to SGD's implicit regularization.

Improved scaling laws in linear regression using data reuse.

problem Sustainability of neural scaling laws when running out of new data.
method Data reuse in multi-pass stochastic gradient descent (multi-pass SGD) for MM-dimensional linear models trained on NN data with sketched features.
result Multi-pass SGD achieves a test error of Θ(M1b+L(1b)/a)Θ(M^{1-b} + L^{(1-b)/a}) with L>NL>N, improving scaling laws in data-constrained regimes.

Paper finds linear laws in Bitcoin price changes, aiding anomaly detection.

problem Detecting anomalies in Bitcoin price changes.
method Time embedding of autocorrelation function, binary series generation, stepped time windows.
result Linear laws became more complex before major market events, suggesting price manipulation.

SignSGD outperforms SGD in linear regression with optimal scaling laws under PLRF model.

problem Improving linear regression performance with signSGD under power-law random features.
method Analysis of signSGD risk under PLRF model, comparison with SGD, identification of unique effects.
result SignSGD can have a steeper compute-optimal slope than SGD in noisy regimes, especially with WSD schedule.

We study higher-order conservation laws of the non-linearizable elliptic Poisson equation 2uzzˉ=f(u) \frac{{\partial}^2 u}{\partial z \partial \bar{z}} = -f(u) as elements of the characteristic cohomology of the associated exterior differential system. The theory of characteristic cohomology determines a normal form for diffe…

2009-06-17abs ↗pdf ↗

This work extends the scaling law to multiple and kernel regression, challenging traditional machine learning principles.

problem Challenging traditional machine learning wisdom with scaling law in large practical models.
method Demonstrates the scaling law in multiple and kernel regression settings.
result The scaling law extends to multiple and kernel regression, providing deeper insights into LLMs.

This work analyzes neural scaling laws using power-law data spectra and derives analytical expressions for generalization error.

problem Understanding how neural network performance scales with key factors like data size and model complexity.
method Statistical mechanics techniques applied to one-pass stochastic gradient descent in a student-teacher framework.
result Derivation of analytical expressions for generalization error under power-law data spectra and identification of conditions for power-law scaling.

This work studies scaling laws for low-precision training in high-dimensional linear regression.

problem Optimizing trade-off between model quality and training costs in high-dimensional linear regression.
method Theoretical study of scaling laws for low-precision training within a high-dimensional sketched linear regression framework, analyzing multiplicative and additive quantization.
result Multiplicative quantization maintains full-precision model size, while additive quantization reduces effective model size.

SAEs struggle with curved activation manifolds, revealing layer-dependent scaling laws.

problem Sparse autoencoders' reconstruction error varies across layers, not fitting existing scaling laws.
method Cross-layer study of 844 SAE checkpoints, fitting and regressing on manifold geometry.
result Manifold geometry predicts layer-dependent width exponents in SAEs, with transferable coefficients.

Noether's First Theorem yields conservation laws for Lagrangians with a variational symmetry group. The explicit formulae for the laws are well known and the symmetry group is known to act on the linear space generated by the conservation laws. In recent work the authors showed the mathematical structure behind both th…

2011-06-20abs ↗pdf ↗

Advocates a local feedback approach for RL in unknown systems.

problem Finding optimal feedback laws in unknown nonlinear dynamical systems.
method Searches over a local feedback representation consisting of an open-loop sequence and an optimal linear feedback law.
result Results in highly efficient training and superior performance compared to global methods.

New findings on how certain functionals behave in random variable spaces.

problem Understanding when law-invariant convex functionals simplify to the mean.
method Analyzing a broad class of random variable spaces and mild semicontinuity assumptions.
result The expectation functional is the only law-invariant convex functional that collapses to the mean under certain conditions.

The study uncovers universality laws for Gaussian mixtures in generalized linear models.

problem Understanding the asymptotic behavior of estimators in Gaussian mixture models.
method Investigates the asymptotic joint statistics of generalized linear estimators from empirical risk minimization and Gibbs sampling.
result Characterizes conditions under which the joint statistics depend only on means and covariances of class conditional features.

Study on neural scaling laws for solving linear systems in-context.

problem Theoretical guarantees for solving linear systems using a linear transformer architecture.
method Neural scaling laws and task diversity for in-domain and out-of-domain generalization.
result Novel notion of task diversity for necessary and sufficient condition of generalization under task shifts.

ALT improves TSC by capturing complex patterns in time series data.

problem Challenges in traditional TSC methods with time series complexity and variability.
method ALT incorporates variable-length shifted time windows to enhance LLT for better feature representation.
result ALT achieves state-of-the-art performance with few hyperparameters.

Auto-regressive conditionally heteroskedastic (ARCH) family models are still used, by practitioners in business and economic policy making, as a conditional volatility forecasting models. Furthermore ARCH models still are attracting an interest of the researchers. In this contribution we consider the well known GARCH(1…

2014-12-19abs ↗pdf ↗

We succeed in writing 2-dimensional conformally invariant non-linear elliptic PDE (harmonic map equation, prescribed mean curvature equations...etc) in divergence form. This divergence free quantities generalize to target manifolds without symmetries the well known conservation laws for harmonic maps into homogeneous s…

2006-03-15abs ↗pdf ↗

Study uncovers scaling laws and spectral properties of shallow neural networks.

problem Understanding scaling laws and spectral properties of shallow neural networks.
method Leveraging connections with matrix compressed sensing and LASSO, derived a phase diagram for excess risk.
result Uncovered crossovers between scaling regimes and plateau behaviors, validated empirical observations.

Improves machine learning models by incorporating physical laws into feature maps.

problem Lack of model interpretability in classical machine learning approaches.
method Physics-informed feature maps constructed from physical laws and dimensional analysis.
result Enhanced model interpretability and potential discovery of new physical equations.

Noether's Theorem yields conservation laws for a Lagrangian with a variational symmetry group. The explicit formulae for the laws are well known and the symmetry group is known to act on the linear space generated by the conservation laws. The aim of this paper is to explain the mathematical structure of both the Euler…

2010-06-23abs ↗pdf ↗

Conservation laws vanishing along characteristic directions of a given system of PDEs are known as characteristic conservation laws, or characteristic integrals. In 2D, they play an important role in the theory of Darboux-integrable equations. In this paper we discuss characteristic integrals in 3D and demonstrate that…

2013-12-18abs ↗pdf ↗

The paper explores neural scaling laws for deep operator networks, offering a theoretical foundation.

problem Understanding neural scaling laws in deep operator networks.
method Theoretical analysis of approximation and generalization errors.
result Established a theoretical framework to quantify neural scaling laws for deep operator networks.

Researchers use quantum chaos and RMT to analyze turbulence, revealing unique scaling laws.

problem Understanding the statistical structure and scaling laws of turbulence.
method Applied tools from quantum chaos and Random Matrix Theory to analyze turbulence datasets.
result Turbulence Gram matrices exhibit power-law scalings distinct from classical chaos and random data.

Deep learning improves image reconstruction, but scaling up training sets doesn't significantly boost performance.

problem Understanding the impact of training set size on deep learning image reconstruction.
method Empirical study and analytical characterization of performance scaling laws.
result Scaling up training set size does not significantly improve reconstruction quality for deep learning.

Consider a Riemannian metric on two-torus. We prove that the question of existence of polynomial first integrals leads naturally to a remarkable system of quasi-linear equations which turns out to be a Rich system of conservation laws. This reduces the question of integrability to the question of existence of smooth (q…

2009-07-29abs ↗pdf ↗

Investigates ways to train larger models with fewer resources, finding that test loss depends only on the actual number of trainable parameters.

problem Training larger models for cheaper under hardware constraints.
method Emulates an increase in effective parameters using frozen random parameters or fast structured transforms.
result Scaling laws cannot be deceived by spurious parameters; test loss depends only on the actual number of trainable parameters.

We apply an asymmetric version of Kirman's herding model to volatile financial markets. In the relation between returns and agent concentration we use the square root law proposed by Zhang. This can be derived by extending the idea of a critical mean field theory suggested by Plerou et al. We show that this model is eq…

2005-08-12abs ↗pdf ↗

Unified analysis of parameter norms in overparameterized linear models, revealing scaling laws and thresholds.

problem Understanding the scaling of parameter norms in overparameterized linear models.
method Simple dual-ray analysis revealing competition between signal spike and bulk of null coordinates.
result Unified closed-form predictions for parameter norm scaling, including elbow and threshold laws.

We study the residual bootstrap (RB) method in the context of high-dimensional linear regression. Specifically, we analyze the distributional approximation of linear contrasts c(β^ρβ)c^{\top} (\hatβ_ρ-β), where β^ρ\hatβ_ρ is a ridge-regression estimator. When regression coefficients are estimated via least squares, classical…

2016-07-04abs ↗pdf ↗

New framework finds more efficient linear layers over structured matrices.

problem Efficient alternatives for dense linear layers in neural networks.
method Unified framework searching over all linear operators, developing a taxonomy based on computational and algebraic properties.
result BTT-MoE provides substantial compute-efficiency gains over dense layers and standard MoE.

New insights into how depth and width affect in-context learning in deep models.

problem Understanding how various resources impact in-context learning in deep models.
method Analyzed linear regression in a deep linear self-attention model, varying resources like depth, width, context length, and training steps.
result Increasing depth improves in-context learning even at infinite context length, contrary to previous findings.

We report analytical results for the development of the viscous fingering instability in a cylindrical Hele-Shaw cell of radius a and thickness b. We derive a generalized version of Darcy's law in such cylindrical background, and find it recovers the usual Darcy's law for flow in flat, rectangular cells, with correctio…

2002-01-31abs ↗pdf ↗

Among econophysics investigations, studies of religious groups have been of interest. On one hand, the present paper concerns the Antoinist community financial reports, - a community which appeared at the end of the 19-th century in Belgium. Several growth-decay regimes have been previously found over different time sp…

2012-08-29abs ↗pdf ↗

Dimension reduction is the process of embedding high-dimensional data into a lower dimensional space to facilitate its analysis. In the Euclidean setting, one fundamental technique for dimension reduction is to apply a random linear map to the data. This dimension reduction procedure succeeds when it preserves certain …

2015-11-30abs ↗pdf ↗