Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

1122 · Jun 200219922001200920172026
20 results for early-time

Early neural networks can be simplified to linear models, revealing surprising simplicity.

problem Complexity of neural network learning dynamics.
method Formal proof and empirical verification of early-time learning dynamics of neural networks.
result Early learning dynamics of neural networks can be approximated by simple linear models.

Deep learning calibrates CO2 storage formations from seismic and well data.

problem Uncertainty in CO2 storage formation properties.
method Two deep learning models for well and seismic data, integrated into MCMC history matching.
result Significant uncertainty reduction in key parameters and accurate CO2 plume predictions.

An exclusion particle model is considered as a highly simplified model of a limit order market. Its price behavior reproduces the well known crossover from over-diffusion (Hurst exponent H>1/2) to diffusion (H=1/2) when the time horizon is increased, provided that orders are allowed to be canceled. For early times a ma…

2002-06-24abs ↗pdf ↗

The study analyzes sharpness dynamics in neural networks, revealing mechanisms and conditions.

problem Understanding sharpness in neural network training.
method Fixed point analysis and edge of stability analysis in a simplified 2-layer linear network.
result Reveals mechanisms behind sharpness trends, conditions for edge of stability, and a period-doubling route to chaos.

Enhances the Bishop-Gromov theorem for curved spaces, especially at late times.

problem The Bishop-Gromov theorem's volume growth upperbound is often too loose, especially at late times.
method Identified and quantified the effect of shear, using higher curvature invariants to improve the upperbound.
result Tighter upper bounds on late-time growth rates of geodesic balls in homogeneous spaces with non-positive sectional curvature.

Systems with long-range persistence and memory are shown to exhibit different precursory as well as recovery patterns in response to shocks of exogeneous versus endogeneous origins. By endogeneous, we envision either fluctuations resulting from an underlying chaotic dynamics or from a stochastic forcing origin which ma…

2002-06-05abs ↗pdf ↗

Diffusion models generalize well until a threshold is reached, preventing memorization.

problem Understanding why diffusion models don't memorize training data.
method Investigation of training dynamics and two timescales: τgenτ_\mathrm{gen} and τmemτ_\mathrm{mem}.
result The threshold τmemτ_\mathrm{mem} increases linearly with training set size nn, preventing memorization.

The paper introduces a novel method for training neural network Stein critics with staged L2L^2-regularization.

problem Learning to differentiate model distributions from observed data in high-dimensional settings.
method Developed a novel staging procedure for L2L^2 regularization over training time, leveraging the advantages of highly-regularized training at early times.
result Theoretical guarantees and empirical validation show that the method improves the approximation of the training dynamic by the kernel optimization, leading to faster convergence and better performance.

Gradient descent aligns neural feature matrices with pre-activation tangent features.

problem Understanding neural feature learning mechanisms.
method Analytical proof of alignment between weight matrices and pre-activation tangent features.
result Derivative alignment occurs almost surely in high-dimensional settings.

Early time series classification (eTSC) is the problem of classifying a time series after as few measurements as possible with the highest possible accuracy. The most critical issue of any eTSC method is to decide when enough data of a time series has been seen to take a decision: Waiting for more data points usually m…

2019-08-09abs ↗pdf ↗

Study geodesics on random hyperbolic surfaces, finding variance similar to prime number theory.

problem Distribution of closed geodesics on random hyperbolic surfaces.
method Investigate random variable counting geodesics with norms in short intervals, comparing to prime number theory.
result Establishes variance of geodesic counting function is asymptotic to \(2H \log X\).

We introduce a new weight-decay scaling rule to maintain sublayer gains across different widths in modern scale-invariant architectures.

problem In modern scale-invariant architectures, training quickly enters a steady state where normalization layers create backward scale sensitivity, degrading learning-rate transfer.
method We introduce a weight-decay scaling rule for AdamW that preserves sublayer gain across widths by equalizing the effective learning rate.
result Our empirical weight-decay scaling rule λ2dλ_2\propto \sqrt{d} approximately keeps sublayer gains width invariant, enabling zero-shot transfer of learning rate and weight decay.

Logarithmic corrections to Price's law near black hole event horizon.

problem Failure of smooth null infinity in black hole spacetimes.
method Analyzing linear wave equation on Schwarzschild background with specific initial conditions.
result Leading-order asymptotics of solutions near future null infinity and event horizon are logarithmically modified.