Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

12.5%25.0%37.5%50.0% · Dec 199319922001200920172026
48 results for moment decay parameters

In several recently proposed stochastic optimization methods (e.g. RMSProp, Adam, Adadelta), parameter updates are scaled by the inverse square roots of exponential moving averages of squared past gradients. Maintaining these per-parameter second-moment estimators requires memory equal to the number of parameters. For …

2018-04-11abs ↗pdf ↗

AdamNX improves Adam's stability by adjusting its learning rate.

problem Adam's tendency to converge to non-flat minima in large-scale models.
method Proposes a novel exponential decay mechanism for Adam's second-order moment estimate.
result AdamNX outperforms Adam and its variants in stability and performance.

Adam optimization algorithm can have non-zero average regret under certain conditions.

problem Non-zero average regret in Adam optimization algorithm.
method Used a three-periodic sequence of linear functions on [-1,1] with slopes c, -1, -1, and analyzed Adam variants.
result Adam optimization algorithm can have non-zero average regret under certain conditions.

Weibull weight-scale parameter λλ evolves during AdamW training, with alignment, injection, and decay forces driving its growth and relaxation.

problem Understanding the evolution of the Weibull weight-scale parameter λλ during AdamW training.
method Deriving a leading-order three-force decomposition of the squared weight norm from AdamW updates.
result The alignment force dominates the rise phase, contributing 88-94% of the absolute force budget across four random seeds.

Study examines robust regression in high dimensions with heavy-tailed data.

problem Analyzing robust regression in high-dimensional settings with heavy-tailed data.
method Sharp asymptotic characterisation of M-estimators and ridge regression in elliptical distributions.
result Ridge regression is optimal and universal for finite second moments but can decay faster without them.

Suppose kk centers are fit to mm points by heuristically minimizing the kk-means cost; what is the corresponding fit over the source distribution? This question is resolved here for distributions with p4p\geq 4 bounded moments; in particular, the difference between the sample cost and distribution cost decays with $…

2013-11-08abs ↗pdf ↗

We propose a simple stochastic volatility model which is analytically tractable, very easy to simulate and which captures some relevant stylized facts of financial assets, including scaling properties. In particular, the model displays a crossover in the log-return distribution from power-law tails (small time) to a Ga…

2010-06-01abs ↗pdf ↗

Study on U-statistics with heavy-tailed samples, providing tail bounds and LDP.

problem Deviation of U-statistics with heavy-tailed samples.
method Exponential tail bounds and Large Deviation Principle (LDP) for U-statistics.
result Obtained an exponential upper bound for U-statistics tail decay, showing two regions of decay.

Study on massless Vlasov equation on Reissner-Nordström spacetimes, showing decay rates and non-decay phenomena.

problem Analyzing decay and non-decay rates of solutions to the massless Vlasov equation on Reissner-Nordström spacetimes.
method Quantitative analysis of geodesic flow and comparison to wave equation instability results.
result Exponential decay rates in subextremal cases and polynomial rates in extremal cases, with non-decay of transversal derivatives in extremal cases.

Regularization in the optimization of deep neural networks is often critical to avoid undesirable over-fitting leading to better generalization of model. One of the most popular regularization algorithms is to impose L-2 penalty on the model parameters resulting in the decay of parameters, called weight-decay, and the …

2019-07-21abs ↗pdf ↗

Novel Adam-family method with decoupled weight decay for training neural networks.

problem Training nonsmooth neural networks with weight decay.
method Proposes a novel Adam-family method with decoupled weight decay, establishing convergence properties and demonstrating superior performance.
result Asymptotically approximates SGD and enhances generalization performance.

Nonlinear SGD achieves high-probability rates in non-convex optimization with heavy-tailed noise.

problem Optimization in non-convex problems with heavy-tailed noise.
method General nonlinear framework for SGD, including symmetrization techniques.
result Achieves O~(t1/2)\widetilde{\mathcal{O}}(t^{-1/2}) rate for heavy-tailed noise.

New framework for Adam-type algorithms with constant β1, improving regret analysis.

problem Theoretical vs. practical use of Adam and variants with constant β1.
method Proposed a novel framework to derive optimal, data-dependent regret bounds with constant β1.
result Optimal, data-dependent regret bounds with constant β1 are achievable without further assumptions.

We propose a method of moments (MoM) algorithm for training large-scale implicit generative models. Moment estimation in this setting encounters two problems: it is often difficult to define the millions of moments needed to learn the model parameters, and it is hard to determine which properties are useful when specif…

2018-06-28abs ↗pdf ↗

New bounds on generalization error using information density moments.

problem Bounding the generalization error of randomized learning algorithms.
method Derives bounds on average and tail probabilities of generalization error using mth central moments of the information density.
result Explicit bounds on generalization error are derived, showing better dependence on confidence level with higher-order information density moments.

Weight decay stabilizes training dynamics by slowing progressive sharpening.

problem Understanding how weight decay affects training stability in deep learning models.
method Analyzing weight decay effects at the Edge of Stability, developing a mathematical framework.
result Weight decay dampens oscillations and stabilizes sharpness in CNNs, causing a phase transition in MLPs.

Using a relationship between the moments of the probability distribution of times between the two consecutive trades (intertrade time distribution) and the moments of the distribution of a daily number of trades we show, that the underlying point process is essentially non-markovian. A detailed analysis of all trades i…

2003-03-12abs ↗pdf ↗

This paper investigates the effectiveness of decoupled weight decay at the start of training.

problem The traditional approach to weight decay is not effective throughout training.
method The authors investigate decoupled weight decay, applying it only at the start of training.
result Applying weight decay only at the start of training stabilizes network weights and improves performance.

A new method for estimating causal parameters from observables reduces the need for finite moment conditions.

problem Estimating causal parameters from observational data with unknown or infinite moment conditions.
method Variational Method of Moments (VMM) for a general class of estimators, including kernel and neural net-based methods.
result VMM estimators are consistent, asymptotically normal, and semiparametrically efficient.

This article investigates parameter estimation of affine term structure models by means of the generalized method of moments. Exact moments of the affine latent process as well as of the yields are obtained by using results derived for p-polynomial processes. Then the generalized method of moments, combined with Quasi-…

2015-08-07abs ↗pdf ↗

The non-gaussianity of processes observed in financial markets and relatively good performance of gaussian models can be reconciled by replacing the Brownian motion with Levy processes whose Levy densities decay as exp(-lambda|x|) or faster, where lambda>0 is large. This leads to asymptotic pricing models. The leading …

2002-12-11abs ↗pdf ↗

Study on martingale property and moment explosions in signature volatility models.

problem Analyzing the martingale property and moment explosions in signature volatility models.
method Fine analysis of the explosion time of a signature stochastic differential equation.
result The price process is a true martingale if and only if the order of the linear form is odd and a correlation parameter is negative.

Study on eigenvalue distribution of correlated time series, showing deformation of Marchenko-Pastur distribution.

problem Eigenvalue distribution of Wishart matrix with temporal correlation.
method Analysis of moments and convergence to deformed Marchenko-Pastur distribution for Gaussian process with temporal correlation.
result Eigenvalue distribution converges to deformed Marchenko-Pastur distribution with longer tail and higher peak.

We introduce a new weight-decay scaling rule to maintain sublayer gains across different widths in modern scale-invariant architectures.

problem In modern scale-invariant architectures, training quickly enters a steady state where normalization layers create backward scale sensitivity, degrading learning-rate transfer.
method We introduce a weight-decay scaling rule for AdamW that preserves sublayer gain across widths by equalizing the effective learning rate.
result Our empirical weight-decay scaling rule λ2dλ_2\propto \sqrt{d} approximately keeps sublayer gains width invariant, enabling zero-shot transfer of learning rate and weight decay.

This work develops efficient methods for computing moments of Gaussian mixtures.

problem Efficient computation of moments for Gaussian mixtures with large dimensions.
method Theory and numerical methods for implicit computations with moment tensors of Gaussian mixtures.
result Reduced computational and storage costs for moment tensors of Gaussian mixtures.

The paper analyzes Teukolsky equations on Kerr backgrounds, proving boundedness and decay of solutions.

problem Analyzing boundedness and decay of solutions to Teukolsky equations on Kerr backgrounds.
method Frequency space analysis of transformed Teukolsky equations on Kerr backgrounds.
result Fixed frequency solutions remain bounded and decay in time for subextremal Kerr backgrounds.

Stochastic Kronecker graphs supply a parsimonious model for large sparse real world graphs. They can specify the distribution of a large random graph using only three or four parameters. Those parameters have however proved difficult to choose in specific applications. This article looks at method of moments estimators…

2011-06-08abs ↗pdf ↗

Gradient descent outperforms ridge regression under certain covariance matrix decay conditions.

problem Comparing the performance of gradient descent and ridge regression in linear models.
method Investigated gradient descent and ridge regression for linear regression with random isotropic ground truth.
result Gradient descent outperforms ridge regression under specific covariance matrix decay conditions.

Mixture modeling is a general technique for making any simple model more expressive through weighted combination. This generality and simplicity in part explains the success of the Expectation Maximization (EM) algorithm, in which updates are easy to derive for a wide class of mixture models. However, the likelihood of…

2016-03-28abs ↗pdf ↗

We present a detailed analysis of \emph{observable} moments based parameter estimators for the Heston SDEs jointly driving the rate of returns RtR_t and the squared volatilities VtV_t. Since volatilities are not directly observable, our parameter estimators are constructed from empirical moments of realized volatilitie…

2017-06-14abs ↗pdf ↗

The paper analyzes multivariate Hawkes processes and their induced population processes.

problem Analyzing the time-dependent joint probability distribution of multivariate Hawkes processes.
method Exact and asymptotic analysis of general multivariate Hawkes processes and their induced population processes.
result Full characterization of the time-dependent joint transform of the multivariate population process and its intensity process.

Empower efficient representation of distributions through moment-preserving methods.

problem Representing high-dimensional probability measures efficiently and accurately.
method Empower efficient representation of distributions through moment-preserving methods.
result Empowers efficient and accurate representation of high-dimensional probability measures.

In this work we prove an universality result regarding the equidistribution of zeros of random holomorphic sections associated to a sequence of singular Hermitian holomorphic line bundles on a compact Kähler complex space XX. Namely, under mild moment assumptions, we show that the asymptotic distribution of zeros of r…

2017-09-27abs ↗pdf ↗

The study analyzes prediction errors in systems with memory kernels, providing bounds and stability results.

problem Prediction errors in stochastic dynamical systems with memory kernels.
method Analysis of generalized Langevin equations (GLEs) with Volterra equations, integrating synchronized noise coupling and weighted norms.
result Prediction discrepancies decay at a rate determined by the memory kernel's decay, quantitatively bounded by kernel estimation errors.

We propose a stochastic process driven by memory effect with novel distributions including both exponential and leptokurtic heavy-tailed distributions. A class of distribution is analytically derived from the continuum limit of the discrete binary process with the renormalized auto-correlation and the closed form momen…

2012-01-27abs ↗pdf ↗