In several recently proposed stochastic optimization methods (e.g. RMSProp, Adam, Adadelta), parameter updates are scaled by the inverse square roots of exponential moving averages of squared past gradients. Maintaining these per-parameter second-moment estimators requires memory equal to the number of parameters. For …
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
AdamNX improves Adam's stability by adjusting its learning rate.
Adam optimization algorithm can have non-zero average regret under certain conditions.
Weibull weight-scale parameter evolves during AdamW training, with alignment, injection, and decay forces driving its growth and relaxation.
Study examines robust regression in high dimensions with heavy-tailed data.
Suppose centers are fit to points by heuristically minimizing the -means cost; what is the corresponding fit over the source distribution? This question is resolved here for distributions with bounded moments; in particular, the difference between the sample cost and distribution cost decays with $…
We propose a simple stochastic volatility model which is analytically tractable, very easy to simulate and which captures some relevant stylized facts of financial assets, including scaling properties. In particular, the model displays a crossover in the log-return distribution from power-law tails (small time) to a Ga…
Study on U-statistics with heavy-tailed samples, providing tail bounds and LDP.
Study on massless Vlasov equation on Reissner-Nordström spacetimes, showing decay rates and non-decay phenomena.
Regularization in the optimization of deep neural networks is often critical to avoid undesirable over-fitting leading to better generalization of model. One of the most popular regularization algorithms is to impose L-2 penalty on the model parameters resulting in the decay of parameters, called weight-decay, and the …
Cautious Weight Decay modifies weight decay for better optimization.
Novel Adam-family method with decoupled weight decay for training neural networks.
Nonlinear SGD achieves high-probability rates in non-convex optimization with heavy-tailed noise.
New framework for Adam-type algorithms with constant β1, improving regret analysis.
We propose a method of moments (MoM) algorithm for training large-scale implicit generative models. Moment estimation in this setting encounters two problems: it is often difficult to define the millions of moments needed to learn the model parameters, and it is hard to determine which properties are useful when specif…
Classifies scalar-flat toric Kähler instantons in 4D.
New bounds on generalization error using information density moments.
Weight decay stabilizes training dynamics by slowing progressive sharpening.
New method tightens sub-Gaussian concentration inequalities.
Using a relationship between the moments of the probability distribution of times between the two consecutive trades (intertrade time distribution) and the moments of the distribution of a daily number of trades we show, that the underlying point process is essentially non-markovian. A detailed analysis of all trades i…
This paper investigates the effectiveness of decoupled weight decay at the start of training.
A new method for estimating causal parameters from observables reduces the need for finite moment conditions.
Developed moment estimators for affine stochastic volatility models.
This article investigates parameter estimation of affine term structure models by means of the generalized method of moments. Exact moments of the affine latent process as well as of the yields are obtained by using results derived for p-polynomial processes. Then the generalized method of moments, combined with Quasi-…
The non-gaussianity of processes observed in financial markets and relatively good performance of gaussian models can be reconciled by replacing the Brownian motion with Levy processes whose Levy densities decay as exp(-lambda|x|) or faster, where lambda>0 is large. This leads to asymptotic pricing models. The leading …
Study on martingale property and moment explosions in signature volatility models.
Study on eigenvalue distribution of correlated time series, showing deformation of Marchenko-Pastur distribution.
We introduce a new weight-decay scaling rule to maintain sublayer gains across different widths in modern scale-invariant architectures.
We propose a simulation method for multidimensional Hawkes processes based on superposition theory of point processes. This formulation allows us to design efficient simulations for Hawkes processes with differing exponentially decaying intensities. We demonstrate that inter-arrival times can be decomposed into simpler…
A new method to improve deep neural networks using weight rescaling.
We propose NovoGrad, an adaptive stochastic gradient descent method with layer-wise gradient normalization and decoupled weight decay. In our experiments on neural networks for image classification, speech recognition, machine translation, and language modeling, it performs on par or better than well tuned SGD with mom…
This paper studies the Fisher-Rao geometry on the parameter space of beta distributions. We derive the geodesic equations and the sectional curvature, and prove that it is negative. This leads to uniqueness for the Riemannian centroid in that space. We use this Riemannian structure to study canonical moments, an intrin…
This work develops efficient methods for computing moments of Gaussian mixtures.
Bayesian framework uses AI-generated data to improve parameter estimation.
New KCM tests improve specification testing via RKHS.
The paper analyzes Teukolsky equations on Kerr backgrounds, proving boundedness and decay of solutions.
Stochastic Kronecker graphs supply a parsimonious model for large sparse real world graphs. They can specify the distribution of a large random graph using only three or four parameters. Those parameters have however proved difficult to choose in specific applications. This article looks at method of moments estimators…
Gradient descent outperforms ridge regression under certain covariance matrix decay conditions.
This paper presents sampling-based speech parameter generation using moment-matching networks for Deep Neural Network (DNN)-based speech synthesis. Although people never produce exactly the same speech even if we try to express the same linguistic and para-linguistic information, typical statistical speech synthesis pr…
Python package ajdmom simplifies moment formula derivation for jump diffusions.
Mixture modeling is a general technique for making any simple model more expressive through weighted combination. This generality and simplicity in part explains the success of the Expectation Maximization (EM) algorithm, in which updates are easy to derive for a wide class of mixture models. However, the likelihood of…
We present a detailed analysis of \emph{observable} moments based parameter estimators for the Heston SDEs jointly driving the rate of returns and the squared volatilities . Since volatilities are not directly observable, our parameter estimators are constructed from empirical moments of realized volatilitie…
The paper analyzes multivariate Hawkes processes and their induced population processes.
Empower efficient representation of distributions through moment-preserving methods.
We study the problem of subsampling in differential privacy (DP), a question that is the centerpiece behind many successful differentially private machine learning algorithms. Specifically, we provide a tight upper bound on the Rényi Differential Privacy (RDP) (Mironov, 2017) parameters for algorithms that: (1) subsamp…
In this work we prove an universality result regarding the equidistribution of zeros of random holomorphic sections associated to a sequence of singular Hermitian holomorphic line bundles on a compact Kähler complex space . Namely, under mild moment assumptions, we show that the asymptotic distribution of zeros of r…
The study analyzes prediction errors in systems with memory kernels, providing bounds and stability results.
We propose a stochastic process driven by memory effect with novel distributions including both exponential and leptokurtic heavy-tailed distributions. A class of distribution is analytically derived from the continuum limit of the discrete binary process with the renormalized auto-correlation and the closed form momen…