BatchNorm helps train quantized networks by avoiding gradient explosion.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Noise injection before gradient steps helps in regularization for neural networks.
We conduct mathematical analysis on the effect of batch normalization (BN) on gradient backpropogation in residual network training, which is believed to play a critical role in addressing the gradient vanishing/explosion problem, in this work. By analyzing the mean and variance behavior of the input and the gradient i…
We show that the moment explosion time in the rough Heston model [El Euch, Rosenbaum 2016, arxiv:1609.02108] is finite if and only if it is finite for the classical Heston model. Upper and lower bounds for the explosion time are established, as well as an algorithm to compute the explosion time (under some restrictions…
Lipschitz normalization boosts deep attention models, especially for graph neural networks.
New paradigm for Neural ODEs stabilizes training and improves model performance.
Gravilon improves gradient descent for neural networks.
Study on martingale property and moment explosions in signature volatility models.
Stochastic gradient methods are dominant in nonconvex optimization especially for deep models but have low asymptotical convergence due to the fixed smoothness. To address this problem, we propose a simple yet effective method for improving stochastic gradient methods named predictive local smoothness (PLS). First, we …
When and why can a neural network be successfully trained? This article provides an overview of optimization algorithms and theory for training neural networks. First, we discuss the issue of gradient explosion/vanishing and the more general issue of undesirable spectrum, and then discuss practical solutions including …
Early training phase affects deep neural network optimization and generalization.
We study the problem of non-explosion of diffusion processes on a manifold with time-dependent Riemannian metric. In particular we obtain that Brownian motion cannot explode in finite time if the metric evolves under backwards Ricci flow. Our result makes it possible to remove the assumption of non-explosion in the pat…
Gradient oversmoothing and expansion hinder deep GNN training, solved with normalization.
We study the explosion of the solutions of the SDE in the quasi-Gaussian HJM model with a CEV-type volatility. The quasi-Gaussian HJM models are a popular approach for modeling the dynamics of the yield curve. This is due to their low dimensional Markovian representation which simplifies their numerical implementation …
A coupling by reflection of a time-inhomogeneous diffusion process on a manifold are studied. The condition we assume is a natural time-inhomogeneous extension of lower Ricci curvature bounds. In particular, it includes the case of backward Ricci flow. As in time-homogeneous cases, our coupling provides a gradient esti…
New method reduces communication costs in distributed nonconvex optimization.
Bayesian model improves categorization of explosions from sparse data.
Spectral normalization stabilizes GANs by controlling gradient explosion and vanishing.
We propose a randomised version of the Heston model-a widely used stochastic volatility model in mathematical finance-assuming that the starting point of the variance process is a random variable. In such a system, we study the small-and large-time behaviours of the implied volatility, and show that the proposed random…
In this paper, we establish sample path large and moderate deviation principles for log-price processes in Gaussian stochastic volatility models, and study the asymptotic behavior of exit probabilities, call pricing functions, and the implied volatility. In addition, we prove that if the volatility function in an uncor…
Ripple Walk Training tackles graph neural network training issues for large and deep graphs.
Guarantees for tuning step size using meta-gradient descent.
Recently, there is an explosive growth of activities to understand stringy properties of orbifolds. In this article, we survey some of recent developments.
Quasi-Gaussian HJM models are a popular approach for modeling the dynamics of the yield curve. This is due to their low dimensional Markovian representation, which greatly simplifies their numerical implementation. We present a qualitative study of the solutions of the quasi-Gaussian log-normal HJM model. Using a small…
Proposes a normalization technique for manifold valued data.
Study on VIX options pricing in SABR model, showing infinite prices due to volatility explosion.
RestoreAI predicts landmine risk from patterns, improving clearance efficiency.
RNNs struggle with chaotic dynamics due to exploding gradients, but we found a way to optimize training.
xRFM improves tabular data inference with better accuracy and scalability.
We prove that the quotient space of a variationally complete group action is a good Riemannian orbifold. The result is generalized to singular Riemannian foliations without horizontal conjugate points.
Wide-band Electromagnetic Induction Sensors (WEMI) have been used for a number of years in subsurface detection of explosive hazards. While WEMI sensors have proven effective at localizing objects exhibiting large magnetic responses, detecting objects lacking or containing very low amounts of conductive materials can b…
Training a neural network using backpropagation algorithm requires passing error gradients sequentially through the network. The backward locking prevents us from updating network layers in parallel and fully leveraging the computing resources. Recently, there are several works trying to decouple and parallelize the ba…
Several recent trends in machine learning theory and practice, from the design of state-of-the-art Gaussian Process to the convergence analysis of deep neural nets (DNNs) under stochastic gradient descent (SGD), have found it fruitful to study wide random neural networks. Central to these approaches are certain scaling…
In the LIBOR market model, forward interest rates are log-normal under their respective forward measures. This note shows that their distributions under the other forward measures of the tenor structure have approximately log-normal tails.
We present a number of related comparison results, which allow to compare moment explosion times, moment generating functions and critical moments between rough and non-rough Heston models of stochastic volatility. All results are based on a comparison principle for certain non-linear Volterra integral equations. Our u…
Recurrent neural networks (RNNs) are particularly well-suited for modeling long-term dependencies in sequential data, but are notoriously hard to train because the error backpropagated in time either vanishes or explodes at an exponential rate. While a number of works attempt to mitigate this effect through gated recur…
We consider a class of asset pricing models, where the risk-neutral joint process of log-price and its stochastic variance is an affine process in the sense of Duffie, Filipovic and Schachermayer [2003]. First we obtain conditions for the price process to be conservative and a martingale. Then we present some results o…
The search for higher-order feature interactions that are statistically significantly associated with a class variable is of high relevance in fields such as Genetics or Healthcare, but the combinatorial explosion of the candidate space makes this problem extremely challenging in terms of computational efficiency and p…
In recent years, plenty of metrics have been proposed to identify networks that are free of gradient explosion and vanishing. However, due to the diversity of network components and complex serial-parallel hybrid connections in modern DNNs, the evaluation of existing metrics usually requires strong assumptions, complex…
How can local-search methods such as stochastic gradient descent (SGD) avoid bad local minima in training multi-layer neural networks? Why can they fit random labels even given non-convex and non-smooth architectures? Most existing theory only covers networks with one hidden layer, so can we go deeper? In this paper, w…
Scaling ResNets requires careful consideration of the layer depth and output scaling factors.
We exhibit sufficient conditions such that components of a multidimensional SDE giving rise to a local martingale are strict local martingales or martingales. We assume that the equations have diffusion coefficients of the form with being a stochastic volatility term.
Federated machine learning systems have been widely used to facilitate the joint data analytics across the distributed datasets owned by the different parties that do not trust each others. In this paper, we proposed a novel Gradient Boosting Machines (GBM) framework SecureGBM built-up with a multi-party computation mo…
Discrete time random walks on a finite set naturally translate via a one-to-one correspondence to discrete Laplace operators. Typically, Ollivier curvature has been investigated via random walks. We first extend the definition of Ollivier curvature to general weighted graphs and then give a strikingly simple representa…
At present, there is an explosion of practical interest in the pricing of interest rate (IR) derivatives. Textbook pricing methods do not take into account the leptokurticity of the underlying IR process. In this paper, such a leptokurtic behaviour is illustrated using LIBOR data, and a possible martingale pricing sche…
This paper deals with a natural stochastic optimization procedure derived from the so-called Heavy-ball method differential equation, which was introduced by Polyak in the 1960s with his seminal contribution [Pol64]. The Heavy-ball method is a second-order dynamics that was investigated to minimize convex functions f .…
Feedback alignment methods need to be evaluated for accuracy and gradient cosine similarity.
New neural networks model complex phenomena with fewer parameters.