Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

93186278371 · Jun 202019922001200920172026
48 results for gradient explosion

Noise injection before gradient steps helps in regularization for neural networks.

problem Improving generalization in overparametrized neural networks.
method Injecting small noise perturbations before computing gradient steps, especially in layer-wise fashion.
result Small noise perturbations can explicitly regularize neural networks without variance explosion.

We conduct mathematical analysis on the effect of batch normalization (BN) on gradient backpropogation in residual network training, which is believed to play a critical role in addressing the gradient vanishing/explosion problem, in this work. By analyzing the mean and variance behavior of the input and the gradient i…

2018-12-02abs ↗pdf ↗

We show that the moment explosion time in the rough Heston model [El Euch, Rosenbaum 2016, arxiv:1609.02108] is finite if and only if it is finite for the classical Heston model. Upper and lower bounds for the explosion time are established, as well as an algorithm to compute the explosion time (under some restrictions…

2018-01-29abs ↗pdf ↗

Lipschitz normalization boosts deep attention models, especially for graph neural networks.

problem Gradient explosion in deep graph attention networks leads to poor performance.
method Enforcing Lipschitz continuity by normalizing attention scores.
result Deep GAT models with LipschitzNorm achieve state-of-the-art results for tasks with long-range dependencies.

Study on martingale property and moment explosions in signature volatility models.

problem Analyzing the martingale property and moment explosions in signature volatility models.
method Fine analysis of the explosion time of a signature stochastic differential equation.
result The price process is a true martingale if and only if the order of the linear form is odd and a correlation parameter is negative.

Stochastic gradient methods are dominant in nonconvex optimization especially for deep models but have low asymptotical convergence due to the fixed smoothness. To address this problem, we propose a simple yet effective method for improving stochastic gradient methods named predictive local smoothness (PLS). First, we …

2018-05-23abs ↗pdf ↗

When and why can a neural network be successfully trained? This article provides an overview of optimization algorithms and theory for training neural networks. First, we discuss the issue of gradient explosion/vanishing and the more general issue of undesirable spectrum, and then discuss practical solutions including …

2019-12-19abs ↗pdf ↗

Early training phase affects deep neural network optimization and generalization.

problem The choice of learning rate influences generalization in deep learning models.
method Showed that SGD implicitly penalizes the trace of the Fisher Information Matrix (FIM) from the start of training, and explicitly penalizing the trace of FIM improves generalization.
result Catastrophic Fisher explosion (large trace of FIM early in training) is linked to poor generalization.

Gradient oversmoothing and expansion hinder deep GNN training, solved with normalization.

problem Gradient oversmoothing and expansion prevent deep GNN training.
method Proposed normalization method to constrain the Lipschitz bound of each layer.
result Residual GNNs with hundreds of layers can be efficiently trained with the proposed normalization.

We study the explosion of the solutions of the SDE in the quasi-Gaussian HJM model with a CEV-type volatility. The quasi-Gaussian HJM models are a popular approach for modeling the dynamics of the yield curve. This is due to their low dimensional Markovian representation which simplifies their numerical implementation …

2019-08-19abs ↗pdf ↗

New method reduces communication costs in distributed nonconvex optimization.

problem Large communication costs between central server and local workers in distributed learning.
method Communication-compressed AMSGrad for distributed nonconvex optimization.
result Converges to first-order stationary point with same iteration complexity as vanilla AMSGrad.

Bayesian model improves categorization of explosions from sparse data.

problem Challenges in categorizing explosions from limited data.
method Bayesian update to Event Categorization Matrix model with Bayesian Decision Theory.
result Consistent gains in overall accuracy and lower false negative rates.

Spectral normalization stabilizes GANs by controlling gradient explosion and vanishing.

problem Stability and sample quality issues in GAN training.
method Spectral normalization controls gradient explosion and vanishing, improving GAN training stability and sample quality.
result Bidirectional Scaled Spectral Normalization (BSSN) outperforms standard spectral normalization in sample quality and training stability.

We propose a randomised version of the Heston model-a widely used stochastic volatility model in mathematical finance-assuming that the starting point of the variance process is a random variable. In such a system, we study the small-and large-time behaviours of the implied volatility, and show that the proposed random…

2016-08-25abs ↗pdf ↗

Ripple Walk Training tackles graph neural network training issues for large and deep graphs.

problem Neighbors explosion, node dependence, and oversmoothing in large and deep GNNs.
method Subgraph-based training framework with Ripple Walk Sampler for high-quality subgraph sampling.
result RWT improves training efficiency and reduces space complexity for deep and large GNNs.

Recently, there is an explosive growth of activities to understand stringy properties of orbifolds. In this article, we survey some of recent developments.

2002-01-15abs ↗pdf ↗

Quasi-Gaussian HJM models are a popular approach for modeling the dynamics of the yield curve. This is due to their low dimensional Markovian representation, which greatly simplifies their numerical implementation. We present a qualitative study of the solutions of the quasi-Gaussian log-normal HJM model. Using a small…

2019-08-19abs ↗pdf ↗

Study on VIX options pricing in SABR model, showing infinite prices due to volatility explosion.

problem Infinite VIX futures and call prices due to volatility explosion in SABR model.
method Analyzing SABR model, showing vtv_t as unique solution to diffusion process, proving explosion using Feller test, proposing capped volatility process.
result VIX futures and call prices are infinite for any maturity due to volatility explosion, but capped volatility process mitigates this issue.

RestoreAI predicts landmine risk from patterns, improving clearance efficiency.

problem Predicting landmine risk from spatial patterns to enhance clearance efficiency.
method RestoreAI uses landmine patterns for risk prediction, implementing three deminers: linear, curved, and Bayesian.
result RestoreAI significantly boosts clearance efficiency, achieving a 14.37 percentage point increase in cleared landmines per timestep.

RNNs struggle with chaotic dynamics due to exploding gradients, but we found a way to optimize training.

problem Challenging training of RNNs with chaotic dynamics due to exploding gradients.
method Relating loss gradients to Lyapunov spectrum to optimize training on chaotic data.
result RNNs with chaotic dynamics always have diverging gradients, while stable ones have bounded gradients.

xRFM improves tabular data inference with better accuracy and scalability.

problem Inference from tabular data remains challenging and underdeveloped compared to other AI areas.
method Combines feature learning kernel machines with a tree structure.
result xRFM outperforms other methods across 100 regression and 200 classification datasets.

Wide-band Electromagnetic Induction Sensors (WEMI) have been used for a number of years in subsurface detection of explosive hazards. While WEMI sensors have proven effective at localizing objects exhibiting large magnetic responses, detecting objects lacking or containing very low amounts of conductive materials can b…

2019-03-22abs ↗pdf ↗

Training a neural network using backpropagation algorithm requires passing error gradients sequentially through the network. The backward locking prevents us from updating network layers in parallel and fully leveraging the computing resources. Recently, there are several works trying to decouple and parallelize the ba…

2018-07-12abs ↗pdf ↗

In the LIBOR market model, forward interest rates are log-normal under their respective forward measures. This note shows that their distributions under the other forward measures of the tenor structure have approximately log-normal tails.

2010-08-12abs ↗pdf ↗

How can local-search methods such as stochastic gradient descent (SGD) avoid bad local minima in training multi-layer neural networks? Why can they fit random labels even given non-convex and non-smooth architectures? Most existing theory only covers networks with one hidden layer, so can we go deeper? In this paper, w…

2018-10-29abs ↗pdf ↗

Scaling ResNets requires careful consideration of the layer depth and output scaling factors.

problem Avoiding vanishing or exploding gradients in deep ResNets as depth increases.
method Probabilistic analysis and continuous-time limit interpretation of ResNets.
result The optimal scaling factor is αL=1Lα_L = \frac{1}{\sqrt{L}} for standard i.i.d. initializations.

We exhibit sufficient conditions such that components of a multidimensional SDE giving rise to a local martingale MM are strict local martingales or martingales. We assume that the equations have diffusion coefficients of the form σ(Mt,vt),σ(M_t,v_t), with vtv_t being a stochastic volatility term.

2019-03-06abs ↗pdf ↗

Federated machine learning systems have been widely used to facilitate the joint data analytics across the distributed datasets owned by the different parties that do not trust each others. In this paper, we proposed a novel Gradient Boosting Machines (GBM) framework SecureGBM built-up with a multi-party computation mo…

2019-11-27abs ↗pdf ↗

At present, there is an explosion of practical interest in the pricing of interest rate (IR) derivatives. Textbook pricing methods do not take into account the leptokurticity of the underlying IR process. In this paper, such a leptokurtic behaviour is illustrated using LIBOR data, and a possible martingale pricing sche…

2004-01-23abs ↗pdf ↗

This paper deals with a natural stochastic optimization procedure derived from the so-called Heavy-ball method differential equation, which was introduced by Polyak in the 1960s with his seminal contribution [Pol64]. The Heavy-ball method is a second-order dynamics that was investigated to minimize convex functions f .…

2016-09-14abs ↗pdf ↗

New neural networks model complex phenomena with fewer parameters.

problem Challenges in studying higher-order interactions in neural networks.
method Introducing curved neural networks using the maximum entropy principle.
result Curved neural networks accelerate memory retrieval and exhibit explosive phase transitions.