New neural network approach mitigates vanishing/exploding gradients.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
TSGO optimizes gradients in tensor networks to avoid vanishing/exploding issues.
The exploding and vanishing gradient problem has been the major conceptual principle behind most architecture and training improvements in recurrent neural networks (RNNs) during the last decade. In this paper, we argue that this principle, while powerful, might need some refinement to explain recent developments. We r…
Blind Descent avoids gradient issues, using a different learning approach.
BatchNorm helps BNNs avoid exploding gradients during training.
This work tackles exploding inverses in INNs, revealing and mitigating their numerical non-invertibility.
Maxout networks study gradients and propose initialization strategies.
Volume-preserving neural networks prevent gradient issues.
ADMMiRNN solves RNN training issues with stable convergence.
A new RNN model tackles long-time dependencies with fast, invertible, and memory-efficient hidden states.
ENRNN uses eigenvalue normalization for short-term memory in RNNs.
A simple gating mechanism improves deep learning convergence.
Gradient flossing stabilizes RNN training by controlling Lyapunov exponents.
Hamiltonian RNN controls hidden states gradient for long-term dependencies.
Novel BSG method for efficient stochastic optimization.
Vanishing nodes cause hidden nodes to behave similarly, complicating deep neural network training.
Stable ResNet stabilizes gradients in deep networks.
This paper extends de Rham theory of smooth manifolds to exploded manifolds. Included are versions of Stokes' theorem, De Rham cohomology, Poincare duality, and integration along the fiber. The resulting cohomology theory is used to define Gromov Witten invariants of exploded manifolds in a separate paper.
RNNs struggle with chaotic dynamics due to exploding gradients, but we found a way to optimize training.
New RNN model handles long-term dependencies in irregularly-sampled time series.
Normalization layers are widely used in deep neural networks to stabilize training. In this paper, we consider the training of convolutional neural networks with gradient descent on a single training example. This optimization problem arises in recent approaches for solving inverse problems such as the deep image prior…
Vanishing and exploding gradients are two of the main obstacles in training deep neural networks, especially in capturing long range dependencies in recurrent neural networks~(RNNs). In this paper, we present an efficient parametrization of the transition matrix of an RNN that allows us to stabilize the gradients that …
This paper analyzes challenges and solutions in deep learning optimization.
We give a rigorous analysis of the statistical behavior of gradients in a randomly initialized fully connected network N with ReLU activations. Our results show that the empirical variance of the squares of the entries in the input-output Jacobian of N is exponential in a simple architecture-dependent constant beta, gi…
Recurrent neural networks (RNNs) are particularly well-suited for modeling long-term dependencies in sequential data, but are notoriously hard to train because the error backpropagated in time either vanishes or explodes at an exponential rate. While a number of works attempt to mitigate this effect through gated recur…
A new method reduces variance in training early-stage rankers for large-scale search systems.
Scaling ResNets requires careful consideration of the layer depth and output scaling factors.
Convolutional neural network is a very important model of deep learning. It can help avoid the exploding/vanishing gradient problem and improve the generalizability of a neural network if the singular values of the Jacobian of a layer are bounded around in the training process. We propose a new penalty function for…
Notes for a short lecture series, covering exploded manifolds, the moduli stack of curves in exploded manifolds, and a tropical gluing formula for Gromov-Witten invariants: a gluing formula providing a degeneration formula for Gromov-Witten invariants in normal-crossing degenerations. I gave the original lecture series…
Recurrent Neural Networks (RNN), Long Short-Term Memory Networks (LSTM), and Memory Networks which contain memory are popularly used to learn patterns in sequential data. Sequential data has long sequences that hold relationships. RNN can handle long sequences but suffers from the vanishing and exploding gradient probl…
A well-conditioned Jacobian spectrum has a vital role in preventing exploding or vanishing gradients and speeding up learning of deep neural networks. Free probability theory helps us to understand and handle the Jacobian spectrum. We rigorously show almost sure asymptotic freeness of layer-wise Jacobians of deep neura…
AuON is a linear-time optimizer that improves upon Muon's performance without approximate orthogonal matrices.
Gradient control plays an important role in feed-forward networks applied to various computer vision tasks. Previous work has shown that Recurrent Highway Networks minimize the problem of vanishing or exploding gradients. They achieve this by setting the eigenvalues of the temporal Jacobian to 1 across the time steps. …
Residual Network (ResNet) is the state-of-the-art architecture that realizes successful training of really deep neural network. It is also known that good weight initialization of neural network avoids problem of vanishing/exploding gradients. In this paper, simplified models of ResNets are analyzed. We argue that good…
One of the difficulties of training deep neural networks is caused by improper scaling between layers. Scaling issues introduce exploding / gradient problems, and have typically been addressed by careful scale-preserving initialization. We investigate the value of preserving scale, or isometry, beyond the initial weigh…
A common problem in training neural networks is the vanishing and/or exploding gradient problem which is more prominently seen in training of Recurrent Neural Networks (RNNs). Thus several algorithms have been proposed for training RNNs. This paper proposes a novel adaptive stochastic Nesterov accelerated quasiNewton (…
A long-standing obstacle to progress in deep learning is the problem of vanishing and exploding gradients. Although, the problem has largely been overcome via carefully constructed initializations and batch normalization, architectures incorporating skip-connections such as highway and resnets perform much better than …
Convolutional neural network is an important model in deep learning. To avoid exploding/vanishing gradient problems and to improve the generalizability of a neural network, it is desirable to have a convolution operation that nearly preserves the norm, or to have the singular values of the transformation matrix corresp…
We present a gluing formula for Gromov-Witten invariants in the case of a triple product. This gluing formula is a simple case of a much more general gluing formula proved and stated using exploded manifolds. We present this simple case because it is relatively easy to explain without any knowledge of exploded manifold…
It is a known fact that training recurrent neural networks for tasks that have long term dependencies is challenging. One of the main reasons is the vanishing or exploding gradient problem, which prevents gradient information from propagating to early layers. In this paper we propose a simple recurrent architecture, th…
A novel RNN model with shuffled hidden states.
A new RNN model based on coupled oscillators mitigates gradient issues.
Proposes methods to add constraints to neural networks to improve stability and generalization.
A new gradient estimator reduces variance near boundaries for binary latent variables.
LEM efficiently models long-term sequences with gradients.
Training recurrent neural networks (RNNs) is a hard problem due to degeneracies in the optimization landscape, a problem also known as vanishing/exploding gradients. Short of designing new RNN architectures, previous methods for dealing with this problem usually boil down to orthogonalization of the recurrent dynamics,…
Batch normalization (batch norm) is often used in an attempt to stabilize and accelerate training in deep neural networks. In many cases it indeed decreases the number of parameter updates required to achieve low training error. However, it also reduces robustness to small adversarial input perturbations and noise by d…
Paper proves multiplicative weight updates can train neural networks without learning rate tuning.