Extends LIBOR market model to reduce exploding scenarios.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper extends de Rham theory of smooth manifolds to exploded manifolds. Included are versions of Stokes' theorem, De Rham cohomology, Poincare duality, and integration along the fiber. The resulting cohomology theory is used to define Gromov Witten invariants of exploded manifolds in a separate paper.
Notes for a short lecture series, covering exploded manifolds, the moduli stack of curves in exploded manifolds, and a tropical gluing formula for Gromov-Witten invariants: a gluing formula providing a degeneration formula for Gromov-Witten invariants in normal-crossing degenerations. I gave the original lecture series…
New neural network approach mitigates vanishing/exploding gradients.
We present a gluing formula for Gromov-Witten invariants in the case of a triple product. This gluing formula is a simple case of a much more general gluing formula proved and stated using exploded manifolds. We present this simple case because it is relatively easy to explain without any knowledge of exploded manifold…
This work tackles exploding inverses in INNs, revealing and mitigating their numerical non-invertibility.
The exploding and vanishing gradient problem has been the major conceptual principle behind most architecture and training improvements in recurrent neural networks (RNNs) during the last decade. In this paper, we argue that this principle, while powerful, might need some refinement to explain recent developments. We r…
Derives a new formula for optimal stopping problems with exploding derivatives.
This paper considers multi-dimensional affine processes with continuous sample paths. By analyzing the Riccati system, which is associated with affine processes via the transform formula, we fully characterize the regions of exponents in which exponential moments of a given process do not explode at any time or explode…
We define Gromov--Witten invariants of exploded manifolds. The technical heart of this paper is a construction of a virtual fundamental class of any Kuranishi category (which is a simplified, more general version of an embedded Kuranishi structure.) We also show how to integrate differential…
Blind Descent avoids gradient issues, using a different learning approach.
We propose a novel approach to addressing the vanishing (or exploding) gradient problem in deep neural networks. We construct a new architecture for deep neural networks where all layers (except the output layer) of the network are a combination of rotation, permutation, diagonal, and activation sublayers which are all…
Maxout networks study gradients and propose initialization strategies.
A new RNN model tackles long-time dependencies with fast, invertible, and memory-efficient hidden states.
ADMMiRNN solves RNN training issues with stable convergence.
In \cite{FOinteger}, Fukaya and Ono outlined a way of counting pseudo-holomorphic curves in a general compact symplectic manifold to obtain integer valued invariants. This paper contains the details of Fukaya and Ono's suggested construction for any compact symplectic manifold and a large class of exploded manifolds.
A simple gating mechanism improves deep learning convergence.
It is well known that the problem of vanishing/exploding gradients is a challenge when training deep networks. In this paper, we describe another phenomenon, called vanishing nodes, that also increases the difficulty of training deep neural networks. As the depth of a neural network increases, the network's hidden node…
The gradient-based optimization method for deep machine learning models suffers from gradient vanishing and exploding problems, particularly when the computational graph becomes deep. In this work, we propose the tangent-space gradient optimization (TSGO) for the probabilistic models to keep the gradients from vanishin…
Recurrent neural networks (RNNs) have gained a great deal of attention in solving sequential learning problems. The learning of long-term dependencies, however, remains challenging due to the problem of a vanishing or exploding hidden states gradient. By exploring further the recently established connections between RN…
Recurrent neural networks (RNNs) are particularly well-suited for modeling long-term dependencies in sequential data, but are notoriously hard to train because the error backpropagated in time either vanishes or explodes at an exponential rate. While a number of works attempt to mitigate this effect through gated recur…
iEFM trains CNF models from unnormalized densities efficiently.
We prove two tropical gluing formulae for Gromov-Witten invariants of exploded manifolds, useful for calculating Gromov-Witten invariants of a symplectic manifold using a normal-crossing degeneration. The first formula generalizes the symplectic-sum formula for Gromov-Witten invariants. The second formula is stronger, …
Improves diffusion models by controlling total variance and signal-to-noise-ratio.
Monte Carlo method trains deep neural networks without gradients.
New RNN model handles long-term dependencies in irregularly-sampled time series.
A well-conditioned Jacobian spectrum has a vital role in preventing exploding or vanishing gradients and speeding up learning of deep neural networks. Free probability theory helps us to understand and handle the Jacobian spectrum. We rigorously show almost sure asymptotic freeness of layer-wise Jacobians of deep neura…
Recurrent Neural Networks (RNN), Long Short-Term Memory Networks (LSTM), and Memory Networks which contain memory are popularly used to learn patterns in sequential data. Sequential data has long sequences that hold relationships. RNN can handle long sequences but suffers from the vanishing and exploding gradient probl…
New method improves DAG learning by using large coefficients for higher-order terms.
Conformally invariant functionals on the space of knots are introduced via extrinsic conformal geometry of the knot and integral geometry on the space of spheres. Our functionals are expressed in terms of a complex-valued 2-form which can be considered as the cross-ratio of a pair of infinitesimal segments of the knot.…
Residual Network (ResNet) is the state-of-the-art architecture that realizes successful training of really deep neural network. It is also known that good weight initialization of neural network avoids problem of vanishing/exploding gradients. In this paper, simplified models of ResNets are analyzed. We argue that good…
Binary Neural Networks (BNNs) are difficult to train, and suffer from drop of accuracy. It appears in practice that BNNs fail to train in the absence of Batch Normalization (BatchNorm) layer. We find the main role of BatchNorm is to avoid exploding gradients in the case of BNNs. This finding suggests that the common in…
One of the difficulties of training deep neural networks is caused by improper scaling between layers. Scaling issues introduce exploding / gradient problems, and have typically been addressed by careful scale-preserving initialization. We investigate the value of preserving scale, or isometry, beyond the initial weigh…
Scaling ResNets requires careful consideration of the layer depth and output scaling factors.
Novel BSG method for efficient stochastic optimization.
AuON is a linear-time optimizer that improves upon Muon's performance without approximate orthogonal matrices.
Vanishing and exploding gradients are two of the main obstacles in training deep neural networks, especially in capturing long range dependencies in recurrent neural networks~(RNNs). In this paper, we present an efficient parametrization of the transition matrix of an RNN that allows us to stabilize the gradients that …
RNNs struggle with chaotic dynamics due to exploding gradients, but we found a way to optimize training.
Training recurrent neural networks (RNNs) is a hard problem due to degeneracies in the optimization landscape, a problem also known as vanishing/exploding gradients. Short of designing new RNN architectures, previous methods for dealing with this problem usually boil down to orthogonalization of the recurrent dynamics,…
Quasi-Gaussian HJM models are a popular approach for modeling the dynamics of the yield curve. This is due to their low dimensional Markovian representation, which greatly simplifies their numerical implementation. We present a qualitative study of the solutions of the quasi-Gaussian log-normal HJM model. Using a small…
Batch normalization (batch norm) is often used in an attempt to stabilize and accelerate training in deep neural networks. In many cases it indeed decreases the number of parameter updates required to achieve low training error. However, it also reduces robustness to small adversarial input perturbations and noise by d…
Gradient flossing stabilizes RNN training by controlling Lyapunov exponents.
Stable ResNet stabilizes gradients in deep networks.
This paper analyzes challenges and solutions in deep learning optimization.
The so called Jenkins-Serrin problem is a kind of Dirichlet problem for graphs with prescribed mean curvature that combines, at the same time, continuous boundary data with regions of the boundary where the boundary values explodes either to or to We give a survey on the development of Jenkins-Serr…
A new method reduces variance in training early-stage rankers for large-scale search systems.
We study the problem of non-explosion of diffusion processes on a manifold with time-dependent Riemannian metric. In particular we obtain that Brownian motion cannot explode in finite time if the metric evolves under backwards Ricci flow. Our result makes it possible to remove the assumption of non-explosion in the pat…
Classical (Itô diffusions) stochastic volatility models are not able to capture the steepness of small-maturity implied volatility smiles. Jumps, in particular exponential Lévy and affine models, which exhibit small-maturity exploding smiles, have historically been proposed to remedy this (see \cite{Tank} for an overvi…