Exploded manifolds and tropical gluing formula for Gromov-Witten invariants.
problem Calculating Gromov-Witten invariants in normal-crossing degenerations.
method Exploded manifolds and tropical gluing formula.
result Provides a degeneration formula for Gromov-Witten invariants.
This paper extends de Rham theory of smooth manifolds to exploded manifolds. Included are versions of Stokes' theorem, De Rham cohomology, Poincare duality, and integration along the fiber. The resulting cohomology theory is used to define Gromov Witten invariants of exploded manifolds in a separate paper.
Formula calculates Gromov-Witten invariants for triple products.
problem Calculating Gromov-Witten invariants for complex geometrical structures.
method Developed a gluing formula for Gromov-Witten invariants in a triple product.
result Simplified gluing formula for easy explanation.
This paper defines Gromov-Witten invariants for exploded manifolds.
problem Defining Gromov-Witten invariants for a new class of manifolds.
method Construction of a virtual fundamental class for Kuranishi categories.
result Independence and compatibility of the invariants with various operations.
Two tropical gluing formulas help calculate Gromov-Witten invariants.
problem Calculating Gromov-Witten invariants of symplectic manifolds.
method Tropical geometry applied to exploded manifolds.
result Generalizes existing formulas for Gromov-Witten invariants.
ERNNs evolve hidden states on an ODE's equilibrium manifold to mitigate vanishing and exploding gradients.
problem Vanishing and exploding gradients in RNNs.
method Develop a novel family of RNNs (ERNNs) that evolve hidden states on the equilibrium manifold of an ODE.
result ERNNs achieve state-of-the-art accuracy with 3-10x speedups and 1.5-3x model size reduction.
In \cite{FOinteger}, Fukaya and Ono outlined a way of counting pseudo-holomorphic curves in a general compact symplectic manifold to obtain integer valued invariants. This paper contains the details of Fukaya and Ono's suggested construction for any compact symplectic manifold and a large class of exploded manifolds.
Paper refines RNN training by analyzing smoothness and attractors.
problem Exploding and vanishing gradients in RNNs.
method Refined concept of exploding gradients using cost function smoothness.
result RNNs need to learn attractors to fully use their power.
New neural network approach mitigates vanishing/exploding gradients.
problem Vanishing and exploding gradients in neural networks.
method Gaussian-Poincaré normalized functions and orthogonal weight matrices.
result High-dimensional probability theory shows gradients disappear with high probability in wide neural networks.
This work tackles exploding inverses in INNs, revealing and mitigating their numerical non-invertibility.
problem Exploding inverses in INNs cause numerical non-invertibility, leading to failures in various tasks.
method Derived bi-Lipschitz properties of INN building blocks, proposed regularizers for local invertibility, and stable INN designs for global invertibility.
result Bi-Lipschitz properties and stable INN designs are crucial for addressing numerical non-invertibility.
TSGO optimizes gradients in tensor networks to avoid vanishing/exploding issues.
problem Gradient vanishing and exploding problems in deep learning models.
method TSGO rotates parameters towards gradient direction in tangent space of normalized state.
result TSGO naturally determines learning rate based on angle between parameters and gradient.
Analysis shows exploding and vanishing gradients in neural nets with ReLU activations.
problem Exploding and vanishing gradients in neural nets with ReLU activations.
method Rigorous analysis of gradient behavior in randomly initialized fully connected networks with ReLU activations.
result Empirical variance of gradients is exponential in beta, the sum of reciprocals of hidden layer widths.
Derives a new formula for optimal stopping problems with exploding derivatives.
problem Optimal stopping problems with complex boundary conditions.
method Develops a change of variable formula for functions with exploding derivatives near a surface.
result Derives a formula similar to Itô's but with less restrictive conditions.
Extends LIBOR market model to reduce exploding scenarios.
problem Exploding scenarios in market-consistent guarantees valuation.
method Mean-field extension of the LIBOR market model.
result Existence and uniqueness of MF-LMM proved.
BatchNorm helps BNNs avoid exploding gradients during training.
problem Difficulty in training Binary Neural Networks (BNNs).
method Theoretical study and numerical experiments on the role of BatchNorm in BNNs.
result BatchNorm is crucial for avoiding exploding gradients in BNNs.
New methods maintain scale to speed up deep learning.
problem Improper scaling between layers causes exploding gradients in deep neural networks.
method Two methods of maintaining isometry (exact and stochastic) are proposed.
result Maintaining scale speeds up learning, especially in the early stages.
This paper considers multi-dimensional affine processes with continuous sample paths. By analyzing the Riccati system, which is associated with affine processes via the transform formula, we fully characterize the regions of exponents in which exponential moments of a given process do not explode at any time or explode…
PIPPS solves deep learning's exploding gradient problem by reparameterization gradients.
problem Exploding gradients in deep learning and model-based RL.
method Develops PIPPS framework, a flexible policy search method robust to chaos-like gradients.
result PIPPS improves over reparameterization gradients by up to 10^6 times.
Solves exploding and vanishing gradient problem in LSTMs.
problem Exploding and vanishing gradient problem in LSTM optimization.
method Introduces a simple stochastic algorithm (h-detach) to prevent suppression of gradient components through the cell state path in LSTM.
result Significant improvements in convergence speed, robustness, and generalization over vanilla LSTM training.
We study the problem of non-explosion of diffusion processes on a manifold with time-dependent Riemannian metric. In particular we obtain that Brownian motion cannot explode in finite time if the metric evolves under backwards Ricci flow. Our result makes it possible to remove the assumption of non-explosion in the pat…
Blind Descent avoids gradient issues, using a different learning approach.
problem Gradient issues like exploding and vanishing gradients.
method Does not use gradients to guide learning; instead, it is a more fundamental learning process.
result Gradient descent is a specific case of Blind Descent.
Study shows short rate can explode to infinity in HJM model, impacting Eurodollar futures.
problem Exploding short rate in HJM model affecting Eurodollar futures.
method Small-noise deterministic limit analysis.
result Explicit explosion criteria derived for short rate under mild assumptions.
LTM tackles long sequence language modeling by avoiding vanishing and exploding gradients.
problem Language models struggle with long sequences due to gradient issues.
method Introduces Long Term Memory network (LTM) that scales memory and weights input, avoiding overfitting.
result LTM achieves state-of-the-art perplexity results with fewer cells than previous models.
The abstract discusses graphs with prescribed mean curvature on Riemannian manifolds, including existence and translating graphs.
problem Graphs with prescribed mean curvature on Riemannian manifolds with boundary conditions at infinity.
method Survey and proof of existence theorems for Jenkins-Serrin graphs.
result Existence of translating Jenkins-Serrin graphs.
Maxout networks study gradients and propose initialization strategies.
problem Complexity in input-output Jacobian distribution complicates stable parameter initialization.
method Obtained bounds on moments of gradients and formulated initialization strategies.
result Parameter initialization strategies improve training of deep maxout networks.
Batch-normalized RHN improves gradient control in recurrent networks.
problem Gradient vanishing or exploding in recurrent networks.
method Batch normalization applied at each recurrence loop in RHN.
result Batch-normalized RHN converges faster and performs better.
Volume-preserving neural networks prevent gradient issues.
problem Vanishing and exploding gradients in deep neural networks.
method A new neural network architecture with volume-preserving sublayers.
result Volume-preserving neural networks maintain gradient stability.
A new RNN model tackles long-time dependencies with fast, invertible, and memory-efficient hidden states.
problem Challenges in processing sequential inputs with long-time dependencies in RNNs.
method A novel RNN architecture based on a Hamiltonian system of oscillators.
result The proposed RNN mitigates exploding and vanishing gradient problems, providing state-of-the-art performance.
ADMMiRNN solves RNN training issues with stable convergence.
problem Training RNN with stable convergence and avoiding gradient issues.
method Built ADMMiRNN framework on unfolded RNN, providing novel update rules and theoretical analysis.
result ADMMiRNN achieves convergent results and outperforms baselines.
Fixup replaces normalization in deep networks, achieving similar stability and performance.
problem The effectiveness of normalization layers in deep neural networks.
method Fixed-update initialization (Fixup) to solve exploding and vanishing gradient problems.
result Residual networks trained with Fixup achieve state-of-the-art performance without normalization.
AntisymmetricRNN improves RNN trainability without extra computation.
problem Difficulty in learning long-term dependencies in RNNs.
method Connecting RNNs to ordinary differential equations and proposing AntisymmetricRNN.
result AntisymmetricRNN captures long-term dependencies more predictably and efficiently.
Non-normal RNNs outperform orthogonal ones in sequential tasks.
problem Vanishing/exploding gradients in RNNs training.
method Investigate non-normal RNNs with non-normal recurrent connectivity matrix.
result Non-normal RNNs outperform orthogonal ones in various benchmarks.
ENRNN uses eigenvalue normalization for short-term memory in RNNs.
problem Vanishing/exploding gradient problem and long-term dependency modeling.
method Eigenvalue normalization of recurrent matrix to simulate short-term memory.
result ENRNN outperforms existing RNN variants in experiments.
A simple gating mechanism improves deep learning convergence.
problem Vanishing or exploding gradients in deep networks.
method Introducing a zero-initialized parameter to each residual connection.
result Training deep networks (up to 120 layers) with fast convergence and better performance.
Solves Multivariate-MAB problem with path planning and Thompson sampling.
problem Exponential exploding issue in multivariate Multi-Armed Bandit (Multivariate-MAB) problem.
method Path planning framework using decision graphs and Thompson sampling for heuristic arm selection.
result Achieves faster convergence speed, better efficient arm allocation, and lower cumulative regret.
Unitary RNNs tackle vanishing/exploding gradients in long-term dependencies.
problem Vanishing and exploding gradients in RNNs for long-term dependencies.
method Proposes a unitary weight matrix with eigenvalues of absolute value 1, parametrized by structured matrices.
result Achieves state-of-the-art results in tasks with very long-term dependencies.
New method prevents 'shattered gradients' in deep networks, improving training of very deep models.
problem Vanishing and exploding gradients in deep learning networks.
method Introducing 'looks linear' (LL) initialization to prevent gradient shattering.
result Gradient shattering is prevented, allowing training of very deep networks without skip-connections.
A basic RNN model with bounded solutions and fast dynamics.
problem Time-series regression and gradient vanishing/exploding issues.
method State space viewpoint, CLM, CoV, gradient descent, co-state dynamics.
result Successfully performs regression tracking of time-series with quantified gradient issues.
Vanishing nodes cause hidden nodes to behave similarly, complicating deep neural network training.
problem Complicating deep neural network training due to hidden nodes behaving similarly.
method Introduced vanishing node indicator (VNI) to quantify node redundancy and correlated behavior.
result The effective number of nodes vanishes to one as VNI increases, indicating a challenging training scenario.
Study on preventing early training failure modes in deep ReLU nets.
problem Early training failure modes in deep ReLU nets.
method Proved and avoided two failure modes: exploding/vanishing mean activation length and exponentially large variance of activation length.
result Correct initialization and architecture can prevent early training failure modes in deep ReLU nets.
iEFM trains CNF models from unnormalized densities efficiently.
problem Training generators from energy functions or unnormalized densities.
method Iterated energy-based flow matching (iEFM) with simulation-free objective.
result iEFM outperforms existing methods in probabilistic modeling.
LRA trains deep networks robustly with less sensitivity to initial weights.
problem Training deep networks is challenging due to issues like exploding and vanishing gradients.
method Local Representation Alignment (LRA) is a training procedure less sensitive to initial weights.
result LRA can train networks robustly, even with null initial weights, and outperforms other methods.
Improves diffusion models by controlling total variance and signal-to-noise-ratio.
problem Long sampling time in diffusion models.
method Total-Variance/Signal-to-Noise-Ratio (TV/SNR) disentangled framework.
result Improves generation performance by controlling TV and SNR independently.
New RNN model handles long-term dependencies in irregularly-sampled time series.
problem Handling long-term dependencies in irregularly-sampled time series data.
method Designing ODE-LSTMs that separate memory from continuous-time state.
result ODE-LSTMs outperform other RNN-based models on non-uniformly sampled data with long-term dependencies.
New method improves DAG learning by using large coefficients for higher-order terms.
problem Recovering DAG structures from observational data is challenging due to combinatorial optimization.
method Proposes truncated matrix power iteration to approximate DAG constraints efficiently.
result Empirically outperforms previous methods by a factor of 3 or more in structural Hamming distance.
Hamiltonian RNN controls hidden states gradient for long-term dependencies.
problem Challenges in learning long-term dependencies in RNNs.
method Symplectic discretization of Hamiltonian system to control gradient.
result Hamiltonian RNN outperforms other RNNs without hyperparameter optimization.
Deep CNNs can be trained without special architectures.
problem Training extremely deep CNNs (10,000 layers) is challenging due to vanishing/exploding gradients.
method Developed a mean field theory for signal propagation and conditions for dynamical isometry. Derived an algorithm for generating random initial orthogonal convolution kernels.
result Vanilla CNNs with ten thousand layers can be efficiently trained using appropriate initialization schemes.
Conformally invariant functionals on the space of knots are introduced via extrinsic conformal geometry of the knot and integral geometry on the space of spheres. Our functionals are expressed in terms of a complex-valued 2-form which can be considered as the cross-ratio of a pair of infinitesimal segments of the knot.…