In this paper, we will give a local version of the Hamilton-Ivey type pinching estimate of the gradient shrinking soliton with vanishing Weyl tensor, and then give a complete classification on gradient shrinking solitons with vanishing Weyl tensor.
New classification of gradient steady Ricci solitons with vanishing D-tensor.
problem Classifying gradient steady Ricci solitons with specific properties.
method Extending Cao-Chen's work on Bach-flat gradient Ricci solitons, proving properties for D-flat solitons. result Any n-dimensional complete noncompact gradient steady Ricci soliton with vanishing D-tensor is either Ricci-flat or isometric to the Bryant soliton. Sigmoid-type networks avoid vanishing gradients with regularization and rescaling.
problem Vanishing gradients in sigmoid-type networks.
method Mathematical arguments and two remedies: regularization and rescaling.
result Demonstrates effectiveness of regularization and rescaling in practice.
Vanishing gradients hinder reinforcement finetuning of language models.
problem Vanishing gradients impede the optimization of language models using reinforcement finetuning.
method The study identifies vanishing gradients as a fundamental optimization obstacle in reinforcement finetuning and proposes an initial supervised finetuning phase to mitigate this issue.
result An initial supervised finetuning phase is crucial for successful reinforcement finetuning of language models, as it helps prevent vanishing gradients and maximizes rewards.
In this article, we mathematically study several GAN related topics, including Inception score, label smoothing, gradient vanishing and the -log(D(x)) alternative. --- An advanced version is included in arXiv:1703.02000 "Activation Maximization Generative Adversarial Nets". Please refer Section 6 in 1703.02000 for deta…
It is well known that the problem of vanishing/exploding gradients is a challenge when training deep networks. In this paper, we describe another phenomenon, called vanishing nodes, that also increases the difficulty of training deep neural networks. As the depth of a neural network increases, the network's hidden node…
Normalization layers are widely used in deep neural networks to stabilize training. In this paper, we consider the training of convolutional neural networks with gradient descent on a single training example. This optimization problem arises in recent approaches for solving inverse problems such as the deep image prior…
The natural gradient of ELBO vanishes in unconstrained optimization, simplifying learning.
problem The gap between evidence and ELBO has a vanishing natural gradient.
method Analyzes the Fisher-Rao gradient of ELBO and its implications for learning.
result Maximizing ELBO is equivalent to minimizing KL divergence, simplifying learning.
New neural network approach mitigates vanishing/exploding gradients.
problem Vanishing and exploding gradients in neural networks.
method Gaussian-Poincaré normalized functions and orthogonal weight matrices.
result High-dimensional probability theory shows gradients disappear with high probability in wide neural networks.
We know SGAN may have a risk of gradient vanishing. A significant improvement is WGAN, with the help of 1-Lipschitz constraint on discriminator to prevent from gradient vanishing. Is there any GAN having no gradient vanishing and no 1-Lipschitz constraint on discriminator? We do find one, called GAN-QP. To construct a …
Analyzes self-attention in recurrent networks, proving it mitigates vanishing gradients.
problem Vanishing gradients in recurrent networks when capturing long-term dependencies.
method Formal analysis of self-attention's effect on gradient propagation, proposing a relevancy screening mechanism.
result Self-attention mitigates vanishing gradients in recurrent networks, providing guarantees.
Gradient amplification boosts deep learning model performance without increasing training time.
problem Vanishing gradients in deep neural networks.
method Gradient amplification approach to prevent vanishing gradients and training strategy to enable/disable across epochs.
result Improves performance of deep learning models with reduced training time.
We classify complete gradient Ricci solitons satisfying a fourth-order vanishing condition on the Weyl tensor, improving previously known results. More precisely, we show that any n-dimensional (n≥4) gradient shrinking Ricci soliton with fourth order divergence-free Weyl tensor is either Einstein, or a finite q…
Blind Descent avoids gradient issues, using a different learning approach.
problem Gradient issues like exploding and vanishing gradients.
method Does not use gradients to guide learning; instead, it is a more fundamental learning process.
result Gradient descent is a specific case of Blind Descent.
The exploding and vanishing gradient problem has been the major conceptual principle behind most architecture and training improvements in recurrent neural networks (RNNs) during the last decade. In this paper, we argue that this principle, while powerful, might need some refinement to explain recent developments. We r…
We conduct mathematical analysis on the effect of batch normalization (BN) on gradient backpropogation in residual network training, which is believed to play a critical role in addressing the gradient vanishing/explosion problem, in this work. By analyzing the mean and variance behavior of the input and the gradient i…
Paper proves properties of minimal hypersurfaces in specific solitons.
problem Characterizing minimal hypersurfaces in shrinking gradient Ricci solitons.
method Analyzes stable minimal hypersurfaces with specific curvature conditions.
result Minimal hypersurfaces in these solitons have zero second fundamental form and normal Ricci curvature.
This paper classifies solitons under specific tensor conditions.
problem Classifying solitons under vanishing conditions on the Weyl, Cotton, and Cao-Chen tensors.
method Analyzing complete conformal gradient solitons and using tensor conditions.
result Classification of complete nontrivial locally conformally flat conformal gradient solitons.
Noise causes learning plateaus in neural networks.
problem Plateau phenomena in online learning due to vanishing gradients.
method Analysis of stochastic gradient descent in multi-layer perceptrons.
result Noise induces synchronisation leading to strong plateaus.
A new GAN model α-GAN with tunable loss function addresses gradient vanishing and mode collapse issues.
problem Addressing vanishing gradients and mode collapse in GANs.
method Introduced a tunable GAN α-GAN using a supervised α-loss function. result Holistic understanding of α-GAN related to Arimoto divergence and convergence properties. Maxout networks study gradients and propose initialization strategies.
problem Complexity in input-output Jacobian distribution complicates stable parameter initialization.
method Obtained bounds on moments of gradients and formulated initialization strategies.
result Parameter initialization strategies improve training of deep maxout networks.
New findings on shrinking Ricci solitons with vanishing Bach-like tensors.
problem Characterizing gradient shrinking Ricci solitons with vanishing Bach-like tensors.
method Defining and analyzing Bach-like tensors, proving rigidity results, and deriving variational formulas.
result Vanishing Bach-like tensors force solitons to be either Einstein or isometric to the Gaussian soliton.
Infinitesimal gradient boosting is a new algorithm derived from gradient boosting.
problem Improving the efficiency and smoothness of gradient boosting.
method Introduced a new class of randomized regression trees and used a limit process in vanishing-learning-rate asymptotic.
result Convergence of the stochastic algorithm and characterization of the limiting procedure as a unique solution of a nonlinear ODE.
Spectral normalization stabilizes GANs by controlling gradient explosion and vanishing.
problem Stability and sample quality issues in GAN training.
method Spectral normalization controls gradient explosion and vanishing, improving GAN training stability and sample quality.
result Bidirectional Scaled Spectral Normalization (BSSN) outperforms standard spectral normalization in sample quality and training stability.
Study the properties of SGD in non-vanishing learning rate regime.
problem Understanding the noise and fluctuation in SGD with finite learning rates.
method Derive exact solvable results for discrete-time SGD in quadratic loss functions.
result Fluctuation caused by discrete-time dynamics is larger than continuous-time theory predicts.
We obtain expressions for the shear and the vorticity tensors of perfect-fluid spacetimes, in terms of the divergence of the Weyl tensor. For such spacetimes, we prove that if the gradient of the energy density is parallel to the velocity, then either the expansion rate is zero, or the vorticity vanishes. This statemen…
We show that the only complete shrinking gradient Ricci solitons with vanishing Weyl tensor are quotients of the standard ones. This gives a new proof of the Hamilton-Ivey-Perel'man classification of 3-dimensional shrinking gradient solitons. We also prove a classification for expanding gradient Ricci solitons with con…
The gradient-based optimization method for deep machine learning models suffers from gradient vanishing and exploding problems, particularly when the computational graph becomes deep. In this work, we propose the tangent-space gradient optimization (TSGO) for the probabilistic models to keep the gradients from vanishin…
Improved GNN handles long-range dependencies in multi-relational graphs.
problem Vanishing gradients in GNNs for multi-relational graphs.
method Proposes a Gated Graph Neural Network with improved long-range dependency handling.
result Outperforms popular GNN models in synthetic tasks.
The paper studies topological properties of Ricci shrinkers using weighted L2 cohomology.
problem Proving topological results for smooth gradient Ricci shrinkers.
method Weighted L2 cohomology and extensions to mean curvature flow self-shrinkers. result Establishes upper bounds for Betti numbers, vanishing theorem for cohomology, and dichotomy for ends.
We propose a novel Shapley value approach to help address neural networks' interpretability and "vanishing gradient" problems. Our method is based on an accurate analytical approximation to the Shapley value of a neuron with ReLU activation. This analytical approximation admits a linear propagation of relevance across …
In this paper, we introduce the concept of quasi Yamabe gradient solitons, which generalizes the concept of Yamabe gradient solitons. By using some ideas in [7,8], we prove that n-dimensional (n≥3) complete quasi Yamabe gradient solitons with vanishing Weyl curvature tensor and positive sectional curvature must …
Stagewise boosting improves gradient boosting for distributional regression.
problem Vanishing gradient in gradient boosting for distributional regression leads to suboptimal models.
method Proposes a stagewise boosting-type algorithm for distributional regression, combining stagewise regression ideas with gradient boosting and incorporating a novel regularization method, correlation filtering.
result The proposed algorithm provides better results, especially for complex distributions, by reducing the risk of being trapped in a local optimum.
ADMMiRNN solves RNN training issues with stable convergence.
problem Training RNN with stable convergence and avoiding gradient issues.
method Built ADMMiRNN framework on unfolded RNN, providing novel update rules and theoretical analysis.
result ADMMiRNN achieves convergent results and outperforms baselines.
Generative Adversarial Networks (GAN) have shown promising results on a wide variety of complex tasks. Recent experiments show adversarial training provides useful gradients to the generator that helps attain better performance. In this paper, we intend to theoretically analyze whether supervised learning with adversar…
LCW reduces activation shift in neural networks, improving training efficiency and generalization.
problem Activation shift in neural networks leading to non-zero mean preactivation values.
method Linearly constrained weights (LCW) to reduce activation shift in fully connected and convolutional layers.
result LCW resolves the vanishing gradient problem and improves generalization of neural networks.
In the last decade, the approximate vanishing ideal and its basis construction algorithms have been extensively studied in computer algebra and machine learning as a general model to reconstruct the algebraic variety on which noisy data approximately lie. In particular, the basis construction algorithms developed in ma…
A new RNN model tackles long-time dependencies with fast, invertible, and memory-efficient hidden states.
problem Challenges in processing sequential inputs with long-time dependencies in RNNs.
method A novel RNN architecture based on a Hamiltonian system of oscillators.
result The proposed RNN mitigates exploding and vanishing gradient problems, providing state-of-the-art performance.
New method approximates sampling from smooth potential distributions using a vanishing penalty.
problem Sampling from smooth potential distributions on high-dimensional spaces.
method Penalized Langevin dynamics (PLD) with vanishing penalty.
result Established upper bound on Wasserstein-2 distance for PLD approximation.
Modelling long-term dependencies is a challenge for recurrent neural networks. This is primarily due to the fact that gradients vanish during training, as the sequence length increases. Gradients can be attenuated by transition operators and are attenuated or dropped by activation functions. Canonical architectures lik…
Gradient flossing stabilizes RNN training by controlling Lyapunov exponents.
problem Gradient instability in RNNs leading to exploding and vanishing gradients.
method Regularizing Lyapunov exponents through backpropagation using differentiable linear algebra.
result Gradient flossing improves RNN training success rate and convergence speed.
Compact gradient ρ-Einstein solitons are isometric to Euclidean spheres.
problem Characterizing gradient ρ-Einstein solitons in Riemannian manifolds.
method Proved isometry by showing constant scalar curvature for compact cases and vanishing scalar curvature for non-compact cases with integral conditions.
result Compact gradient ρ-Einstein solitons are isometric to Euclidean spheres.
The main purpose of this article is to provide an alternate proof to a result of Perelman on gradient shrinking solitons. In dimension three we also generalize the result by removing the κ-non-collapsing assumption. In high dimension this new method allows us to prove a classification result on gradient shrinking sol…
Paper introduces a new G⋆ regret measure for online convex optimization with smooth losses.
problem Online convex optimization with smooth losses.
method Introduces a new G⋆ regret measure that depends on the cumulative squared gradient norm. result The G⋆ regret can be arbitrarily sharper than existing measures when losses have vanishing curvature. New framework analyzes effectiveness of neural network-based combinatorial problem solvers.
problem Analyzing neural network-based methods for combinatorial optimization problems.
method Introducing a theoretical framework to assess the effectiveness of solution-samplers using policy-gradient methods.
result Positive theoretical answer to the existence of expressive, tractable, and benign optimization landscapes for combinatorial problems.
Graph manifolds are manifolds that decompose along tori into pieces with a tame S1-structure. In this paper, we prove that the simplicial volume of graph manifolds (which is known to be zero) can be approximated by integral simplicial volumes of their finite coverings. This gives a uniform proof of the vanishing of …
New methods optimize training VQAs without barren plateaus, improving efficiency and applicability.
problem Barren plateaus in training variational quantum algorithms.
method Derive adaptive learning rates and use Gaussian kernels to optimize movement in parameter space.
result Optimized training methods outperform other routines and can train VQAs free of barren plateaus.
Study compares adversarial regularization to sole supervision in machine learning.
problem Understanding when adversarial regularization outperforms sole supervision.
method Examines vanishing gradient, iteration complexity, gradient flow, and convergence in both paradigms.
result Adversarial regularization accelerates gradient descent and improves generalization.