Proposes a continuous flow model to understand and control instability in gradient descent for deep learning.
problem Understanding and controlling the instability of gradient descent in deep learning.
method Introduces the Principal Flow (PF), a continuous time flow that approximates gradient descent dynamics.
result The PF captures divergent and oscillatory behaviors of gradient descent, including escaping local minima and saddle points.
Large learning rates cause parameter instability, leading to better generalization.
problem Understanding why deep neural networks perform well despite operating outside the traditional stability regime.
method Analyzing the effect of large learning rates on the orientation of Hessian eigenvectors and parameter exploration.
result Large learning rates induce parameter instability, leading to better generalization through exploration of flatter regions of the loss landscape.
Deep neural networks struggle with numerical instability during training.
problem Numerical instability in gradient descent training of deep neural networks.
method Analysis of floating-point arithmetic and gradient descent in ReLU neural networks.
result It is highly unlikely for ReLU networks to maintain a superlinear number of affine pieces during training.
We develop a general framework for the description of instabilities on soap films using the Björling representation of minimal surfaces. The construction is naturally geometric and the instability has the interpretation as being specified by its amplitude and transverse gradient along any curve lying in the minimal sur…
Gradient flossing stabilizes RNN training by controlling Lyapunov exponents.
problem Gradient instability in RNNs leading to exploding and vanishing gradients.
method Regularizing Lyapunov exponents through backpropagation using differentiable linear algebra.
result Gradient flossing improves RNN training success rate and convergence speed.
Gradient-based meta-RL fails with incorrect task distributions, leading to instability and poor performance.
problem Gradient-based meta-RL's sensitivity to task distributions causes instability and poor performance.
method Proposes meta Active Domain Randomization (meta-ADR) to learn task distributions for gradient-based meta-RL.
result Meta-ADR improves stability and generalization of MAML on simulated locomotion and navigation tasks.
New findings on optimal transport gradient for generative models, addressing numerical instabilities.
problem Numerical instabilities in training Wasserstein Generative Adversarial Networks (WGAN).
method Valid differentiation theorem for entropic regularized transport, semi-discrete gradient formulation, and optimization algorithm.
result Existence of optimal transport gradient for generative models under specified conditions.
Fractal learning rate schedules accelerate vanilla gradient descent.
problem Difficulty in tuning learning rates in iterative optimization.
method Introduce Chebyshev learning rate schedule for gradient descent.
result Locally unstable updates can lead to convergence in deep learning.
Momentum affects optimization differently at small vs large batch sizes near instability.
problem Understanding how momentum impacts optimization near the edge of stability.
method Demonstrated through batch-size dependent behavior of SGD with momentum.
result Momentum operates in two distinct regimes: amplifying stochastic fluctuations at small batch sizes and stabilizing at large batch sizes.
This work tackles GAN training instability through parallel tempering.
problem Training instability and mode collapse in GANs.
method Introduces a parallel tempering framework to stabilize GAN training.
result Significantly reduces gradient variance and improves training efficiency.
Simplifies RL training with fewer techniques, reducing bias and instability.
problem Training instabilities and high sample complexity in RL.
method Introduced a simple deterministic policy gradient, used propensity estimation, and delayed policy updates.
result Improved performance and reduced sample complexity through these techniques.
Study shows instability of Kähler Ricci solitons and stability of orbifold singularities.
problem Linear stability and instability of Kähler Ricci solitons.
method Extending the approach of \cite{chi04} and \cite{hm11}, via recent work \cite{cm21} on gradient shrinking Ricci solitons.
result Linear instability of the BCCD shrinking soliton and stability of orbifold singularities of Kähler solitons.
Study stabilizes adversarial training in neural networks over infinite-dimensional spaces.
problem Stability issues in adversarial training of neural networks.
method Functional analysis of minimax optimization over infinite-dimensional spaces of continuous functions and probability measures.
result Convergence property of minimax problems under certain conditions, interpreted as stabilization techniques.
While Generative Adversarial Networks (GANs) have seen huge successes in image synthesis tasks, they are notoriously difficult to adapt to different datasets, in part due to instability during training and sensitivity to hyperparameters. One commonly accepted reason for this instability is that gradients passing from t…
Introduces TDRC to balance TD's ease and soundness.
problem TD learning's instability and divergence issues.
method Gradient Temporal-Difference Learning with Regularized Corrections (TDRC).
result TDRC performs as well as TD when TD works, but is sound in divergent cases.
Proposes a new adaptive gradient method based on gradient differences.
problem Manual tuning of stepsize in vanilla gradient methods.
method Adaptation driven by cumulative squared norms of gradient differences.
result More robust than AdaGrad in various settings.
BERT fine-tuning is unstable due to optimization issues, not forgetting or dataset size.
problem Stability of fine-tuning BERT-based models across different random seeds.
method Analysis of BERT, RoBERTa, and ALBERT fine-tuned on GLUE datasets, identifying optimization difficulties as the cause of instability.
result Fine-tuning instability is due to optimization difficulties leading to vanishing gradients, not forgetting or dataset size.
Gradient descent at edge of stability stabilizes implicitly, following projected gradient descent.
problem Gradient descent's stability and sharpness behavior at the edge of instability.
method Cubic Taylor expansion analysis of gradient descent dynamics.
result Gradient descent at edge of stability implicitly follows projected gradient descent.
We present an approach for efficiently training Gaussian Mixture Model (GMM) by Stochastic Gradient Descent (SGD) with non-stationary, high-dimensional streaming data. Our training scheme does not require data-driven parameter initialization (e.g., k-means) and can thus be trained based on a random initialization. Furt…
Accelerated gradient method's stability deteriorates exponentially with steps.
problem Algorithmic stability of Nesterov's accelerated gradient method.
method Analysis of two notions of algorithmic stability for Nesterov's accelerated gradient method.
result Stability of Nesterov's accelerated method deteriorates exponentially with the number of gradient steps.
This study explains how adversarial interaction creates non-homogeneous patterns using a pseudo-Reaction-Diffusion model.
problem Understanding how adversarial interaction leads to non-homogeneous patterns in systems.
method Developed a pseudo-Reaction-Diffusion model to explain the mechanism.
result Turing instability is involved in creating non-homogeneous patterns.
We improve current instability-based methods for the selection of the number of clusters k in cluster analysis by developing a normalized cluster instability measure that corrects for the distribution of cluster sizes, a previously unaccounted driver of cluster instability. We show that our normalized instability mea…
Modeling HFT interactions reveals market instability.
problem Market instability caused by HFT dynamic coupling.
method Developed a recurrence relations framework to model HFT interactions.
result Unexpected latency and feedback can trigger market instability.
New proof of instability for certain Einstein metrics.
problem Einstein metrics on specific 4-manifolds.
method Proving instability of conformally Kähler, Einstein metrics.
result Proven instability of certain Einstein metrics.
New method stabilizes saddle-point optimization with unbounded gradients.
problem Stochastic saddle-point optimization faces instability due to large gradients.
method Proposes a regularization technique to stabilize iterates.
result Yields meaningful performance guarantees even with unbounded gradients.
Proposes a differentiable LSE-ICNN for modeling multi-well potentials.
problem Modeling multi-well potentials in various scientific domains.
method Log-sum-exponential (LSE) mixture of input convex neural network (ICNN) modes.
result Smooth surrogate that retains convexity within basins and allows gradient-based learning.
Characterizes corridors in loss surfaces for gradient-based optimization.
problem Understanding and mitigating training instabilities in gradient-based optimization.
method Characterizes corridors as regions where gradient descent and gradient flow trajectories are linearly related.
result Corridors indicate regions without implicit regularization effects, leading to better learning rate adaptation schemes.
Study stability and instability of Ricci-flat metrics under generalized Ricci flow.
problem Stability and instability of Ricci-flat metrics under Ricci flow.
method Analysis of generalized Ricci flow for Ricci-flat metrics and vanishing 3-forms.
result Dynamical stability and instability results for Ricci-flat metrics and vanishing 3-forms.
Some exotic compact objects possess evanescent ergosurfaces: timelike submanifolds on which a Killing vector field, which is timelike everywhere else, becomes null. We show that any manifold possessing an evanescent ergosurface but no event horizon exhibits a linear instability of a peculiar kind: either there are solu…
There is widespread sentiment that it is not possible to effectively utilize fast gradient methods (e.g. Nesterov's acceleration, conjugate gradient, heavy ball) for the purposes of stochastic optimization due to their instability and error accumulation, a notion made precise in d'Aspremont 2008 and Devolder, Glineur, …
New recommendations improve Gaussian process accuracy and stability.
problem Numerical instabilities and poor test likelihoods in iterative Gaussian process learning.
method Investigated CG tolerance, preconditioner rank, and Lanczos decomposition rank. Recommended small CG tolerance and large root decomposition size.
result L-BFGS-B optimizer achieves convergence with fewer gradient updates, improving Gaussian process accuracy.
This work addresses the instability in asynchronous data parallel optimization. It does so by introducing a novel distributed optimizer which is able to efficiently optimize a centralized model under communication constraints. The optimizer achieves this by pushing a normalized sequence of first-order gradients to a pa…
Interval Neural Networks detect instabilities in image reconstructions.
problem Detecting instabilities in deep learning image reconstructions.
method Employed uncertainty quantification methods with Interval Neural Networks.
result Interval Neural Networks effectively reveal image reconstruction instabilities.
Note on instabilities in super-time-stepping methods for Heston model.
problem Instabilities in super-time-stepping methods applied to Heston model.
method Exploration of explicit super-time-stepping schemes (RK-Chebyshev, RK-Legendre) for Heston model.
result Relevance of stability remarks beyond super-time-stepping schemes.
The paper explores how word embeddings affect the stability of downstream NLP models.
problem Small changes in training data can cause significant changes in model predictions.
method Empirical and theoretical analysis of embedding instability, including the introduction of eigenspace instability measure.
result Increasing embedding memory can reduce the disagreement in predictions by 5% to 37%.
Proposes a normalization technique for manifold valued data.
problem Instability in optimization for manifold valued data.
method Develops a general normalization technique for manifold valued data.
result Demonstrates performance gain in synthetic and real datasets.
Binary perceptron's instability linked to replica symmetry breaking.
problem Understanding the relationship between algorithmic instability and replica symmetry breaking in binary perceptron learning.
method Established the connection between algorithmic instability and replica symmetry breaking by comparing the instability condition around the fixed point to the instability for breaking the replica symmetric solution of the free energy function.
result The instability condition around the algorithmic fixed point is identical to the instability for breaking the replica symmetric saddle point solution of the free energy function.
A central area of research in nonlinear science is the study of instabilities that drive the emergence of extreme events. Unfortunately, experimental techniques for measuring such phenomena often provide only partial characterization. For example, real-time studies of instabilities in nonlinear fibre optics frequently …
Behavior cloning training instabilities amplified by SGD noise over long horizons.
problem Training instabilities in behavior cloning with deep neural networks.
method Empirical dissection of minibatch SGD updates and their effects on long-horizon rewards.
result Exponential moving average (EMA) of iterates effectively mitigates gradient variance amplification (GVA).
New research shows hyperbolic embeddings are useful for global consistency tasks in graphs.
problem The usefulness of hyperbolic representations in graph learning tasks.
method Computed hyperbolic embeddings for node classification and link prediction tasks, addressing optimization issues at zero curvature.
result Hyperbolic embeddings are more effective for tasks requiring global consistency, while Euclidean models are superior for other tasks.
Stochastic gradient descent procedures have gained popularity for parameter estimation from large data sets. However, their statistical properties are not well understood, in theory. And in practice, avoiding numerical instability requires careful tuning of key parameters. Here, we introduce implicit stochastic gradien…
The study examines stability and instability of Poincaré-Einstein metrics using Ricci flow.
problem Stability and instability of Poincaré-Einstein metrics.
method Variant of expander entropy for asymptotically hyperbolic manifolds, local positive mass theorem, volume comparison.
result Characterization of stability and instability in terms of local positive mass theorem and volume comparison.
New method stabilizes GAN training by solving ODEs.
problem Stability issues in GAN training.
method Solving ordinary differential equations (ODEs) to stabilize GAN training.
result Well-known ODE solvers can stabilize GAN training.
A method for online tensor dictionary learning is proposed. With the assumption of separable dictionaries, tensor contraction is used to diminish a N-way model of O(LN) into a simple matrix equation of O(NL2) with a real-time capability. To avoid numerical instability d…
Study shows instability of naked singularities in perfect fluid models.
problem Instability of naked singularities in Einstein equations coupled with isothermal perfect fluid.
method Investigated spherically symmetric self-similar naked singularities under C1,α perturbations of an external massless scalar field. result Spherically symmetric self-similar naked singularities are unstable to trapped surface formation.
Study on spectral stability of Riemannian coverings.
problem Stability of eigenvalues in Riemannian coverings.
method Analysis of Laplacian eigenvalues under finite coverings.
result Necessary conditions for spectral stability or instability.
Clinical models can be unstable, leading to unreliable predictions.
problem Stability of clinical prediction models developed using statistical or machine learning methods.
method Simulation and case studies of statistical and machine learning approaches to show instability in model predictions.
result Model instability often leads to miscalibration of predictions in new data.
Previously, the exploding gradient problem has been explained to be central in deep learning and model-based reinforcement learning, because it causes numerical issues and instability in optimization. Our experiments in model-based reinforcement learning imply that the problem is not just a numerical issue, but it may …