Many industrial machine learning (ML) systems require frequent retraining to keep up-to-date with constantly changing data. This retraining exacerbates a large challenge facing ML systems today: model training is unstable, i.e., small changes in training data can cause significant changes in the model's predictions. In…
Dual-objective GANs reduce training instabilities with tunable α-loss parameters.
problem Training instabilities in Generative Adversarial Networks (GANs).
method Introduce (αD,αG)-GANs with dual objectives modeled using α-loss. result Upper bounds on estimation error show improved performance under certain conditions.
Generative adversarial nets (GANs) are a promising technique for modeling a distribution from samples. It is however well known that GAN training suffers from instability due to the nature of its maximin formulation. In this paper, we explore ways to tackle the instability problem by dualizing the discriminator. We sta…
A central area of research in nonlinear science is the study of instabilities that drive the emergence of extreme events. Unfortunately, experimental techniques for measuring such phenomena often provide only partial characterization. For example, real-time studies of instabilities in nonlinear fibre optics frequently …
Deep neural networks struggle with numerical instability during training.
problem Numerical instability in gradient descent training of deep neural networks.
method Analysis of floating-point arithmetic and gradient descent in ReLU neural networks.
result It is highly unlikely for ReLU networks to maintain a superlinear number of affine pieces during training.
Bayesian neural network predicts planetary instability.
problem Predicting planetary instability in compact systems.
method Novel Bayesian neural network trained on raw orbital elements.
result Model predicts planetary instability times with high accuracy and robust generalization.
Simplifies RL training with fewer techniques, reducing bias and instability.
problem Training instabilities and high sample complexity in RL.
method Introduced a simple deterministic policy gradient, used propensity estimation, and delayed policy updates.
result Improved performance and reduced sample complexity through these techniques.
Computer vision models are unstable due to task symmetries and labelling issues.
problem Instability of computer vision models in classification tasks.
method Analysis of symmetries, categorical nature, and labelling issues.
result Instability is a necessary result of current computer vision formulation.
Improved GAN training stability through tunable classification losses.
problem Training instabilities in GANs.
method Reformulated GAN value function using class probability estimation (CPE) losses, defined (αD,αG)-GANs. result Tuning (αD,αG) can alleviate training instabilities. Large learning rates cause parameter instability, leading to better generalization.
problem Understanding why deep neural networks perform well despite operating outside the traditional stability regime.
method Analyzing the effect of large learning rates on the orientation of Hessian eigenvectors and parameter exploration.
result Large learning rates induce parameter instability, leading to better generalization through exploration of flatter regions of the loss landscape.
This work tackles GAN training instability through parallel tempering.
problem Training instability and mode collapse in GANs.
method Introduces a parallel tempering framework to stabilize GAN training.
result Significantly reduces gradient variance and improves training efficiency.
Proposes a continuous flow model to understand and control instability in gradient descent for deep learning.
problem Understanding and controlling the instability of gradient descent in deep learning.
method Introduces the Principal Flow (PF), a continuous time flow that approximates gradient descent dynamics.
result The PF captures divergent and oscillatory behaviors of gradient descent, including escaping local minima and saddle points.
NGRC shows numerical instabilities with short lags and high-degree polynomials.
problem Numerical instabilities in NGRC feature matrix.
method Combining numerical linear algebra and dynamical systems theory, we study feature matrix conditioning. We evaluate different numerical algorithms for solving the regularized least-squares problem.
result SVD-based training achieves accurate forecasts without regularization, preferable for short lags and high-degree polynomials.
New method stabilizes GAN training by solving ODEs.
problem Stability issues in GAN training.
method Solving ordinary differential equations (ODEs) to stabilize GAN training.
result Well-known ODE solvers can stabilize GAN training.
Improved ANN-based Monte Carlo simulation for Higgs decay events.
problem Accurate simulation of Higgs boson decay events.
method Monte Carlo simulation using an Artificial Neural Network (ANN) with improved training algorithm.
result The ANN simulation of Higgs decay is within 0.7% of the true value and achieves 26% unweighting efficiency.
Novel approach analyzes ReLU networks' training dynamics and proposes GmP for improved optimization.
problem Stochastic optimization instability in ReLU networks impedes convergence and generalization.
method Characteristic activation boundaries analysis and Geometric Parameterization (GmP) technique.
result GmP resolves instability, leading to better optimization, convergence, and generalization.
Study stabilizes adversarial training in neural networks over infinite-dimensional spaces.
problem Stability issues in adversarial training of neural networks.
method Functional analysis of minimax optimization over infinite-dimensional spaces of continuous functions and probability measures.
result Convergence property of minimax problems under certain conditions, interpreted as stabilization techniques.
Gradient flossing stabilizes RNN training by controlling Lyapunov exponents.
problem Gradient instability in RNNs leading to exploding and vanishing gradients.
method Regularizing Lyapunov exponents through backpropagation using differentiable linear algebra.
result Gradient flossing improves RNN training success rate and convergence speed.
We improve current instability-based methods for the selection of the number of clusters k in cluster analysis by developing a normalized cluster instability measure that corrects for the distribution of cluster sizes, a previously unaccounted driver of cluster instability. We show that our normalized instability mea…
We present an approach for efficiently training Gaussian Mixture Model (GMM) by Stochastic Gradient Descent (SGD) with non-stationary, high-dimensional streaming data. Our training scheme does not require data-driven parameter initialization (e.g., k-means) and can thus be trained based on a random initialization. Furt…
New diagnostics detect variability in individual risk estimates from machine learning models in healthcare.
problem Variability in individual risk estimates from machine learning models in healthcare, leading to unreliable treatment decisions.
method Proposed evaluation framework using empirical prediction interval width and empirical decision flip rate diagnostics.
result Randomness in optimization and initialization can lead to substantial individual-level variability in risk estimates, affecting clinical decisions.
Improved continuous-time consistency models for large-scale image generation.
problem Training instability and discretization errors in existing diffusion models.
method Unified theoretical framework, improved diffusion process, and network architecture.
result Trained continuous-time CMs at 1.5B parameters, achieving state-of-the-art FID scores.
Modeling HFT interactions reveals market instability.
problem Market instability caused by HFT dynamic coupling.
method Developed a recurrence relations framework to model HFT interactions.
result Unexpected latency and feedback can trigger market instability.
BERT fine-tuning is unstable due to optimization issues, not forgetting or dataset size.
problem Stability of fine-tuning BERT-based models across different random seeds.
method Analysis of BERT, RoBERTa, and ALBERT fine-tuned on GLUE datasets, identifying optimization difficulties as the cause of instability.
result Fine-tuning instability is due to optimization difficulties leading to vanishing gradients, not forgetting or dataset size.
New proof of instability for certain Einstein metrics.
problem Einstein metrics on specific 4-manifolds.
method Proving instability of conformally Kähler, Einstein metrics.
result Proven instability of certain Einstein metrics.
New method improves causal effect estimation by addressing imbalance in training data.
problem Imbalance between treatment and control groups in training data.
method Combines distributionally robust optimization and weight regularization.
result Consistent improvements over existing methods in experiments.
Study stability and instability of Ricci-flat metrics under generalized Ricci flow.
problem Stability and instability of Ricci-flat metrics under Ricci flow.
method Analysis of generalized Ricci flow for Ricci-flat metrics and vanishing 3-forms.
result Dynamical stability and instability results for Ricci-flat metrics and vanishing 3-forms.
Some exotic compact objects possess evanescent ergosurfaces: timelike submanifolds on which a Killing vector field, which is timelike everywhere else, becomes null. We show that any manifold possessing an evanescent ergosurface but no event horizon exhibits a linear instability of a peculiar kind: either there are solu…
Interval Neural Networks detect instabilities in image reconstructions.
problem Detecting instabilities in deep learning image reconstructions.
method Employed uncertainty quantification methods with Interval Neural Networks.
result Interval Neural Networks effectively reveal image reconstruction instabilities.
Note on instabilities in super-time-stepping methods for Heston model.
problem Instabilities in super-time-stepping methods applied to Heston model.
method Exploration of explicit super-time-stepping schemes (RK-Chebyshev, RK-Legendre) for Heston model.
result Relevance of stability remarks beyond super-time-stepping schemes.
Batch normalization prevents rank collapse in deep networks, improving training stability.
problem Rank collapse in randomly initialized deep networks with increasing depth.
method Investigates spectral instabilities in random matrices and uses batch normalization to avoid rank collapse.
result Batch normalization prevents rank collapse in both linear and ReLU networks, improving training stability.
Training recurrent neural networks (RNNs) on long sequence tasks is plagued with difficulties arising from the exponential explosion or vanishing of signals as they propagate forward or backward through the network. Many techniques have been proposed to ameliorate these issues, including various algorithmic and archite…
Binary perceptron's instability linked to replica symmetry breaking.
problem Understanding the relationship between algorithmic instability and replica symmetry breaking in binary perceptron learning.
method Established the connection between algorithmic instability and replica symmetry breaking by comparing the instability condition around the fixed point to the instability for breaking the replica symmetric solution of the free energy function.
result The instability condition around the algorithmic fixed point is identical to the instability for breaking the replica symmetric saddle point solution of the free energy function.
Feature Quantization improves GAN training stability.
problem Stability issues in GAN training.
method Feature Quantization (FQ) for the discriminator, embedding true and fake data into a shared discrete space.
result FQ-GAN achieves new state-of-the-art performance on various GAN tasks.
Generative Adversarial Networks are a new family of generative models, frequently used for generating photorealistic images. The theory promises for the GAN to eventually reach an equilibrium where generator produces pictures indistinguishable for the training set. In practice, however, a range of problems frequently p…
The study examines stability and instability of Poincaré-Einstein metrics using Ricci flow.
problem Stability and instability of Poincaré-Einstein metrics.
method Variant of expander entropy for asymptotically hyperbolic manifolds, local positive mass theorem, volume comparison.
result Characterization of stability and instability in terms of local positive mass theorem and volume comparison.
This paper improves forecast stability without sacrificing accuracy using dynamic loss weighting.
problem Rolling origin forecast instability in time series forecasting.
method Dynamic loss weighting algorithms applied to the N-BEATS model.
result Dynamic loss weighting can further improve forecast stability without compromising accuracy.
Fractal learning rate schedules accelerate vanilla gradient descent.
problem Difficulty in tuning learning rates in iterative optimization.
method Introduce Chebyshev learning rate schedule for gradient descent.
result Locally unstable updates can lead to convergence in deep learning.
Study shows instability of naked singularities in perfect fluid models.
problem Instability of naked singularities in Einstein equations coupled with isothermal perfect fluid.
method Investigated spherically symmetric self-similar naked singularities under C1,α perturbations of an external massless scalar field. result Spherically symmetric self-similar naked singularities are unstable to trapped surface formation.
Study on spectral stability of Riemannian coverings.
problem Stability of eigenvalues in Riemannian coverings.
method Analysis of Laplacian eigenvalues under finite coverings.
result Necessary conditions for spectral stability or instability.
Clinical models can be unstable, leading to unreliable predictions.
problem Stability of clinical prediction models developed using statistical or machine learning methods.
method Simulation and case studies of statistical and machine learning approaches to show instability in model predictions.
result Model instability often leads to miscalibration of predictions in new data.
A data-driven approach predicts morphological development under structural instability.
problem Understanding and predicting spatiotemporal complexities of morphogenesis under structural instability.
method Machine-learning framework based on physical modeling of morphogenesis.
result Identification of key bifurcation characteristics and prediction of history-dependent development.
In this paper we consider the training of single hidden layer neural networks by pseudoinversion, which, in spite of its popularity, is sometimes affected by numerical instability issues. Regularization is known to be effective in such cases, so that we introduce, in the framework of Tikhonov regularization, a matricia…
Study shows instability of naked singularities in scalar field models.
problem Stability of naked singularities in spherically symmetric Einstein-Scalar field systems.
method Analysis of a family of incoming null cones becoming increasingly singular.
result Naked singularities are unstable to black hole formation under certain perturbations.
Study shows instability of certain MOTSs with continuous symmetry.
problem Stability of MOTSs with continuous symmetry.
method Analysis of initial data sets with continuous symmetry and non-preserved MOTSs.
result Exotic MOTSs are unstable except in exceptional cases.
New method selects data points for better model performance.
problem Balancing input coverage and model utility in selective prediction.
method Study training dynamics to reject inputs with unstable predictions.
result State-of-the-art selective prediction performance achieved without model modifications.
One of the challenges in the study of generative adversarial networks is the instability of its training. In this paper, we propose a novel weight normalization technique called spectral normalization to stabilize the training of the discriminator. Our new normalization technique is computationally light and easy to in…
Deep convolutional neural networks are known to be unstable during training at high learning rate unless normalization techniques are employed. Normalizing weights or activations allows the use of higher learning rates, resulting in faster convergence and higher test accuracy. Batch normalization requires minibatch sta…