Natural gradients boost performance in non-conjugate Gaussian process models.
problem Improving inference in non-conjugate Gaussian process models.
method Use of natural gradients in non-conjugate stochastic settings with hyperparameter learning.
result Natural gradients significantly improve performance, especially for ill-conditioned posteriors.
TANGO optimizes models with small learning rates, converging to natural gradient.
problem Optimizing models with small learning rates to converge to natural gradient.
method TANGO, a simple algorithm that converges to natural gradient descent.
result TANGO achieves natural gradient descent with small learning rates.
A new natural gradient accounts for correlated variational parameters in variational inference.
problem Traditional natural gradients fail to correct for correlations in variational inference.
method Construct a new natural gradient called the Variational Predictive Natural Gradient (VPNG).
result VPNG accounts for the relationship between model parameters and variational parameters.
Natural gradient simplification for deep learning networks.
problem Efficiency in training deep Bayesian networks.
method Analysis of two geometries of Fisher information matrix and development of a method to simplify natural gradient for the second geometry.
result A method to simplify natural gradient for deep networks using an auxiliary recognition model.
A framework for natural gradient with arbitrary similarity measures.
problem Unclear metric for natural gradient in non-Euclidean spaces.
method Derive a metric for natural gradient given an arbitrary similarity measure.
result General framework for natural gradient in non-Euclidean spaces.
Square-root natural-gradient improves variational inference convergence.
problem Challenges in establishing theoretical convergence guarantees for natural-gradient descent.
method Square-root parameterization for Gaussian covariance.
result Establishes novel convergence guarantees for natural-gradient Gaussian inference.
We extend natural-gradient methods to mixtures of exponential-family distributions, improving inference speed.
problem Complex, multimodal posterior distributions are difficult to approximate with simple exponential-family distributions.
method We use minimal conditional-EF representations and derive simple natural-gradient updates.
result Our natural-gradient method converges faster than black-box methods with reparameterization gradients.
Natural gradient improves deep Q-learning performance.
problem Improving deep Q-learning stability and performance.
method Integrates natural-gradient techniques into deep Q-learning.
result Natural-gradient deep Q-learning (NGDQN) outperforms standard DQN without target networks and performs similarly to DQN with target networks.
Natural-gradient methods improve Bayesian inference in complex models.
problem Computational challenges in Bayesian inference for complex models.
method Derive fast natural-gradient updates for variational inference.
result Natural-gradient methods provide more accurate local approximations.
Natural gradient is equivalent to Kalman filtering for parameter estimation.
problem Parameter estimation in probabilistic models from observations.
method Casting natural gradient as an extended Kalman filter for parameter estimation.
result Exact algebraic correspondence between natural gradient and Kalman filtering.
New methods using natural gradient for structured optimization.
problem Structured optimization problems.
method Structured second-order methods via natural gradient descent.
result Efficiency demonstrated on non-convex and deep learning problems.
The natural gradient allows for more efficient gradient descent by removing dependencies and biases inherent in a function's parameterization. Several papers present the topic thoroughly and precisely. It remains a very difficult idea to get your head around however. The intent of this note is to provide simple intuiti…
A new method uses natural gradients for efficient distribution optimization.
problem Challenges in computing natural gradients for many distributions.
method Reframe optimization as a surrogate distribution with easy natural gradient computation.
result Expands set of distributions efficiently targetable with natural gradients.
Trust-region methods and natural gradients are equivalent in certain policy search scenarios.
problem Improving policy search methods in continuous control tasks.
method Introducing compatible policy search (COPOS) that uses natural parameterization and compatible value function approximation to control entropy loss.
result COPOS yields state-of-the-art results in challenging tasks and reduces entropy loss.
Natural gradient for Wasserstein metric approximated using kernel methods.
problem Optimization of cost functionals over probability distributions.
method Kernelized Wasserstein metric, natural gradient approximation.
result Effective gradient estimator for Wasserstein metric with theoretical guarantees.
Natural gradient optimization improves model parameter estimation in graphical models.
problem Estimating model parameters in graphical models.
method Reformulated as an information geometric optimization problem, introduced natural gradient descent strategy.
result Natural gradient strategy leads to optimal parameter learning without fitting an incorrect distribution.
The natural gradient of ELBO vanishes in unconstrained optimization, simplifying learning.
problem The gap between evidence and ELBO has a vanishing natural gradient.
method Analyzes the Fisher-Rao gradient of ELBO and its implications for learning.
result Maximizing ELBO is equivalent to minimizing KL divergence, simplifying learning.
Quantum Natural Gradient uses quantum geometry for optimization.
problem Optimizing variational quantum circuits efficiently.
method Quantum generalization of Natural Gradient Descent using Quantum Information Geometry.
result Efficient algorithm for computing metric tensor approximations.
Natural gradient descent avoids the magic of model parametrization, leading to different optimization outcomes.
problem Understanding the impact of model parametrization on optimization and generalization in deep learning.
method Characterization of natural gradient flow in deep linear networks and nonlinear neural networks.
result Natural gradient descent fails to generalize in some cases, while gradient descent with the right architecture performs well.
Optimizes graph neural networks using natural gradient descent.
problem Improving efficiency and performance of graph neural networks.
method Employing natural gradient descent to optimize graph neural networks.
result Natural gradient optimization leads to superior performance compared to existing methods.
Researchers propose a non-monotone quantum natural gradient for quantum systems.
problem Applying natural gradient methods to quantum systems without monotonicity.
method Introducing a non-monotone quantum natural gradient (QNG) and demonstrating its superiority over conventional QNG.
result Non-monotone QNG outperforms conventional QNG in terms of convergence speed.
VB uses natural gradients in information geometry.
problem Estimating or computing natural gradients in VB.
method Natural-gradient descent algorithm and Bayesian Learning Rule.
result Simplification of Bayes' rule and generalization of quadratic surrogates.
Natural gradient descent speeds up convergence in neural networks, especially with overparameterization.
problem Mitigating the effects of curvature in neural network optimization.
method Analysis of natural gradient descent on nonlinear neural networks with stability conditions.
result Natural gradient descent converges efficiently under specific conditions for overparameterized networks.
Improved VI method for deep mixed models in finance.
problem Inaccurate and slow variational inference in high dimensions.
method Natural gradient hybrid VI method targeting joint posterior.
result Natural gradient method is faster and more accurate than existing methods.
Improved neural network inference with eigenvalue correction.
problem Inference of flexible variational posteriors is computationally expensive.
method Eigenvalue correction to matrix-variate Gaussian posterior.
result Empirically, the method outperforms existing algorithms.
Noisy natural gradient improves variational inference for Bayesian neural nets.
problem Tradeoff between simple and complex variational families in Bayesian neural nets.
method Adaptive weight noise in natural gradient ascent to implicitly fit variational posteriors.
result Noisy natural gradient algorithms can train full-covariance variational posteriors efficiently.
We develop a coordinate-free approach to natural gradient descent for scalable neural networks.
problem First-order optimization methods are sensitive to model parameterization.
method We construct a coordinate-free natural gradient and analyze its invariance properties for K-FAC.
result K-FAC's natural gradient matches the coordinate-free update, maintaining invariance to affine transformations.
Optimal transport natural gradient improves optimization in statistical models.
problem Improving optimization in statistical models with continuous sample spaces.
method Pulling back the Wasserstein metric tensor to a parameter space, creating a Riemannian manifold.
result Natural gradient descent outperforms standard gradient descent in Wasserstein distance optimization.
Unified perspective on natural gradient methods for GMMs, improving variational inference.
problem Efficiently learning multi-modal approximations of complex distributions.
method Comparison and optimization of VIPS and iBayes-GMM methods for Gaussian mixture models.
result Hybrid approach significantly outperforms both VIPS and iBayes-GMM.
A framework connects Taylor methods with Fisher-efficient NG.
problem Combining Taylor-based methods with Fisher-efficient NG.
method Constructs a theoretical framework linking Taylor approximation and NG.
result Mathematical justification for combining higher order methods with NG.
NGBoost boosts probabilistic predictions using natural gradients.
problem Uncertainty estimation in probabilistic predictions.
method Gradient boosting for probabilistic regression with natural gradient correction.
result NGBoost outperforms existing methods for probabilistic prediction.
Natural gradient learning improves synaptic plasticity in spiking neurons.
problem Parametrization dependence leads to inconsistencies in classical synaptic plasticity theories.
method Proposes natural gradient descent in Riemannian geometry for spiking neurons.
result Derives a synaptic learning rule that explains biological phenomena.
Improved natural gradient boosting with leaf number clipping for faster and better performance.
problem Slower training speed and poor performance on large datasets for natural gradient boosting.
method Leaf number clipping regularization to optimize hyperparameters and improve performance.
result Significant improvement in performance and up to 4.85x speed up on various datasets.
Natural gradient descent is an optimization method traditionally motivated from the perspective of information geometry, and works well for many applications as an alternative to stochastic gradient descent. In this paper we critically analyze this method and its properties, and show how it can be viewed as a type of 2…
Quantum method speeds up VB estimation in machine learning.
problem Prohibitively expensive natural gradient in high dimensions.
method Regression-based natural gradient estimation with quantum matrix inversion.
result Quantum method enables efficient VB estimation.
Efficient methods for training deep neural networks using subsampled Gauss-Newton and natural gradient.
problem Training deep neural networks with large datasets and variables.
method Subsampled Gauss-Newton and natural gradient methods with subsampled gradient estimates.
result Methods converge to a stationary point and are efficient to implement.
A new black-box optimizer using implicit natural gradient.
problem Efficient optimization for complex, computationally intensive problems.
method Stochastic update with implicit natural gradient of an exponential-family distribution.
result Theoretical convergence rate for convex functions and continuous non-differentiable functions.
Disputes the empirical Fisher approximation for natural gradient descent.
problem The empirical Fisher approximation fails to capture second-order information in general.
method Comparison of empirical Fisher and Fisher information matrices.
result The empirical Fisher does not generally approximate the Fisher or Hessian.
The paper proposes a method to train time-varying generative models using natural gradients.
problem Training time-varying generative models efficiently and accurately.
method Projecting generative model parameters onto an exponential family manifold and optimizing using natural gradient descent.
result The proposed method efficiently approximates the natural gradient and can be applied to various exponential family models.
Paper tackles anomaly detection in e-commerce using Bayesian semi-supervised tensor decomposition.
problem Detecting anomalies in seller-reviewer data in e-commerce.
method Bayesian semi-supervised tensor decomposition with Polya-Gamma data augmentation and partial natural gradient learning.
result Semi-supervised approach outperforms state-of-the-art unsupervised baselines.
NGD improves multivariate Gaussian inference by optimizing Fisher information.
problem Efficiently optimizing multivariate Gaussian models.
method Natural Gradient Descent applied to multivariate Gaussian parameters.
result NGD updates are more efficient for symmetric covariance matrices.
Paper extends option-critic architecture to estimate natural gradient for reinforcement learning.
problem Estimating natural gradient in hierarchical reinforcement learning.
method Introduces natural option critic algorithm to estimate natural gradient for option's policy and termination function.
result Improves over vanilla gradient approach in experimental results.
Inversion-free natural gradient method for Riemannian manifolds.
problem Hindered by the need for Euclidean space, Fisher information matrix inversion, and computational cost.
method Intrinsic, inversion-free natural gradient method on Riemannian manifolds, using moving approximation of inverse FIM.
result Almost-sure convergence rates and sub-quadratic storage complexity for large-scale applications.
Improves understanding of stochastic NGVI convergence rates.
problem Lack of knowledge about non-asymptotic convergence rates in stochastic NGVI.
method Proved non-asymptotic convergence rates for conjugate likelihoods and showed implicit optimization for non-conjugate likelihoods.
result First O(T1) non-asymptotic convergence rate for stochastic NGVI in conjugate likelihoods. A new batch optimisation framework using Natural Gradient improves DNN acoustic models.
problem Optimizing DNN acoustic models for better word error rate approximation.
method Proposes a Natural Gradient (NG) approach to sequence training, correcting the gradient based on local curvature of KL-divergence.
result The NG method converges more quickly and can be applied to any sequence discriminative training criterion.
QBVI uses natural gradients for efficient Bayesian learning.
problem Efficient Bayesian learning in complex models.
method Natural gradient updates in a black-box framework for exponential-family distributions.
result QBVI framework is effective for a wide range of Bayesian inference problems.
Proposes a new stochastic optimization method for MLR models.
problem Slow convergence of SGD in big data scenarios.
method Dual Stochastic Natural Gradient Descent (DNSGD) based on manifold optimization.
result DNSGD converges and has linear computational complexity.
BONG optimizes Bayesian inference online with natural gradient descent.
problem Sequential Bayesian inference in online settings.
method Bayesian online natural gradient (BONG) approach based on variational Bayes.
result BONG outperforms other online VB methods in non-conjugate settings.