A new natural gradient accounts for correlated variational parameters in variational inference.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Natural gradient optimization improves model parameter estimation in graphical models.
We prove that the natural principal parameters on a given Weingarten surface are also natural principal parameters for the parallel surfaces of the given one. As a consequence of this result we obtain that the natural PDE of any Weingarten surface is the natural PDE of its parallel surfaces. We show that the linear fra…
New insights into natural exponential families improve regret bounds for bandit problems.
A framework for natural gradient with arbitrary similarity measures.
We develop a more efficient NGD method for structured parameters.
Black box discrete optimization (BBDO) appears in wide range of engineering tasks. Evolutionary or other BBDO approaches have been applied, aiming at automating necessary tuning of system parameters, such as hyper parameter tuning of machine learning based systems when being installed for a specific task. However, auto…
We cast Amari's natural gradient in statistical learning as a specific case of Kalman filtering. Namely, applying an extended Kalman filter to estimate a fixed unknown parameter of a probabilistic model from a series of observations, is rigorously equivalent to estimating this parameter via an online stochastic natural…
A method for converting NIW parameters for better estimation.
We assume a vector bundle with a general linear connection and a classical linear connection $\Lam$ on . We prove that all classical linear connections on the total space naturally given by $(\Lam, K)$ form a 15-parameter family. Further we prove that all connections on naturally given by…
Inversion-free natural gradient method for Riemannian manifolds.
New method uses fewer parameters to match state-of-the-art performance on multiple natural language tasks.
This paper presents sampling-based speech parameter generation using moment-matching networks for Deep Neural Network (DNN)-based speech synthesis. Although people never produce exactly the same speech even if we try to express the same linguistic and para-linguistic information, typical statistical speech synthesis pr…
Improved inference for heterogeneous multi-output Gaussian processes using natural gradient optimization.
New NPG variants ensure parameter convergence in multi-agent learning.
This work proposes a new method for variational inference using Wasserstein gradient descent.
A new method uses natural gradients for efficient distribution optimization.
Improved VI method for deep mixed models in finance.
AOPU stabilizes NN training by approximating natural gradient, improving stability and convergence.
Minimal surfaces of general type in Euclidean 4-space are characterized with the conditions that the ellipse of curvature at any point is centered at this point and has two different principal axes. Any minimal surface of general type locally admits geometrically determined parameters - canonical parameters. In such pa…
We study the existence of natural and projectively equivariant quantizations for differential operators acting between order 1 vector bundles over a smooth manifold M. To that aim, we make use of the Thomas-Whitehead approach of projective structures and construct a Casimir operator depending on a projective Cartan con…
Improves SVGP methods for faster and more accurate Gaussian process inference.
Stochastic gradient methods enable learning probabilistic models from large amounts of data. While large step-sizes (learning rates) have shown to be best for least-squares (e.g., Gaussian noise) once combined with parameter averaging, these are not leading to convergent algorithms in general. In this paper, we conside…
Proposes an alternative method to train RBMs with binary synapses using Bayesian learning rule.
NGBoost boosts probabilistic predictions using natural gradients.
On any timelike surface with zero mean curvature in the four-dimensional Minkowski space we introduce special geometric (canonical) parameters and prove that the Gauss curvature and the normal curvature of the surface satisfy a system of two natural partial differential equations. Conversely, any two solutions to this …
We solve the mean parametrization of von Mises-Fisher distribution.
Bayesian active learning tackles nuisance parameters, leading to bias and dilemmas.
Many tasks in natural language understanding require learning relationships between two sequences for various tasks such as natural language inference, paraphrasing and entailment. These aforementioned tasks are similar in nature, yet they are often modeled individually. Knowledge transfer can be effective for closely …
New perspective on CNNs using Hessian maps reveals their structure.
The paper provides a Weierstrass representation for maximal space-like surfaces in 4D pseudo-Euclidean space.
Classifies geodesic orbit spaces with abelian isotropy subgroups.
New methods using natural gradient for structured optimization.
The recently proposed option-critic architecture Bacon et al. provide a stochastic policy gradient approach to hierarchical reinforcement learning. Specifically, they provide a way to estimate the gradient of the expected discounted return with respect to parameters that define a finite number of temporally extended ac…
The paper finds a Weierstrass representation for a specific type of Lorentzian minimal surface.
Proposes a new stochastic optimization method for MLR models.
Efficiently learns exponential family distributions with i.i.d. samples.
We study geodesics of the form , $X,Y\in \fr{g}=\operatorname{Lie}(G)$, in homogeneous spaces , where is the natural projection. These curves naturally generalise homogeneous geodesics, that is orbits of one-parameter subgroups of (i.e. , $X\in …
The natural gradient allows for more efficient gradient descent by removing dependencies and biases inherent in a function's parameterization. Several papers present the topic thoroughly and precisely. It remains a very difficult idea to get your head around however. The intent of this note is to provide simple intuiti…
Using the fact that any minimal strongly regular surface carries locally canonical principal parameters, we obtain a canonical representation of these surfaces, which makes more precise the Weierstrass representation in canonical principal parameters. This allows us to describe locally the solutions of the natural part…
Improved learning of probabilistic box embeddings by modeling parameters with Gumbel distributions.
Develops parameter-free online mirror descent for optimal dynamic regret.
Improved Bayesian learning rule handles positive-definite constraints efficiently.
On a main class of the almost contact manifolds with B-metric, it is described the family of the linear connections preserving the manifold's structures by 4 parameters. In this family there are determined the canonical-type connection and the connection with zero parameters.
In natural language processing, a lot of the tasks are successfully solved with recurrent neural networks, but such models have a huge number of parameters. The majority of these parameters are often concentrated in the embedding layer, which size grows proportionally to the vocabulary length. We propose a Bayesian spa…
We describe the neural-network training framework used in the Kaldi speech recognition toolkit, which is geared towards training DNNs with large amounts of training data using multiple GPU-equipped or multi-core machines. In order to be as hardware-agnostic as possible, we needed a way to use multiple machines without …
Unified method visualizes curvature on curves and surfaces.
We introduce a simple algorithm, True Asymptotic Natural Gradient Optimization (TANGO), that converges to a true natural gradient descent in the limit of small learning rates, without explicit Fisher matrix estimation. For quadratic models the algorithm is also an instance of averaged stochastic gradient, where the par…