Variational inference transforms posterior inference into parametric optimization thereby enabling the use of latent variable models where otherwise impractical. However, variational inference can be finicky when different variational parameters control variables that are strongly correlated under the model. Traditiona…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We prove that the natural principal parameters on a given Weingarten surface are also natural principal parameters for the parallel surfaces of the given one. As a consequence of this result we obtain that the natural PDE of any Weingarten surface is the natural PDE of its parallel surfaces. We show that the linear fra…
New insights into natural exponential families improve regret bounds for bandit problems.
We develop a more efficient NGD method for structured parameters.
In optimization, the natural gradient method is well-known for likelihood maximization. The method uses the Kullback-Leibler divergence, corresponding infinitesimally to the Fisher-Rao metric, which is pulled back to the parameter space of a family of probability distributions. This way, gradients with respect to the p…
Black box discrete optimization (BBDO) appears in wide range of engineering tasks. Evolutionary or other BBDO approaches have been applied, aiming at automating necessary tuning of system parameters, such as hyper parameter tuning of machine learning based systems when being installed for a specific task. However, auto…
We cast Amari's natural gradient in statistical learning as a specific case of Kalman filtering. Namely, applying an extended Kalman filter to estimate a fixed unknown parameter of a probabilistic model from a series of observations, is rigorously equivalent to estimating this parameter via an online stochastic natural…
We assume a vector bundle with a general linear connection and a classical linear connection $\Lam$ on . We prove that all classical linear connections on the total space naturally given by $(\Lam, K)$ form a 15-parameter family. Further we prove that all connections on naturally given by…
A method for converting NIW parameters for better estimation.
Inversion-free natural gradient method for Riemannian manifolds.
This paper deals with estimating model parameters in graphical models. We reformulate it as an information geometric optimization problem and introduce a natural gradient descent strategy that incorporates additional meta parameters. We show that our approach is a strong alternative to the celebrated EM approach for le…
This paper presents sampling-based speech parameter generation using moment-matching networks for Deep Neural Network (DNN)-based speech synthesis. Although people never produce exactly the same speech even if we try to express the same linguistic and para-linguistic information, typical statistical speech synthesis pr…
New NPG variants ensure parameter convergence in multi-agent learning.
This work proposes a new method for variational inference using Wasserstein gradient descent.
A new method uses natural gradients for efficient distribution optimization.
Improved VI method for deep mixed models in finance.
It is known that any maximal space-like surface without isotropic points in the four-dimensional pseudo-Euclidean space with neutral metric admits locally geometric parameters which are special case of isothermal parameters. With respect to such parameters the surface is determined uniquely up to a motion by the Gauss …
AOPU stabilizes NN training by approximating natural gradient, improving stability and convergence.
Minimal surfaces of general type in Euclidean 4-space are characterized with the conditions that the ellipse of curvature at any point is centered at this point and has two different principal axes. Any minimal surface of general type locally admits geometrically determined parameters - canonical parameters. In such pa…
We study the existence of natural and projectively equivariant quantizations for differential operators acting between order 1 vector bundles over a smooth manifold M. To that aim, we make use of the Thomas-Whitehead approach of projective structures and construct a Casimir operator depending on a projective Cartan con…
Improves SVGP methods for faster and more accurate Gaussian process inference.
Stochastic gradient methods enable learning probabilistic models from large amounts of data. While large step-sizes (learning rates) have shown to be best for least-squares (e.g., Gaussian noise) once combined with parameter averaging, these are not leading to convergent algorithms in general. In this paper, we conside…
Proposes an alternative method to train RBMs with binary synapses using Bayesian learning rule.
On any timelike surface with zero mean curvature in the four-dimensional Minkowski space we introduce special geometric (canonical) parameters and prove that the Gauss curvature and the normal curvature of the surface satisfy a system of two natural partial differential equations. Conversely, any two solutions to this …
We solve the mean parametrization of von Mises-Fisher distribution.
Bayesian active learning tackles nuisance parameters, leading to bias and dilemmas.
Many tasks in natural language understanding require learning relationships between two sequences for various tasks such as natural language inference, paraphrasing and entailment. These aforementioned tasks are similar in nature, yet they are often modeled individually. Knowledge transfer can be effective for closely …
New perspective on CNNs using Hessian maps reveals their structure.
A recent novel extension of multi-output Gaussian processes handles heterogeneous outputs assuming that each output has its own likelihood function. It uses a vector-valued Gaussian process prior to jointly model all likelihoods' parameters as latent functions drawn from a Gaussian process with a linear model of coregi…
Classifies geodesic orbit spaces with abelian isotropy subgroups.
New methods using natural gradient for structured optimization.
The recently proposed option-critic architecture Bacon et al. provide a stochastic policy gradient approach to hierarchical reinforcement learning. Specifically, they provide a way to estimate the gradient of the expected discounted return with respect to parameters that define a finite number of temporally extended ac…
The paper finds a Weierstrass representation for a specific type of Lorentzian minimal surface.
Efficiently learns exponential family distributions with i.i.d. samples.
We study geodesics of the form , $X,Y\in \fr{g}=\operatorname{Lie}(G)$, in homogeneous spaces , where is the natural projection. These curves naturally generalise homogeneous geodesics, that is orbits of one-parameter subgroups of (i.e. , $X\in …
The natural gradient allows for more efficient gradient descent by removing dependencies and biases inherent in a function's parameterization. Several papers present the topic thoroughly and precisely. It remains a very difficult idea to get your head around however. The intent of this note is to provide simple intuiti…
Using the fact that any minimal strongly regular surface carries locally canonical principal parameters, we obtain a canonical representation of these surfaces, which makes more precise the Weierstrass representation in canonical principal parameters. This allows us to describe locally the solutions of the natural part…
The multinomial logistic regression (MLR) model is widely used in statistics and machine learning. Stochastic gradient descent (SGD) is the most common approach for determining the parameters of a MLR model in big data scenarios. However, SGD has slow sub-linear rates of convergence. A way to improve these rates of con…
Improved learning of probabilistic box embeddings by modeling parameters with Gumbel distributions.
Develops parameter-free online mirror descent for optimal dynamic regret.
On a main class of the almost contact manifolds with B-metric, it is described the family of the linear connections preserving the manifold's structures by 4 parameters. In this family there are determined the canonical-type connection and the connection with zero parameters.
Improved Bayesian learning rule handles positive-definite constraints efficiently.
In natural language processing, a lot of the tasks are successfully solved with recurrent neural networks, but such models have a huge number of parameters. The majority of these parameters are often concentrated in the embedding layer, which size grows proportionally to the vocabulary length. We propose a Bayesian spa…
We describe the neural-network training framework used in the Kaldi speech recognition toolkit, which is geared towards training DNNs with large amounts of training data using multiple GPU-equipped or multi-core machines. In order to be as hardware-agnostic as possible, we needed a way to use multiple machines without …
Unified method visualizes curvature on curves and surfaces.
We introduce a simple algorithm, True Asymptotic Natural Gradient Optimization (TANGO), that converges to a true natural gradient descent in the limit of small learning rates, without explicit Fisher matrix estimation. For quadratic models the algorithm is also an instance of averaged stochastic gradient, where the par…
We present Natural Gradient Boosting (NGBoost), an algorithm for generic probabilistic prediction via gradient boosting. Typical regression models return a point estimate, conditional on covariates, but probabilistic regression models output a full probability distribution over the outcome space, conditional on the cov…
We introduce a new embarrassingly parallel parameter learning algorithm for Markov random fields with untied parameters which is efficient for a large class of practical models. Our algorithm parallelizes naturally over cliques and, for graphs of bounded degree, its complexity is linear in the number of cliques. Unlike…