Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

148296443591 · Jun 202019922001200920172026
48 results for Natural Parameters

Variational inference transforms posterior inference into parametric optimization thereby enabling the use of latent variable models where otherwise impractical. However, variational inference can be finicky when different variational parameters control variables that are strongly correlated under the model. Traditiona…

2019-03-07abs ↗pdf ↗

New insights into natural exponential families improve regret bounds for bandit problems.

problem Improving regret bounds for bandit problems with subexponential tails.
method Proving self-concordance for natural exponential families and applying to bandits.
result Optimistic algorithms for generalized linear bandits have second-order regret bounds that are free of an exponential dependence on problem parameters.

We cast Amari's natural gradient in statistical learning as a specific case of Kalman filtering. Namely, applying an extended Kalman filter to estimate a fixed unknown parameter of a probabilistic model from a series of observations, is rigorously equivalent to estimating this parameter via an online stochastic natural…

2017-03-01abs ↗pdf ↗

We assume a vector bundle p:EMp: E\to M with a general linear connection KK and a classical linear connection $\Lam$ on MM. We prove that all classical linear connections on the total space EE naturally given by $(\Lam, K)$ form a 15-parameter family. Further we prove that all connections on J1EJ^1 E naturally given by…

2004-10-21abs ↗pdf ↗

Inversion-free natural gradient method for Riemannian manifolds.

problem Hindered by the need for Euclidean space, Fisher information matrix inversion, and computational cost.
method Intrinsic, inversion-free natural gradient method on Riemannian manifolds, using moving approximation of inverse FIM.
result Almost-sure convergence rates and sub-quadratic storage complexity for large-scale applications.

This paper deals with estimating model parameters in graphical models. We reformulate it as an information geometric optimization problem and introduce a natural gradient descent strategy that incorporates additional meta parameters. We show that our approach is a strong alternative to the celebrated EM approach for le…

2019-05-14abs ↗pdf ↗

New NPG variants ensure parameter convergence in multi-agent learning.

problem Non-convergence of parameters in NPG for multi-agent learning.
method Proposed variants of NPG for multi-agent learning scenarios.
result Global last-iterate parameter convergence guarantees in various multi-agent learning settings.

This work proposes a new method for variational inference using Wasserstein gradient descent.

problem Optimizing variational parameters to match a true posterior distribution.
method Reinterpreting VI as an optimization problem over a variational parameter space, using Wasserstein gradient descent.
result The proposed Wasserstein gradient descent can be seen as a generalization of existing optimization techniques in VI.

A new method uses natural gradients for efficient distribution optimization.

problem Challenges in computing natural gradients for many distributions.
method Reframe optimization as a surrogate distribution with easy natural gradient computation.
result Expands set of distributions efficiently targetable with natural gradients.

AOPU stabilizes NN training by approximating natural gradient, improving stability and convergence.

problem Stability and interpretability in online NN training for industrial soft sensors.
method AOPU truncates gradient backpropagation, optimizing trackable parameters, and approximating natural gradient.
result AOPU achieves stable convergence and superior performance on chemical process datasets.

We study the existence of natural and projectively equivariant quantizations for differential operators acting between order 1 vector bundles over a smooth manifold M. To that aim, we make use of the Thomas-Whitehead approach of projective structures and construct a Casimir operator depending on a projective Cartan con…

2006-01-21abs ↗pdf ↗

Proposes an alternative method to train RBMs with binary synapses using Bayesian learning rule.

problem Training RBMs with binary synapses is challenging due to discrete nature of synapses.
method Proposes an alternative optimization method using the Bayesian learning rule, updating natural parameters instead of expectation parameters.
result No additional clipping is needed as natural parameters take values in the entire real domain.

On any timelike surface with zero mean curvature in the four-dimensional Minkowski space we introduce special geometric (canonical) parameters and prove that the Gauss curvature and the normal curvature of the surface satisfy a system of two natural partial differential equations. Conversely, any two solutions to this …

2011-11-18abs ↗pdf ↗

We solve the mean parametrization of von Mises-Fisher distribution.

problem No closed-form normalization function for mean parameters exists.
method Derived a second-order ODE for mean normalizer and provided approximations.
result Rapid evaluation of densities and natural parameters in terms of mean parameters.

Bayesian active learning tackles nuisance parameters, leading to bias and dilemmas.

problem Bayesian active learning with nuisance parameters leads to bias and dilemmas.
method Characterizes and mitigates negative interference by accurately estimating nuisance parameters.
result The extent of negative interference can be extremely large, and accurate estimation of nuisance parameters is critical.

Many tasks in natural language understanding require learning relationships between two sequences for various tasks such as natural language inference, paraphrasing and entailment. These aforementioned tasks are similar in nature, yet they are often modeled individually. Knowledge transfer can be effective for closely …

2018-04-23abs ↗pdf ↗

New perspective on CNNs using Hessian maps reveals their structure.

problem Understanding the nature of Convolutional Neural Networks (CNNs).
method Developed a framework using Toeplitz representation of CNNs to reveal Hessian structure and prove rank bounds.
result Proved that the Hessian rank of CNNs grows as the square root of the number of parameters.

Classifies geodesic orbit spaces with abelian isotropy subgroups.

problem Characterizing and classifying geodesic orbit spaces with specific isotropy subgroups.
method Simplified study of geodesic orbit metrics on G/S by reducing to submanifolds and generalized flag manifolds, using properties of root systems.
result Geodesic orbit spaces of the form (G/S,g) are naturally reductive.

The recently proposed option-critic architecture Bacon et al. provide a stochastic policy gradient approach to hierarchical reinforcement learning. Specifically, they provide a way to estimate the gradient of the expected discounted return with respect to parameters that define a finite number of temporally extended ac…

2018-12-04abs ↗pdf ↗

The paper finds a Weierstrass representation for a specific type of Lorentzian minimal surface.

problem Minimal Lorentzian surfaces in R24\mathbb{R}^4_2 with certain curvature conditions.
method Weierstrass representation with respect to isothermal and canonical parameters.
result Explicit solution to the system of natural PDEs for general type surfaces.

Efficiently learns exponential family distributions with i.i.d. samples.

problem Learning natural parameters of truncated exponential families efficiently.
method Proposes a novel loss function and computationally efficient estimator.
result Achieves optimal sample complexity and asymptotic normality.

We study geodesics of the form γ(t)=π(exp(tX)exp(tY))γ(t)=π(\exp(tX)\exp(tY)), $X,Y\in \fr{g}=\operatorname{Lie}(G)$, in homogeneous spaces G/KG/K, where π:GG/Kπ:G\rightarrow G/K is the natural projection. These curves naturally generalise homogeneous geodesics, that is orbits of one-parameter subgroups of GG (i.e. γ(t)=π(exp(tX))γ(t)=π(\exp (tX)), $X\in …

2016-11-14abs ↗pdf ↗

Improved learning of probabilistic box embeddings by modeling parameters with Gumbel distributions.

problem Local identifiability issues in geometric embeddings.
method Modeling box parameters with min and max Gumbel distributions, calculating expected intersection volume.
result Improves the ability of probabilistic box embeddings to learn.

Improved Bayesian learning rule handles positive-definite constraints efficiently.

problem Bayesian learning rule struggles with positive-definite constraints.
method Proposes an improved rule using Riemannian gradient methods for block-coordinate natural parameterization.
result Outperforms existing methods without increased computation.

In natural language processing, a lot of the tasks are successfully solved with recurrent neural networks, but such models have a huge number of parameters. The majority of these parameters are often concentrated in the embedding layer, which size grows proportionally to the vocabulary length. We propose a Bayesian spa…

2018-10-25abs ↗pdf ↗

We introduce a simple algorithm, True Asymptotic Natural Gradient Optimization (TANGO), that converges to a true natural gradient descent in the limit of small learning rates, without explicit Fisher matrix estimation. For quadratic models the algorithm is also an instance of averaged stochastic gradient, where the par…

2017-12-22abs ↗pdf ↗

We present Natural Gradient Boosting (NGBoost), an algorithm for generic probabilistic prediction via gradient boosting. Typical regression models return a point estimate, conditional on covariates, but probabilistic regression models output a full probability distribution over the outcome space, conditional on the cov…

2019-10-08abs ↗pdf ↗

We introduce a new embarrassingly parallel parameter learning algorithm for Markov random fields with untied parameters which is efficient for a large class of practical models. Our algorithm parallelizes naturally over cliques and, for graphs of bounded degree, its complexity is linear in the number of cliques. Unlike…

2013-08-29abs ↗pdf ↗