Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

22436586 · May 202619922001200920172026
48 results for Amari's $\alpha$-divergence

Improved Bayesian inference using power priors with historical data.

problem Improving Bayesian inference with historical data.
method Generalized power priors that adapt to the α\alpha parameter of Amari's α\alpha-divergence.
result Improved performance through appropriate choices of the α\alpha parameter.

Characterizes connections on multivariate normal distributions.

problem Characterizing connections on statistical manifold of multivariate normal distributions.
method Analyzes statistical manifold (N,gF,ablaA,ablaA)(\mathcal{N}, g^F, abla^{A}, abla^{A*}) of multivariate normal distributions.
result The Amari-Chentsov connection ablaA abla^{A} is characterized by conjugate symmetry.

The paper evaluates biased methods for alpha-divergence minimization.

problem The impact of bias on solutions found for alpha-divergence minimization.
method Empirical evaluation of biased methods for alpha-divergence minimization, focusing on bias effects and dimensionality.
result Solutions are biased towards KL-divergence minimizers and require impractical computation in high dimensions to minimize alpha-divergence.

The paper presents methods to improve uncertainty calibration in Bayesian Neural Networks.

problem Uncalibrated Bayesian Neural Networks often lead to overconfidence.
method The paper uses alpha-divergences from Information Geometry for calibration.
result Calibration using alpha-divergences provides better uncertainty estimates and is more efficient.

We describe the underlying probabilistic interpretation of alpha and beta divergences. We first show that beta divergences are inherently tied to Tweedie distributions, a particular type of exponential family, known as exponential dispersion models. Starting from the variance function of a Tweedie model, we outline how…

2012-09-19abs ↗pdf ↗

New insights into Markov chain geometry via positive transition measures.

problem Lack of statistical meaning in the space of transition probabilities.
method Constructing an extension of the space of transition probabilities using Amari's theory of positive measures.
result Introduction of a new dually flat structure for the space of positive transition measures.

We investigate the use of alternative divergences to Kullback-Leibler (KL) in variational inference(VI), based on the Variational Dropout \cite{kingma2015}. Stochastic gradient variational Bayes (SGVB) \cite{aevb} is a general framework for estimating the evidence lower bound (ELBO) in Variational Bayes. In this work, …

2017-11-12abs ↗pdf ↗

On the predual of a von Neumann algebra, we define a differentiable manifold structure and affine connections by embeddings into non-commutative L_p-spaces. Using the geometry of uniformly convex Banach spaces and duality of the L_p and L_q spaces for 1/p+1/q=1, we show that we can introduce the α-divergence, for αin (…

2003-11-05abs ↗pdf ↗

This paper introduces the variational Rényi bound (VR) that extends traditional variational inference to Rényi's alpha-divergences. This new family of variational methods unifies a number of existing approaches, and enables a smooth interpolation from the evidence lower-bound to the log (marginal) likelihood that is co…

2016-02-06abs ↗pdf ↗

The paper explores how information geometry impacts classical CR inequalities.

problem Deriving and generalizing CR inequalities using information geometry.
method Examining Eguchi's theory and applying Amari-Nagoaka's theory to KL-divergence, and then extending to other divergences.
result Generalized CR inequalities derived from various divergences.

New α\alpha-divergence loss function improves neural density ratio estimation.

problem Optimization challenges in existing DRE methods, especially overfitting and high sample requirements.
method Derived α\alpha-divergence loss function (α\alpha-Div) for neural density ratio estimation.
result The α\alpha-divergence loss function (α\alpha-Div) offers stable and effective optimization for DRE.

We propose a novel interpretation of the collapsed variational Bayes inference with a zero-order Taylor expansion approximation, called CVB0 inference, for latent Dirichlet allocation (LDA). We clarify the properties of the CVB0 inference by using the alpha-divergence. We show that the CVB0 inference is composed of two…

2012-06-27abs ↗pdf ↗

The paper develops divergences for Gaussian processes and RKHS settings.

problem Estimating divergences in infinite-dimensional spaces.
method Formulations of Alpha Log-Det divergences, continuity in norm, laws of large numbers, consistent estimation from finite samples.
result Infinite-dimensional divergences can be estimated from finite-dimensional versions with dimension-independent sample complexities.

To obtain uncertainty estimates with real-world Bayesian deep learning models, practical inference approximations are needed. Dropout variational inference (VI) for example has been used for machine vision and medical applications, but VI can severely underestimates model uncertainty. Alpha-divergences are alternative …

2017-03-08abs ↗pdf ↗

A method to compute divergences between decomposable models, useful in supervised learning.

problem Computing exact divergences between high-dimensional distributions is intractable.
method Proposes an approach to compute exact alpha-beta divergences between marginal and conditional distributions of decomposable models.
result Tractable computation of marginal and conditional alpha-beta divergences.

This paper introduces a variational approximation framework using direct optimization of what is known as the {\it scale invariant Alpha-Beta divergence} (sAB divergence). This new objective encompasses most variational objectives that use the Kullback-Leibler, the R{é}nyi or the gamma divergences. It also gives access…

2018-05-02abs ↗pdf ↗

In Riemannian geometry geodesics are integral curves of the Riemannian distance gradient. We extend this classical result to the framework of Information Geometry. In particular, we prove that the rays of level-sets defined by a pseudo-distance are generated by the sum of two tangent vectors. By relying on these vector…

2018-06-29abs ↗pdf ↗

Paper formalizes and analyzes a new bound for variational inference.

problem Lack of theoretical guarantees in variational algorithms.
method Introduces VR-IWAE bound, a generalization of IWAE.
result VR-IWAE bound leads to unbiased gradient estimators.

Black box variational inference (BBVI) with reparameterization gradients triggered the exploration of divergence measures other than the Kullback-Leibler (KL) divergence, such as alpha divergences. In this paper, we view BBVI with generalized divergences as a form of estimating the marginal likelihood via biased import…

2017-09-21abs ↗pdf ↗

AES uses α-divergence to select informative points for BO, improving optimization performance.

problem Optimizing complex functions with limited evaluations.
method AES uses α-divergence to select points based on dependency with global maximum.
result AES outperforms other information-based acquisition functions in various experiments.

Study of generalized Csiszár divergences and their application to Cramér-Rao bounds.

problem Deriving lower bounds for estimator variance using generalized divergences.
method Applied Eguchi's theory to derive Fisher information metric and dual affine connections.
result More widely applicable Cramér-Rao inequality for escort distributions.

We consider the nonlinear Kalman filtering problem using Kullback-Leibler (KL) and αα-divergence measures as optimization criteria. Unlike linear Kalman filters, nonlinear Kalman filters do not have closed form Gaussian posteriors because of a lack of conjugacy due to the nonlinearity in the likelihood. In this paper …

2017-05-01abs ↗pdf ↗

The Conant-Ashby theorem is verified for hypergraph observers, leading to unique learning rules.

problem Verifying conditions for hypergraph observers to maintain internal models.
method Formalizing persistent observers, applying the Conant-Ashby theorem, and using natural gradient descent.
result Natural gradient descent is the unique admissible learning rule for hypergraph observers.

Study curvature and torsion in Gaussian distribution's dual coordinate system.

problem Characterize geometric invariants of Gaussian distribution.
method Investigate Riemannian curvature and torsion in a dual coordinate system of Gaussian distribution.
result Explicitly give Amari formulas in the new coordinate system.

Dual affine connections on Riemannian manifolds have played a central role in the field of information geometry since their introduction by Amari. Here I would like to extend the notion of dual connections to general vector bundles with an inner product, in the same way as a unitary connection generalizes a metric affi…

2015-11-24abs ↗pdf ↗

Optimized α\alpha-posteriors reduce KL divergence from true posterior in parametric misspecification.

problem Reduction of KL divergence from true posterior in parametric model misspecification.
method Derivation of Bernstein-von Mises theorem and optimization of α\alpha-posteriors.
result Optimized α\alpha-posteriors minimize KL divergence from true posterior, especially in severe misspecification.

We develop a family of infinite-dimensional (non-parametric) manifolds of probability measures. The latter are defined on underlying Banach spaces, and have densities of class CbkC_b^k with respect to appropriate reference measures. The case k=k=\infty, in which the manifolds are modelled on Fréchet spaces, is included.…

2016-08-13abs ↗pdf ↗

Black-box alpha (BB-αα) is a new approximate inference method based on the minimization of αα-divergences. BB-αα scales to large datasets because it can be implemented using stochastic gradient descent. BB-αα can be applied to complex probabilistic models with little effort since it only requires as input the likel…

2015-11-10abs ↗pdf ↗

A new Weyl prior is proposed for Bayesian statistics, offering a more canonical choice for parameter α.

problem Choosing a prior distribution for Bayesian inference.
method Proposed a new Weyl prior based on the Weyl structure on a statistical manifold.
result The Weyl prior is a special case of the α-parallel prior with α = -n, where n is the dimension of the statistical manifold.

EGAB algorithms improve online portfolio selection.

problem Online portfolio selection problem.
method Generalized exponentiated gradient (EG) updates with Alpha-Beta divergence regularization.
result EGAB algorithms enhance portfolio performance, especially with transaction costs.

Alpha2 discovers logical formulaic alphas using deep reinforcement learning.

problem Discovering interpretable formulaic alphas for better trading strategies.
method Formulating alpha discovery as program construction, using deep reinforcement learning to navigate the search space.
result Empirical experiments show Alpha2 identifies diverse, logical, and effective alphas improving trading strategy performance.

The problem of estimating an unknown discrete distribution from its samples is a fundamental tenet of statistical learning. Over the past decade, it attracted significant research effort and has been solved for a variety of divergence measures. Surprisingly, an equally important problem, estimating an unknown Markov ch…

2018-10-28abs ↗pdf ↗

We review basic notions in the field of information geometry such as Fisher metric on statistical manifold, αα-connection and corresponding curvature following Amari's work . We show application of information geometry to asymptotic statistical inference.

2014-10-09abs ↗pdf ↗