New divergence measures improve KL approximation.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In high-dimensional data, many sparse regression methods have been proposed. However, they may not be robust against outliers. Recently, the use of density power weight has been studied for robust parameter estimation and the corresponding divergences have been discussed. One of such divergences is the -divergence a…
Study on geometric Jensen-Shannon divergence for Gaussian measures in Hilbert space.
Rényi divergence is related to Rényi entropy much like Kullback-Leibler divergence is related to Shannon's entropy, and comes up in many settings. It was introduced by Rényi as a measure of information that satisfies almost the same axioms as Kullback-Leibler divergence, and depends on a parameter that is called its or…
Proposes practical kernel tests for -divergences with theoretical guarantees.
Improved Bayesian inference using power priors with historical data.
Distances are fundamental primitives whose choice significantly impacts the performances of algorithms in machine learning and signal processing. However selecting the most appropriate distance for a given task is an endeavor. Instead of testing one by one the entries of an ever-expanding dictionary of {\em ad hoc} dis…
We investigate the use of alternative divergences to Kullback-Leibler (KL) in variational inference(VI), based on the Variational Dropout \cite{kingma2015}. Stochastic gradient variational Bayes (SGVB) \cite{aevb} is a general framework for estimating the evidence lower bound (ELBO) in Variational Bayes. In this work, …
Proposes a new measure to evaluate stability of statistical parameters under distributional shifts.
Develops a direct debiased machine learning framework using Bregman divergence.
This paper introduces a novel approach for learning to rank (LETOR) based on the notion of monotone retargeting. It involves minimizing a divergence between all monotonic increasing transformations of the training scores and a parameterized prediction function. The minimization is both over the transformations as well …
Automates VI divergence selection for efficient few-shot learning.
This paper introduces a variational approximation framework using direct optimization of what is known as the {\it scale invariant Alpha-Beta divergence} (sAB divergence). This new objective encompasses most variational objectives that use the Kullback-Leibler, the R{é}nyi or the gamma divergences. It also gives access…
We present the very first robust Bayesian Online Changepoint Detection algorithm through General Bayesian Inference (GBI) with -divergences. The resulting inference procedure is doubly robust for both the parameter and the changepoint (CP) posterior, with linear time and constant space complexity. We provide a const…
AES uses α-divergence to select informative points for BO, improving optimization performance.
Rank-statistic method approximates -divergences without density-ratio estimation.
Matrix SMD converges to unique solution minimizing Bregman divergence.
This paper addresses the estimation of the latent dimensionality in nonnegative matrix factorization (NMF) with the β-divergence. The β-divergence is a family of cost functions that includes the squared Euclidean distance, Kullback-Leibler and Itakura-Saito divergences as special cases. Learning the model order is impo…
New method uses KL-divergence to create non-informative priors for multivariate Gaussian.
A novel stepwise VI method using vine copulas for complex latent dependence.
BaM improves BBVI by optimizing a score-based divergence, leading to faster convergence.
We focus on the maximum regularization parameter for anisotropic total-variation denoising. It corresponds to the minimum value of the regularization parameter above which the solution remains constant. While this value is well know for the Lasso, such a critical value has not been investigated in details for the total…
Paper calculates KL divergence for isotropic Gaussian-Markov fields.
We develop a method to combine Markov chain Monte Carlo (MCMC) and variational inference (VI), leveraging the advantages of both inference approaches. Specifically, we improve the variational distribution by running a few MCMC steps. To make inference tractable, we introduce the variational contrastive divergence (VCD)…
We investigate the framework of privacy amplification by iteration, recently proposed by Feldman et al., from an information-theoretic lens. We demonstrate that differential privacy guarantees of iterative mappings can be determined by a direct application of contraction coefficients derived from strong data processing…
This paper proposes a more efficient training method for energy-based models.
EM optimizes tensor density estimation by relaxing -divergence to KL-divergence.
To ensure stability of learning, state-of-the-art generalized policy iteration algorithms augment the policy improvement step with a trust region constraint bounding the information loss. The size of the trust region is commonly determined by the Kullback-Leibler (KL) divergence, which not only captures the notion of d…
AIS algorithm improves heavy-tailed distribution estimation.
There has been a growing interest in mutual information measures due to their wide range of applications in Machine Learning and Computer Vision. In this paper, we present a generalized structured regression framework based on Shama-Mittal divergence, a relative entropy measure, which is introduced to the Machine Learn…
Model financial markets using information theory with a single parameter.
Local mass perspective on Bayesian inference
Unified technique for sequential estimation of convex divergences.
Bayesian surrogate models reduce uncertainty in high-dimensional design optimisation problems.
The paper presents methods to improve uncertainty calibration in Bayesian Neural Networks.
Neural networks estimate statistical divergences with performance guarantees.
A study on -GANs proving convergence and estimation guarantees.
We study -divergence contraction and its privacy implications.
Study compares chi-squared divergence and KL-divergence posteriors for PAC-Bayesian bounds.
The empirical NTK diverges from the NTK in classification problems during overtraining.
Improved UDA framework using -divergence measures.
Generative adversarial networks (GANs) are a family of generative models that do not minimize a single training criterion. Unlike other generative models, the data distribution is learned via a game between a generator (the generative model) and a discriminator (a teacher providing training signal) that each minimize t…
We study 1-parameter families in the space of -invariant, unit volume metrics on a given compact, connected, almost-effective homogeneous space . In particular, we focus on diverging sequences, i.e. which are not contained in any compact subset of , and we prove some structu…
Stochastic variational inference (SVI) plays a key role in Bayesian deep learning. Recently various divergences have been proposed to design the surrogate loss for variational inference. We present a simple upper bound of the evidence as the surrogate loss. This evidence upper bound (EUBO) equals to the log marginal li…
Deep nonlinear models pose a challenge for fitting parameters due to lack of knowledge of the hidden layer and the potentially non-affine relation of the initial and observed layers. In the present work we investigate the use of information theoretic measures such as mutual information and Kullback-Leibler (KL) diverge…
Gaussian mixture models (GMM) are powerful parametric tools with many applications in machine learning and computer vision. Expectation maximization (EM) is the most popular algorithm for estimating the GMM parameters. However, EM guarantees only convergence to a stationary point of the log-likelihood function, which c…
Differentially private statistical inference using -divergence.
Unified framework for debiased machine learning using Riesz representer and Bregman divergence.