Rényi divergence is related to Rényi entropy much like Kullback-Leibler divergence is related to Shannon's entropy, and comes up in many settings. It was introduced by Rényi as a measure of information that satisfies almost the same axioms as Kullback-Leibler divergence, and depends on a parameter that is called its or…
A new method optimizes a generalized Kullback-Leibler divergence for better simulation-based inference.
problem Optimizing likelihood functions when they are only known implicitly.
method Optimizes a generalized Kullback-Leibler divergence that accounts for normalization constants in unnormalized distributions.
result Unified approach that combines Neural Posterior Estimation and Neural Ratio Estimation.
Information-theoretic measures such as the entropy, cross-entropy and the Kullback-Leibler divergence between two mixture models is a core primitive in many signal processing tasks. Since the Kullback-Leibler divergence of mixtures provably does not admit a closed-form formula, it is in practice either estimated using …
New dispersion indices based on inaccuracy and divergence introduced for information measures.
problem Measuring variability in uncertainty measures.
method Introducing new dispersion indices based on Kerridge inaccuracy and Kullback-Leibler divergence.
result Properties, bounds, and examples of new dispersion indices presented.
This paper improves active learning by using robust divergences for committee disagreement.
problem Active learning with high measurement costs.
method Query by committee with Bregman divergence (including Kullback-Leibler divergence as a special case).
result The proposed method is more robust and performs as well as or better than conventional methods.
Paper studies regularized KKL divergence for distributions with disjoint supports.
problem Inability of original KKL divergence to handle distributions with disjoint supports.
method Proposes a regularized variant of KKL divergence, derives bounds, and provides closed-form expression.
result Regularized KKL divergence is well-defined for all distributions and has finite-sample bounds.
Study compares statistical properties and power of divergence measures for credit risk monitoring.
problem Detecting distributional shifts in credit risk models.
method Derives statistical properties and chi-square benchmark values for Jensen-Shannon Divergence and Kullback-Leibler Divergence, demonstrating their applicability in credit risk monitoring.
result Jensen-Shannon Divergence and Kullback-Leibler Divergence follow chi-square distributions and reveal practical trade-offs in minimizing false positives vs. detecting changes.
Paper calculates KL divergence for isotropic Gaussian-Markov fields.
problem Measuring divergence between isotropic Gaussian-Markov fields.
method Derives closed-form KL divergence expressions.
result Develops new similarity measures in image processing.
We present a derivation of the Kullback Leibler (KL)-Divergence (also known as Relative Entropy) for the von Mises Fisher (VMF) Distribution in d d d -dimensions.
Proposes a guaranteed regularization method for maximum likelihood estimation using gauge symmetry in Kullback-Leibler divergence.
problem Overfitting in maximum likelihood estimation.
method Introduces a regularization approach based on gauge symmetry in Kullback-Leibler divergence.
result The method provides a theoretically guaranteed optimal model without frequent hyperparameter tuning.
New ONMF model minimizes KL divergence for better sparse data modeling.
problem Clustering and data modeling with sparse vectors.
method Developed KL-ONMF algorithm based on alternating optimization.
result KL-ONMF outperforms Frobenius-norm ONMF for document classification and hyperspectral image unmixing.
Optimal transport with f f f -divergence regularization using generalized Sinkhorn algorithm.
problem Optimal transport with f f f -divergence regularization. method Generalized Sinkhorn algorithm for solving optimal transport problems with various f f f -divergences. result Strong duality holds, optimums are attained, and convergence to an optimal solution is guaranteed under certain conditions.
The Wasserstein probability metric has received much attention from the machine learning community. Unlike the Kullback-Leibler divergence, which strictly measures change in probability, the Wasserstein metric reflects the underlying geometry between outcomes. The value of being sensitive to this geometry has been demo…
Machine learning classification limits estimated using Kullback-Leibler divergence and Cohen's Kappa.
problem Estimating the best possible performance of machine learning classification algorithms.
method Relating Kullback-Leibler divergence to Cohen's Kappa and using the Chernoff-Stein Lemma to estimate error rates.
result Classification algorithms could not have performed any better due to underlying probability density functions for the two classes.
In this paper, we introduce new classes of divergences by extending the definitions of the Bregman divergence and the skew Jensen divergence. These new divergence classes (g-Bregman divergence and skew g-Jensen divergence) satisfy some properties similar to the Bregman or skew Jensen divergence. We show these g-diverge…
The paper addresses instability in KL divergence estimation using a neural network discriminator.
problem Unstable estimation of KL divergence due to discriminator complexity.
method Using a Reproducing Kernel Hilbert Space (RKHS) to control discriminator complexity.
result Theoretical bound on error probability of KL estimates based on discriminator complexity in RKHS.
Paper relaxes triangle inequality for KL divergence between Gaussian distributions.
problem KL divergence does not satisfy triangle inequality for Gaussian distributions.
method Investigates relaxed triangle inequality and finds supremum.
result Supremum of KL divergence is found and conditions for attaining it are determined.
The paper optimizes distribution estimation with high probability in Kullback-Leibler divergence.
problem Estimating discrete distributions with high probability in Kullback-Leibler divergence.
method Uses online learning techniques for novel estimator construction via online-to-batch conversion.
result Optimal rate of estimation is pinned down up to a doubly logarithmic factor of K.
This paper provides efficient algorithms for computing entropy and KL divergence in Bayesian networks.
problem Computing entropy and KL divergence for Bayesian networks efficiently.
method Leveraging the graphical structure of Bayesian networks, the paper provides computationally efficient algorithms.
result Reduces computational complexity of KL divergence from cubic to quadratic for Gaussian BNs.
The paper proposes a new method to approximate Wasserstein-Fisher-Rao flows using Monte Carlo techniques.
problem Sampling from probability distributions and minimizing Kullback-Leibler divergence.
method Sequential Monte Carlo approximations of Wasserstein-Fisher-Rao gradient flows.
result The proposed method outperforms other Monte Carlo algorithms in certain conditions.
Paper connects rejection learning to Bhattacharyya divergence.
problem Learning models to abstain from predictions.
method Developed a link between rejection and thresholding different statistical divergences, focusing on Bhattacharyya divergence.
result Rejector obtained by joint ideal distribution corresponds to thresholding of skewed Bhattacharyya divergence.
CAKD framework optimizes knowledge transfer by focusing on influential components of distillation.
problem Balancing and optimizing knowledge transfer in distillation models.
method Decouple KL divergence into BCD, SCD, and WCD; prioritize influential components.
result CAKD framework consistently outperforms baseline across diverse models and datasets.
In this paper, we derive a useful lower bound for the Kullback-Leibler divergence (KL-divergence) based on the Hammersley-Chapman-Robbins bound (HCRB). The HCRB states that the variance of an estimator is bounded from below by the Chi-square divergence and the expectation value of the estimator. By using the relation b…
Proposes a new learning method for RBMs that combines strengths of forward and reverse KLD.
problem Underfitting and mode-collapse issues in RBM learning.
method Ratio divergence learning using target energy.
result Significantly outperforms other learning methods in energy function fitting, mode-covering, and stability.
New method uses KL divergence to detect out-of-distribution data effectively.
problem Flow-based models assign higher likelihoods to OOD data than ID data, making OOD detection challenging.
method Proposes a method leveraging KL divergence and local pixel dependence of representations for anomaly detection.
result Demonstrates effectiveness and robustness on prevalent benchmarks.
The paper develops inequalities for log-concave functions and related surface areas.
problem Understanding log-concave functions and their inequalities.
method Establishing new inequalities through f-divergences and functional affine surface areas.
result New inequalities on functional affine surface area and bounds for Kullback-Leibler divergence.
The Kelly Criterion is applied to prediction markets to analyze risk and return.
problem Mean beliefs in prediction markets often differ from actual prices.
method Logarithmic utility and Kullback-Leibler divergence are used to study risk and return adjustments.
result Misjudgment of bias and investment fraction affect portfolio growth rate.
The study explores generalized divergences and exponential families with a focus on sufficient conditions and laws of large numbers.
problem Generalization of Kullback-Leibler divergence and exponential families.
method Investigation of ( h , τ ) (h,τ) ( h , τ ) -divergence and ( h , τ ) (h,τ) ( h , τ ) -exponential families, definition of ( h , τ ) (h,τ) ( h , τ ) -dependence, proof of law of large numbers. result Sufficient condition for ( h , τ ) (h,τ) ( h , τ ) -divergence to induce Hessian structure on ( h , τ ) (h,τ) ( h , τ ) -exponential family, proof of law of large numbers. Develops deep NMF models using β-divergences for feature extraction.
problem Inadequate evaluation metrics for deep NMF on diverse datasets.
method Introduces new deep NMF models using Kullback-Leibler divergence.
result Improves feature extraction quality across different types of data.
Study improves density estimation for compact domains using h h h -lifted KL divergence.
problem Estimating probability density functions on compact domains.
method Introduced h h h -lifted Kullback--Leibler (KL) divergence for risk minimization. result Proved O ( 1 / n ) \mathcal{O}(1/{\sqrt{n}}) O ( 1/ n ) bound on estimation error. New optimization method corrects data-driven optimizer's curse.
problem Over-optimistic evaluation in data-driven optimization.
method Smoothed f f f -Divergence Distributionally Robust Optimization (DRO). result Statistical bound on out-of-sample performance nearly tightest.
E 2 ^2 2 M optimizes tensor density estimation by relaxing α α α -divergence to KL-divergence.
problem Analytical challenges in traditional α α α -divergence optimization for tensor-based density estimation. method E 2 ^2 2 M algorithm: relaxes optimization to KL-divergence, then applies tensor many-body approximation. result Flexible modeling of various low-rank structures and their mixtures.
DAIS minimizes symmetrized KL divergence between initial and target distributions.
problem Optimizing over initial distributions in importance sampling.
method Differentiable annealed importance sampling (DAIS) minimizing symmetrized KL divergence.
result DAIS minimizes symmetrized KL divergence between initial and target distributions.
Learning word representations has garnered greater attention in the recent past due to its diverse text applications. Word embeddings encapsulate the syntactic and semantic regularities of sentences. Modelling word embedding as multi-sense gaussian mixture distributions, will additionally capture uncertainty and polyse…
Generative algorithms learn high-dimensional data efficiently and generate new samples.
problem Learning from scarce high-dimensional data.
method Lipschitz-regularized gradient flows and particle-based algorithms.
result Correctly transports gene expression data points with high dimensionality.
We propose a robust estimator to improve maximum likelihood in probabilistic models.
problem Overfitting and sensitivity to noise in maximum likelihood estimation.
method Distributionally robust maximum likelihood estimator that minimizes worst-case expected log-loss.
result The robust estimator is statistically consistent and performs well in regression and classification tasks.
We introduce a new approximation of f f f -divergences for machine learning.
problem Variational representations of f f f -divergences for machine learning. method Definition and analysis of Moreau-Yosida approximation of f f f -divergences with the Wasserstein-1 metric. result Generalization and relaxation of hard Lipschitz constraints in f f f -divergences. Study measures irreversibility in crypto trends using Kullback-Leibler divergence.
problem Assessing irreversibility in cryptocurrency trends.
method Defined irreversibility index using Kullback-Leibler divergence between uptrend and downtrend distributions.
result Strong irreversibility in all analyzed cryptocurrencies, with trends evolving over time.
Paper proves Jeffrey's update rule minimizes relative entropy.
problem Improving Bayesian learning algorithms.
method More concise proof of Jeffrey's update rule.
result Jeffrey's update rule reduces relative entropy.
Combines expert models using Kullback-Leibler divergence to create a combined model.
problem Combining expert views on stochastic processes.
method Minimizes weighted Kullback-Leibler divergence to create a barycentre model.
result Existence and uniqueness of the barycentre model with explicit representation.
A new method combines VI and IS to improve Bayesian inference accuracy.
problem Bayesian inference often underestimates posterior tails, leading to miscalibration and degeneracy.
method Proposes a novel combination of optimization and sampling techniques using the forward KL divergence.
result The method guarantees asymptotic consistency and fast convergence to optimal IS and variational approximations.
Recently, a method called the Mutual Information Neural Estimator (MINE) that uses neural networks has been proposed to estimate mutual information and more generally the Kullback-Leibler (KL) divergence between two distributions. The method uses the Donsker-Varadhan representation to arrive at the estimate of the KL d…
Optimizes diffusion processes for target distributions.
problem Efficiently generating target distributions from point masses.
method Stochastic interpolant framework with conditional expectation drift.
result Optimal diffusion coefficient minimizes path-space KL divergence.
The paper explores how information geometry impacts classical CR inequalities.
problem Deriving and generalizing CR inequalities using information geometry.
method Examining Eguchi's theory and applying Amari-Nagoaka's theory to KL-divergence, and then extending to other divergences.
result Generalized CR inequalities derived from various divergences.
ETM identifies field-specific keywords in text classification.
problem Unsupervised text classification with field-specific keywords.
method Weighted Lasso penalty and pairwise Kullback-Leibler divergence penalty for topic separation.
result ETM improves topic coherence by 22% and 10% compared to LDA.
Generative models often misrepresent class frequencies; this paper calibrates them.
problem Miscalibration of class frequencies in generative models.
method Formulated as constrained optimization, using surrogate objectives to approximate constraints.
result Significant reduction in calibration error across various models and applications.
New guarantees for VI in symmetric cases, extending previous results.
problem Symmetry in variational inference for complex distributions.
method Analysis of f f f -divergences and their stationary points under symmetry. result Symmetry-matching principles ensure recovery of mean and correlation matrix.
Method uses normalizing flows to efficiently sample from complex target densities.
problem Sampling from complex target densities with zero values in regions of transformation.
method Normalizing flows to address exploding reverse Kullback-Leibler divergence.
result Demonstrated efficient sampling from multi-mode complex density function.