Rényi divergence is related to Rényi entropy much like Kullback-Leibler divergence is related to Shannon's entropy, and comes up in many settings. It was introduced by Rényi as a measure of information that satisfies almost the same axioms as Kullback-Leibler divergence, and depends on a parameter that is called its or…
New dispersion indices based on inaccuracy and divergence introduced for information measures.
problem Measuring variability in uncertainty measures.
method Introducing new dispersion indices based on Kerridge inaccuracy and Kullback-Leibler divergence.
result Properties, bounds, and examples of new dispersion indices presented.
Information-theoretic measures such as the entropy, cross-entropy and the Kullback-Leibler divergence between two mixture models is a core primitive in many signal processing tasks. Since the Kullback-Leibler divergence of mixtures provably does not admit a closed-form formula, it is in practice either estimated using …
This paper improves active learning by using robust divergences for committee disagreement.
problem Active learning with high measurement costs.
method Query by committee with Bregman divergence (including Kullback-Leibler divergence as a special case).
result The proposed method is more robust and performs as well as or better than conventional methods.
Paper studies regularized KKL divergence for distributions with disjoint supports.
problem Inability of original KKL divergence to handle distributions with disjoint supports.
method Proposes a regularized variant of KKL divergence, derives bounds, and provides closed-form expression.
result Regularized KKL divergence is well-defined for all distributions and has finite-sample bounds.
A new method optimizes a generalized Kullback-Leibler divergence for better simulation-based inference.
problem Optimizing likelihood functions when they are only known implicitly.
method Optimizes a generalized Kullback-Leibler divergence that accounts for normalization constants in unnormalized distributions.
result Unified approach that combines Neural Posterior Estimation and Neural Ratio Estimation.
Study compares statistical properties and power of divergence measures for credit risk monitoring.
problem Detecting distributional shifts in credit risk models.
method Derives statistical properties and chi-square benchmark values for Jensen-Shannon Divergence and Kullback-Leibler Divergence, demonstrating their applicability in credit risk monitoring.
result Jensen-Shannon Divergence and Kullback-Leibler Divergence follow chi-square distributions and reveal practical trade-offs in minimizing false positives vs. detecting changes.
We present a derivation of the Kullback Leibler (KL)-Divergence (also known as Relative Entropy) for the von Mises Fisher (VMF) Distribution in d d d -dimensions.
Paper calculates KL divergence for isotropic Gaussian-Markov fields.
problem Measuring divergence between isotropic Gaussian-Markov fields.
method Derives closed-form KL divergence expressions.
result Develops new similarity measures in image processing.
Proposes a guaranteed regularization method for maximum likelihood estimation using gauge symmetry in Kullback-Leibler divergence.
problem Overfitting in maximum likelihood estimation.
method Introduces a regularization approach based on gauge symmetry in Kullback-Leibler divergence.
result The method provides a theoretically guaranteed optimal model without frequent hyperparameter tuning.
New ONMF model minimizes KL divergence for better sparse data modeling.
problem Clustering and data modeling with sparse vectors.
method Developed KL-ONMF algorithm based on alternating optimization.
result KL-ONMF outperforms Frobenius-norm ONMF for document classification and hyperspectral image unmixing.
Machine learning classification limits estimated using Kullback-Leibler divergence and Cohen's Kappa.
problem Estimating the best possible performance of machine learning classification algorithms.
method Relating Kullback-Leibler divergence to Cohen's Kappa and using the Chernoff-Stein Lemma to estimate error rates.
result Classification algorithms could not have performed any better due to underlying probability density functions for the two classes.
The paper optimizes distribution estimation with high probability in Kullback-Leibler divergence.
problem Estimating discrete distributions with high probability in Kullback-Leibler divergence.
method Uses online learning techniques for novel estimator construction via online-to-batch conversion.
result Optimal rate of estimation is pinned down up to a doubly logarithmic factor of K.
This paper provides efficient algorithms for computing entropy and KL divergence in Bayesian networks.
problem Computing entropy and KL divergence for Bayesian networks efficiently.
method Leveraging the graphical structure of Bayesian networks, the paper provides computationally efficient algorithms.
result Reduces computational complexity of KL divergence from cubic to quadratic for Gaussian BNs.
The paper proposes a new method to approximate Wasserstein-Fisher-Rao flows using Monte Carlo techniques.
problem Sampling from probability distributions and minimizing Kullback-Leibler divergence.
method Sequential Monte Carlo approximations of Wasserstein-Fisher-Rao gradient flows.
result The proposed method outperforms other Monte Carlo algorithms in certain conditions.
CAKD framework optimizes knowledge transfer by focusing on influential components of distillation.
problem Balancing and optimizing knowledge transfer in distillation models.
method Decouple KL divergence into BCD, SCD, and WCD; prioritize influential components.
result CAKD framework consistently outperforms baseline across diverse models and datasets.
In this paper, we derive a useful lower bound for the Kullback-Leibler divergence (KL-divergence) based on the Hammersley-Chapman-Robbins bound (HCRB). The HCRB states that the variance of an estimator is bounded from below by the Chi-square divergence and the expectation value of the estimator. By using the relation b…
Proposes a new learning method for RBMs that combines strengths of forward and reverse KLD.
problem Underfitting and mode-collapse issues in RBM learning.
method Ratio divergence learning using target energy.
result Significantly outperforms other learning methods in energy function fitting, mode-covering, and stability.
The paper develops inequalities for log-concave functions and related surface areas.
problem Understanding log-concave functions and their inequalities.
method Establishing new inequalities through f-divergences and functional affine surface areas.
result New inequalities on functional affine surface area and bounds for Kullback-Leibler divergence.
Develops deep NMF models using β-divergences for feature extraction.
problem Inadequate evaluation metrics for deep NMF on diverse datasets.
method Introduces new deep NMF models using Kullback-Leibler divergence.
result Improves feature extraction quality across different types of data.
DAIS minimizes symmetrized KL divergence between initial and target distributions.
problem Optimizing over initial distributions in importance sampling.
method Differentiable annealed importance sampling (DAIS) minimizing symmetrized KL divergence.
result DAIS minimizes symmetrized KL divergence between initial and target distributions.
The Wasserstein probability metric has received much attention from the machine learning community. Unlike the Kullback-Leibler divergence, which strictly measures change in probability, the Wasserstein metric reflects the underlying geometry between outcomes. The value of being sensitive to this geometry has been demo…
Learning word representations has garnered greater attention in the recent past due to its diverse text applications. Word embeddings encapsulate the syntactic and semantic regularities of sentences. Modelling word embedding as multi-sense gaussian mixture distributions, will additionally capture uncertainty and polyse…
Study measures irreversibility in crypto trends using Kullback-Leibler divergence.
problem Assessing irreversibility in cryptocurrency trends.
method Defined irreversibility index using Kullback-Leibler divergence between uptrend and downtrend distributions.
result Strong irreversibility in all analyzed cryptocurrencies, with trends evolving over time.
Paper proves Jeffrey's update rule minimizes relative entropy.
problem Improving Bayesian learning algorithms.
method More concise proof of Jeffrey's update rule.
result Jeffrey's update rule reduces relative entropy.
Combines expert models using Kullback-Leibler divergence to create a combined model.
problem Combining expert views on stochastic processes.
method Minimizes weighted Kullback-Leibler divergence to create a barycentre model.
result Existence and uniqueness of the barycentre model with explicit representation.
A new method combines VI and IS to improve Bayesian inference accuracy.
problem Bayesian inference often underestimates posterior tails, leading to miscalibration and degeneracy.
method Proposes a novel combination of optimization and sampling techniques using the forward KL divergence.
result The method guarantees asymptotic consistency and fast convergence to optimal IS and variational approximations.
Paper relaxes triangle inequality for KL divergence between Gaussian distributions.
problem KL divergence does not satisfy triangle inequality for Gaussian distributions.
method Investigates relaxed triangle inequality and finds supremum.
result Supremum of KL divergence is found and conditions for attaining it are determined.
Paper connects rejection learning to Bhattacharyya divergence.
problem Learning models to abstain from predictions.
method Developed a link between rejection and thresholding different statistical divergences, focusing on Bhattacharyya divergence.
result Rejector obtained by joint ideal distribution corresponds to thresholding of skewed Bhattacharyya divergence.
ETM identifies field-specific keywords in text classification.
problem Unsupervised text classification with field-specific keywords.
method Weighted Lasso penalty and pairwise Kullback-Leibler divergence penalty for topic separation.
result ETM improves topic coherence by 22% and 10% compared to LDA.
In this paper, we introduce new classes of divergences by extending the definitions of the Bregman divergence and the skew Jensen divergence. These new divergence classes (g-Bregman divergence and skew g-Jensen divergence) satisfy some properties similar to the Bregman or skew Jensen divergence. We show these g-diverge…
Optimal transport with f f f -divergence regularization using generalized Sinkhorn algorithm.
problem Optimal transport with f f f -divergence regularization. method Generalized Sinkhorn algorithm for solving optimal transport problems with various f f f -divergences. result Strong duality holds, optimums are attained, and convergence to an optimal solution is guaranteed under certain conditions.
E 2 ^2 2 M optimizes tensor density estimation by relaxing α α α -divergence to KL-divergence.
problem Analytical challenges in traditional α α α -divergence optimization for tensor-based density estimation. method E 2 ^2 2 M algorithm: relaxes optimization to KL-divergence, then applies tensor many-body approximation. result Flexible modeling of various low-rank structures and their mixtures.
Study improves density estimation for compact domains using h h h -lifted KL divergence.
problem Estimating probability density functions on compact domains.
method Introduced h h h -lifted Kullback--Leibler (KL) divergence for risk minimization. result Proved O ( 1 / n ) \mathcal{O}(1/{\sqrt{n}}) O ( 1/ n ) bound on estimation error. New optimization method corrects data-driven optimizer's curse.
problem Over-optimistic evaluation in data-driven optimization.
method Smoothed f f f -Divergence Distributionally Robust Optimization (DRO). result Statistical bound on out-of-sample performance nearly tightest.
Model financial markets using information theory with a single parameter.
problem Capture the complexity of financial markets with a simple model.
method Derive an idealized model based on four information-theoretic assumptions, minimizing surprisal and divergence.
result The model uses squared radial Ornstein-Uhlenbeck processes for state variables and their sums.
The paper develops divergences for Gaussian processes and RKHS settings.
problem Estimating divergences in infinite-dimensional spaces.
method Formulations of Alpha Log-Det divergences, continuity in norm, laws of large numbers, consistent estimation from finite samples.
result Infinite-dimensional divergences can be estimated from finite-dimensional versions with dimension-independent sample complexities.
The Kelly Criterion is applied to prediction markets to analyze risk and return.
problem Mean beliefs in prediction markets often differ from actual prices.
method Logarithmic utility and Kullback-Leibler divergence are used to study risk and return adjustments.
result Misjudgment of bias and investment fraction affect portfolio growth rate.
We introduce a new approximation of f f f -divergences for machine learning.
problem Variational representations of f f f -divergences for machine learning. method Definition and analysis of Moreau-Yosida approximation of f f f -divergences with the Wasserstein-1 metric. result Generalization and relaxation of hard Lipschitz constraints in f f f -divergences. This paper provides performance guarantees for neural estimation of statistical distances.
problem Developing performance guarantees for neural estimation of statistical distances.
method Non-asymptotic error bounds using function approximation theorems and empirical process theory.
result Established a fundamental tradeoff between approximation and estimation errors in neural estimation of statistical distances.
Recently, a method called the Mutual Information Neural Estimator (MINE) that uses neural networks has been proposed to estimate mutual information and more generally the Kullback-Leibler (KL) divergence between two distributions. The method uses the Donsker-Varadhan representation to arrive at the estimate of the KL d…
New method for factor analysis using nuclear and ℓ 0 \ell_0 ℓ 0 norms.
problem Finding a low-rank plus sparse decomposition from noisy covariance matrix.
method Formulated an optimization problem with nuclear norm, ℓ 0 \ell_0 ℓ 0 norm, and KL divergence. Used alternating minimization algorithm. result Algorithm effectively decomposes covariance matrices in synthetic and real datasets.
Due to the success of the bag-of-word modeling paradigm, clustering histograms has become an important ingredient of modern information processing. Clustering histograms can be performed using the celebrated k k k -means centroid-based algorithm. From the viewpoint of applications, it is usually required to deal with symm…
We propose a robust estimator to improve maximum likelihood in probabilistic models.
problem Overfitting and sensitivity to noise in maximum likelihood estimation.
method Distributionally robust maximum likelihood estimator that minimizes worst-case expected log-loss.
result The robust estimator is statistically consistent and performs well in regression and classification tasks.
We review recent results about the maximal values of the Kullback-Leibler information divergence from statistical models defined by neural networks, including naive Bayes models, restricted Boltzmann machines, deep belief networks, and various classes of exponential families. We illustrate approaches to compute the max…
The study explores generalized divergences and exponential families with a focus on sufficient conditions and laws of large numbers.
problem Generalization of Kullback-Leibler divergence and exponential families.
method Investigation of ( h , τ ) (h,τ) ( h , τ ) -divergence and ( h , τ ) (h,τ) ( h , τ ) -exponential families, definition of ( h , τ ) (h,τ) ( h , τ ) -dependence, proof of law of large numbers. result Sufficient condition for ( h , τ ) (h,τ) ( h , τ ) -divergence to induce Hessian structure on ( h , τ ) (h,τ) ( h , τ ) -exponential family, proof of law of large numbers. Simulates risk-neutral markets using neural spline flows.
problem Creating realistic risk-neutral market simulations.
method Developed a low-dimensional martingale representation and used neural spline flows for sampling.
result The calibrated simulator is closest to historical data with respect to Kullback-Leibler divergence.
Paper improves variational inference on Boolean hypercube using quantum methods.
problem Improving variational inference for pairwise Markov random fields on the Boolean hypercube.
method Quantum relaxations of the Kullback-Leibler divergence for upper-bounds, primal-dual optimization, and greedy selection of hierarchies.
result Efficient algorithm and improved bounds for variational inference.