Regularizes f-divergences with MMD to analyze Wasserstein flows.
problem Limitations of f-divergences in measures' support. method Rewriting MMD regularization as Moreau envelope in RKHS, analyzing gradients.
result Analysis of Wasserstein flows of MMD-regularized f-divergences. In this report, we derive a non-negative series expansion for the Jensen-Shannon divergence (JSD) between two probability distributions. This series expansion is shown to be useful for numerical calculations of the JSD, when the probability distributions are nearly equal, and for which, consequently, small numerical er…
Paper presents ERM with f-divergence regularization and its properties.
problem Minimizing empirical risk with f-divergence constraints. method Introduces normalization function and solves ERM-fDR via ODE. result Characterizes difference between empirical risks and provides numerical algorithm.
New method improves imitation learning from expert observations.
problem Challenges in imitation learning from observation setting.
method Reparameterized Variational Divergence Minimization.
result Our method outperforms baseline approaches in low-dimensional tasks.
New divergences improve estimation and GAN training performance.
problem Improving estimation and training in machine learning models.
method Function-space regularized Rényi divergences.
result New divergences reduce variance and improve training performance.
Proposes a new learning method for RBMs that combines strengths of forward and reverse KLD.
problem Underfitting and mode-collapse issues in RBM learning.
method Ratio divergence learning using target energy.
result Significantly outperforms other learning methods in energy function fitting, mode-covering, and stability.
Estimating divergences in a consistent way is of great importance in many machine learning tasks. Although this is a fundamental problem in nonparametric statistics, to the best of our knowledge there has been no finite sample exponential inequality convergence bound derived for any divergence estimators. The main cont…
Improved hypothesis testing and change-point detection using diffusion-based methods.
problem Limited power of score-based hypothesis tests and change-point detection.
method Extending score-based Fisher divergence to diffusion-divergence by multiplying score functions with a matrix-valued function or weight matrix.
result Theoretical quantification and demonstration of optimal performance of diffusion-based algorithms.
In this paper, we derive a useful lower bound for the Kullback-Leibler divergence (KL-divergence) based on the Hammersley-Chapman-Robbins bound (HCRB). The HCRB states that the variance of an estimator is bounded from below by the Chi-square divergence and the expectation value of the estimator. By using the relation b…
New algorithm optimizes unimodal bandits using empirical divergence.
problem Optimizing decisions in multi-armed bandit problems with unimodal distributions.
method Indexed Minimum Empirical Divergence (IMED) adapted for unimodal structure.
result IMED-UB algorithm optimally exploits unimodal structure.
Dynamic Vocabulary Pruning stabilizes LLM training by removing low-probability tokens.
problem Training Large Language Models (LLMs) with Reinforcement Learning (RL) causes numerical divergence between inference and training.
method Dynamic Vocabulary Pruning (DVP) constrains the RL objective to a safe vocabulary that excludes low-probability tokens.
result DVP stabilizes training by reducing systematic bias introduced by the extreme tail of the token distribution.
New method tightens variational representations of divergences for faster learning.
problem Improving tightness of variational representations of divergences for faster statistical estimation.
method Improved objective functionals constructed via an auxiliary optimization problem, leveraging neural network approximation.
result Tighter variational representations can result in significantly faster learning and more accurate estimation of divergences.
New α-divergence loss function improves neural density ratio estimation.
problem Optimization challenges in existing DRE methods, especially overfitting and high sample requirements.
method Derived α-divergence loss function (α-Div) for neural density ratio estimation. result The α-divergence loss function (α-Div) offers stable and effective optimization for DRE. Study local invariants of divergence-free webs in geometry.
problem Characterize triviality of divergence-free webs.
method Introduce two local invariants: differential and geometric.
result Triviality of either invariant characterizes trivial divergence-free web-germs.
Paper explores how generative models can be made more creative.
problem Limitation of generative models in diverging from original data distribution.
method Proposes a novel training objective called Bounded Adversarial Divergence (BAD) to enable creative divergence.
result Preliminary results suggest BAD can enable creative divergence in generative models.
The paper introduces a new divergence for portfolio management to outperform a benchmark.
problem Maximizing expected utility of outperformance over a benchmark with constraints.
method Uses α-Bregman-Wasserstein divergence to penalize underperformance more than overperformance. result Proves existence and uniqueness of optimal portfolio strategy and conditions for constraints binding.
Optimal payoff choice constrained by Bregman-Wasserstein divergence.
problem Maximizing utility under a deviation constraint from a benchmark.
method Solving the problem using Bregman-Wasserstein divergence with a convex function φ.
result Provided the optimal payoff choice in this setting.
New tools quantify deep generative models' performance.
problem Measuring the quality-diversity trade-off in deep generative models.
method Established non-asymptotic bounds on sample complexity and introduced frontier integrals.
result Smoothed estimators improve convergence rates of divergence frontiers.
Optimal transport with f-divergence regularization using generalized Sinkhorn algorithm.
problem Optimal transport with f-divergence regularization. method Generalized Sinkhorn algorithm for solving optimal transport problems with various f-divergences. result Strong duality holds, optimums are attained, and convergence to an optimal solution is guaranteed under certain conditions.
In high-dimensional data, many sparse regression methods have been proposed. However, they may not be robust against outliers. Recently, the use of density power weight has been studied for robust parameter estimation and the corresponding divergences have been discussed. One of such divergences is the γ-divergence a…
New method uses SoS densities and α-divergences for efficient sequential transport maps.
problem Efficiently generating samples from approximated densities.
method Sequential transport maps using Sum-of-Squares (SoS) densities and α-divergences.
result Convex optimization problems with efficient semidefinite programming solutions.
Researchers establish bounds for SGMs' KL and Wasserstein divergences under various noise schedules.
problem Estimating the error between target and estimated distributions in SGMs.
method Established upper bounds for KL divergence and Wasserstein distance, incorporating target distribution properties and SGM hyperparameters.
result Optimal noise schedules identified for SGMs, improving generative quality.
A new gradient flow for MMD with closed-form implementation.
problem Existing gradient flows either lack tractable numerical implementation or require strong assumptions.
method Introduces a (de)-regularized Maximum Mean Discrepancy (DrMMD) and its gradient flow.
result Guarantees near-global convergence for a broad class of targets in both continuous and discrete time.
A density ratio is defined by the ratio of two probability densities. We study the inference problem of density ratios and apply a semi-parametric density-ratio estimator to the two-sample homogeneity test. In the proposed test procedure, the f-divergence between two probability densities is estimated using a density-r…
New Lie-group methods preserve geometric divergence-free features on manifolds.
problem Designing divergence-free Lie-group methods on manifolds.
method Introducing planar aromatic trees to span the free tracial post-Lie-Rinehart algebra.
result New Lie-group methods derived for high-order accuracy.
Reduced sample complexity for group-invariant distributions.
problem Improving sample complexity for estimating divergences of group-invariant distributions.
method Quantified reduction in sample complexity for Wasserstein-1 metric and Lipschitz-regularized α-divergences under finite and infinite groups.
result Sample complexity reduction proportional to group size for finite groups, and convergence rate depends on intrinsic dimension for infinite groups.
New variational formula for Rényi divergences improves neural network estimation in high dimensions.
problem Estimating Rényi divergences in high-dimensional systems.
method Derive and apply a variational formula for Rényi divergences over various function spaces.
result Neural network estimators of Rényi divergences are consistent under certain conditions.
Generative models tackle incompressible fluid flows by enforcing divergence-free constraints.
problem Simulating incompressible fluid flows with generative models.
method Score-based diffusion models with divergence-free constraint.
result Models can reproduce Kolmogorov turbulence characteristics.
The paper develops divergences for Gaussian processes and RKHS settings.
problem Estimating divergences in infinite-dimensional spaces.
method Formulations of Alpha Log-Det divergences, continuity in norm, laws of large numbers, consistent estimation from finite samples.
result Infinite-dimensional divergences can be estimated from finite-dimensional versions with dimension-independent sample complexities.
New loss functions based on f-divergences improve language model performance.
problem Improving multiclass classification and language modeling performance.
method Constructing new convex loss functions using f-divergences and deriving an operator for computation.
result The α-divergence loss function with α=1.5 performs well across various tasks. New method minimizes robust density power-based divergences for general parametric densities.
problem Computational complexity of minimizing DPD for general parametric densities.
method Stochastic approach to minimize DPD for general parametric density models.
result Proposed method can be applied to minimize other density power-based γ-divergences.
New findings show score matching's accuracy doesn't ensure numerical stability in diffusion sampling.
problem Numerical stability issues in diffusion sampling despite small forward-marginal error.
method Constructing a smooth score field with arbitrarily small forward-marginal L2 error, showing nonexplosive behavior and moments of every order. result Euler--Maruyama discretizations can converge in probability even when moments diverge, demonstrating failure of weak convergence.
Flow matching KL divergence bound derived for smooth distributions.
problem Estimating smooth distributions efficiently.
method Deterministic upper bound on KL divergence derived from flow-matching loss.
result Flow matching achieves nearly minimax-optimal efficiency under TV distance.
Paper analyzes sample complexity for offline f-divergence-regularized contextual bandits.
problem Lack of tight analyses for sample complexity in offline reinforcement learning.
method Novel pessimism-based analysis for reverse KL divergence, establishing ildeO(ε−1) sample complexity. result Achieves ildeO(ε−1) sample complexity for reverse KL divergence, surpassing existing bounds. Divergence estimators based on direct approximation of density-ratios without going through separate approximation of numerator and denominator densities have been successfully applied to machine learning tasks that involve distribution comparison such as outlier detection, transfer learning, and two-sample homogeneity…
Stochastic variational inference (SVI) plays a key role in Bayesian deep learning. Recently various divergences have been proposed to design the surrogate loss for variational inference. We present a simple upper bound of the evidence as the surrogate loss. This evidence upper bound (EUBO) equals to the log marginal li…
Positive definite matrices abound in a dazzling variety of applications. This ubiquity can be in part attributed to their rich geometric structure: positive definite matrices form a self-dual convex cone whose strict interior is a Riemannian manifold. The manifold view is endowed with a "natural" distance function whil…
This paper provides performance guarantees for neural estimation of statistical distances.
problem Developing performance guarantees for neural estimation of statistical distances.
method Non-asymptotic error bounds using function approximation theorems and empirical process theory.
result Established a fundamental tradeoff between approximation and estimation errors in neural estimation of statistical distances.
Paper shows robust generative learning with minimal assumptions on target distributions.
problem Learning generative models with minimal assumptions on target distributions.
method Lipschitz-regularized α-divergences with minimal assumptions. result Stable learning across various target distributions with minimal assumptions.
This paper provides efficient algorithms for computing entropy and KL divergence in Bayesian networks.
problem Computing entropy and KL divergence for Bayesian networks efficiently.
method Leveraging the graphical structure of Bayesian networks, the paper provides computationally efficient algorithms.
result Reduces computational complexity of KL divergence from cubic to quadratic for Gaussian BNs.
We propose a method to decrease the number of hidden units of the restricted Boltzmann machine while avoiding decrease of the performance measured by the Kullback-Leibler divergence. Then, we demonstrate our algorithm by using numerical simulations.
A new distribution addresses scalability and numerical stability issues of the vMF.
problem Scalability and numerical stability issues in sampling from the von Mises-Fisher (vMF) distribution.
method Proposes the Power Spherical distribution, retaining vMF's properties but addressing its drawbacks.
result Demonstrates the stability of Power Spherical distributions and applies it to a variational auto-encoder.
Generative adversarial network (GAN) is a minimax game between a generator mimicking the true model and a discriminator distinguishing the samples produced by the generator from the real training samples. Given an unconstrained discriminator able to approximate any function, this game reduces to finding the generative …
The paper introduces a new method for estimating optimal policies in dynamic treatment regimes using information geometry.
problem Estimating optimal policies in dynamic treatment regimes.
method Minimum information divergence method based on γ-power divergence. result The γ-power divergence method effectively seeks the optimal policy by vanishing the divergence between policy-equivalent Q-functions. Gradient flows of neural networks converge to optimal values or diverge, with thresholds and asymptotic behaviors.
problem Understanding the convergence and divergence of gradient flows in neural networks.
method Analysis of gradient flows on loss landscapes of neural networks using o-minimal structures.
result Gradient flows either converge to optimal values or diverge to infinity, with thresholds and asymptotic behaviors.
New f-Betas for portfolio optimization using f-divergence risk measures.
problem Optimizing portfolio performance under varying market conditions.
method Derive f-Betas and Hellinger-Betas, using f-divergence risk measures.
result Demonstrated new Beta metrics provide better performance under stress.
Enhanced 3D shape analysis using information geometry.
problem Challenges in comparing 3D point clouds due to their unstructured nature and complex geometry.
method Information geometric framework for 3D point cloud shape analysis using Gaussian Mixture Models (GMMs) on a statistical manifold. Proposed MSKL divergence with upper and lower bounds.
result MSKL provides stable and monotonically varying values that directly reflect geometric variation, outperforming traditional distances and existing KL approximations.
On non-Kähler manifolds the notion of harmonic maps is modified to that of Hermitian harmonic maps in order to be compatible with the complex structure. The resulting semilinear elliptic system is {\it not} in divergence form. The case of noncompact complete preimage and target manifolds is considered. We give conditio…