f-divergences are a general class of divergences between probability measures which include as special cases many commonly used divergences in probability, mathematical statistics and information theory such as Kullback-Leibler divergence, chi-squared divergence, squared Hellinger distance, total variation distance e…
Innovative inequalities for divergences with applications in PAC-Bayesian bounds and Monte Carlo.
problem Developing new inequalities for divergences.
method Introducing novel change of measure inequalities for f-divergences and α-divergences. result Applications in PAC-Bayesian bounds and Monte Carlo estimates.
New PAC-Bayes bounds derived using Legendre transform and f-divergences.
problem Deriving PAC-Bayes bounds under various assumptions.
method Combining Legendre transform and Fenchel--Young inequality to derive change-of-measure inequalities.
result Extended PAC-Bayesian guarantees under tailored assumptions.
New bounds for Neyman-Pearson region using f-divergences.
problem Bounding the Neyman-Pearson region for hypothesis testing.
method Establishing novel lower and upper bounds using f-divergences. result Best possible lower bound for the Neyman-Pearson boundary using hockey-stick f-divergences. In this paper, we introduce new classes of divergences by extending the definitions of the Bregman divergence and the skew Jensen divergence. These new divergence classes (g-Bregman divergence and skew g-Jensen divergence) satisfy some properties similar to the Bregman or skew Jensen divergence. We show these g-diverge…
We unify f-divergences, Bregman divergences, surrogate loss bounds (regret bounds), proper scoring rules, matching losses, cost curves, ROC-curves and information. We do this by systematically studying integral and variational representations of these objects and in so doing identify their primitives which all are rela…
The paper develops inequalities for log-concave functions and related surface areas.
problem Understanding log-concave functions and their inequalities.
method Establishing new inequalities through f-divergences and functional affine surface areas.
result New inequalities on functional affine surface area and bounds for Kullback-Leibler divergence.
We investigate the framework of privacy amplification by iteration, recently proposed by Feldman et al., from an information-theoretic lens. We demonstrate that differential privacy guarantees of iterative mappings can be determined by a direct application of contraction coefficients derived from strong data processing…
Study on Lp affine surface areas and their inequalities for convex bodies.
problem Understanding weighted Lp affine surface areas in convex bodies. method Investigating valuations, isoperimetric inequalities, and connections to f divergences. result Established isoperimetric inequalities for weighted Lp affine surface areas. New bounds for estimating partition functions under bounded f-divergence.
problem Estimating partition functions with limited sample access.
method Information-theoretic characterization using integrated coverage profile and f-divergences. result Sharp phase transitions in sample complexity under f-divergences. Technical report on f-divergences and f-GAN training properties.
problem Understanding and optimizing f-divergences for GAN training.
method Elementary derivation and detailed expressions of f-divergences and their variational lower bounds.
result Informative properties of f-divergences and f-GAN training, including gradient matching and stability improvements.
Regularizes f-divergences with MMD to analyze Wasserstein flows.
problem Limitations of f-divergences in measures' support. method Rewriting MMD regularization as Moreau envelope in RKHS, analyzing gradients.
result Analysis of Wasserstein flows of MMD-regularized f-divergences. ERM with f-divergence regularization yields unique solution.
problem Optimizing empirical risk with f-divergence. method Mild conditions on f lead to unique optimal measure. result Equivalence of ERM-fDR to different f-divergence regularization. Paper introduces f-divergence variational inference for broader application.
problem Variational inference limited to specific divergences.
method Generalizes variational inference to all f-divergences using f-divergence minimization.
result Unified framework for variational inference with arbitrary f-divergences.
New f-divergence measures improve robustness in noisy label learning.
problem Improving robustness in learning with noisy labels.
method Derived decoupling property of f-divergence measures under label noise. result Properly defined f-divergence measures are robust with label noise. Proposes practical kernel tests for f-divergences with theoretical guarantees.
problem Two-sample testing and machine unlearning evaluation.
method Regularized f-divergence kernel tests, adaptive to hyperparameters. result Different f-divergences highlight localized differences. Paper presents ERM with f-divergence regularization and its properties.
problem Minimizing empirical risk with f-divergence constraints. method Introduces normalization function and solves ERM-fDR via ODE. result Characterizes difference between empirical risks and provides numerical algorithm.
Paper proposes f-EBM for training deep EBMs using various f-divergences.
problem Training deep EBMs with intractable partition functions.
method Introduces f-EBM framework and optimization algorithm for any f-divergence.
result f-EBM outperforms contrastive divergence and other f-divergences.
Rank-statistic method approximates f-divergences without density-ratio estimation.
problem Approximating f-divergences without explicit density-ratio estimation. method Mapping distribution rank histograms to discrete f-divergence and averaging over random projections. result The rank-statistic estimator is a lower bound of the true f-divergence and converges under mild conditions. New loss functions based on f-divergences improve language model performance.
problem Improving multiclass classification and language modeling performance.
method Constructing new convex loss functions using f-divergences and deriving an operator for computation.
result The α-divergence loss function with α=1.5 performs well across various tasks. Optimal transport with f-divergence regularization using generalized Sinkhorn algorithm.
problem Optimal transport with f-divergence regularization. method Generalized Sinkhorn algorithm for solving optimal transport problems with various f-divergences. result Strong duality holds, optimums are attained, and convergence to an optimal solution is guaranteed under certain conditions.
We introduce a new approximation of f-divergences for machine learning.
problem Variational representations of f-divergences for machine learning. method Definition and analysis of Moreau-Yosida approximation of f-divergences with the Wasserstein-1 metric. result Generalization and relaxation of hard Lipschitz constraints in f-divergences. The paper analyzes the statistical properties of GANs using f-divergence.
problem Understanding the statistical behavior of GANs and comparing different f-divergences. method Asymptotic analysis of f-divergence GANs, including Kullback-Leibler divergence. result Asymptotically equivalent GANs with the same discriminator classes for correctly specified models.
This work develops a unified framework for RLHF with general f-divergence regularization.
problem Theoretical understanding of general f-divergence regularization in RLHF. method Holistic approach across f-divergence class, two algorithms based on distinct sampling principles. result Provably efficient algorithms with O(logT) regret and O(1/T) sub-optimality gap. Paper analyzes sample complexity for offline f-divergence-regularized contextual bandits.
problem Lack of tight analyses for sample complexity in offline reinforcement learning.
method Novel pessimism-based analysis for reverse KL divergence, establishing ildeO(ε−1) sample complexity. result Achieves ildeO(ε−1) sample complexity for reverse KL divergence, surpassing existing bounds. Probabilistic models are often trained by maximum likelihood, which corresponds to minimizing a specific f-divergence between the model and data distribution. In light of recent successes in training Generative Adversarial Networks, alternative non-likelihood training criteria have been proposed. Whilst not necessarily…
The paper improves semi-supervised learning using f-divergences and α-Rényi divergences.
problem Improving semi-supervised learning with noisy pseudo-labels.
method Inspired by f-divergences and α-Rényi divergences, the paper develops new empirical risk functions and regularization techniques. result The new methods show better performance than traditional self-training methods, especially in noisy pseudo-label scenarios.
The paper analyzes the reward improvement of aligned policies in large language models.
problem Optimizing policies in large language models while staying close to a reference policy.
method Information-theoretic analysis and reduction to exponential order statistics.
result Information-theoretic upper bounds on reward improvement are derived.
A density ratio is defined by the ratio of two probability densities. We study the inference problem of density ratios and apply a semi-parametric density-ratio estimator to the two-sample homogeneity test. In the proposed test procedure, the f-divergence between two probability densities is estimated using a density-r…
Formula derived for sample complexity in binary hypothesis testing.
problem Determine the minimum number of samples to distinguish between two distributions.
method Developed a formula for sample complexity in both prior-free and Bayesian settings, using Jensen-Shannon and Hellinger divergences.
result Formula characterizes sample complexity for a wide range of error parameters, up to multiplicative constants.
We show that the variational representations for f-divergences currently used in the literature can be tightened. This has implications to a number of methods recently proposed based on this representation. As an example application we use our tighter representation to derive a general f-divergence estimator based on t…
Unified framework for generative models incorporating VAE and GAN.
problem Flexible incorporation of diverse measures of probability distance in generative models.
method Unified f-divergence generative model (f-GM) that incorporates both VAE and f-GAN.
result Unified f-GM enables flexible design of f-divergence functions without changing network structure.
New f-Betas for portfolio optimization using f-divergence risk measures.
problem Optimizing portfolio performance under varying market conditions.
method Derive f-Betas and Hellinger-Betas, using f-divergence risk measures.
result Demonstrated new Beta metrics provide better performance under stress.
New theory explains GAN's high quality but low diversity.
problem Lack of theoretical justification for non-saturating GAN training.
method Showed non-saturating GAN training approximately minimizes a specific f-divergence.
result Non-saturating GAN training minimizes a particular f-divergence.
Improved UDA framework using f-divergence measures.
problem Addressing distribution shifts in machine learning.
method Refined f-divergence-based discrepancy and f-domain discrepancy. result Novel target error and sample complexity bounds.
The paper improves model robustness by regularizing posterior differences.
problem Improving model robustness in noisy input scenarios.
method Posterior differential regularization with f-divergence. result Regularizing with f-divergence improves model robustness. Unified framework for deriving generalization bounds in supervised learning.
problem Generalization error bounds in supervised learning.
method Data Processing Inequality PAC-Bayesian framework.
result Unified bounds on binary Kullback-Leibler generalization gap for various divergences.
The t-distributed Stochastic Neighbor Embedding (t-SNE) is a powerful and popular method for visualizing high-dimensional data. It minimizes the Kullback-Leibler (KL) divergence between the original and embedded data distributions. In this work, we propose extending this method to other f-divergences. We analytically a…
New method estimates velocity fields for minimizing f-divergences without overfitting.
problem Minimizing statistical discrepancies between target and particle distributions.
method Directly estimate velocity fields using interpolation techniques, proving consistency under mild conditions.
result Consistent estimators of velocity fields improve accuracy in applications like domain adaptation and missing data imputation.
New method improves imitation learning from expert observations.
problem Challenges in imitation learning from observation setting.
method Reparameterized Variational Divergence Minimization.
result Our method outperforms baseline approaches in low-dimensional tasks.
We study exponential Levy models with change-point which is a random variable, independent from initial Levy processes. On canonical space with initially enlarged filtration we describe all equivalent martingale measures for change-point model and we give the conditions for the existence of f-divergence minimal equival…
New dual formulation reduces generalization error for ERM-fDR.
problem Generalization error in constrained optimization problems.
method Introduces a dual formulation of ERM-fDR using Legendre-Fenchel transform and implicit function theorem.
result Explicit characterizations of generalization error for algorithms under mild conditions.
New method for bidirectional generative modeling using adversarial gradient estimation.
problem Bidirectional generative modeling with various f-divergences. method Adversarial gradient estimation for f-divergence optimization. result Similar algorithms for different f-divergences with varying scaling. Dual optimization connects ERM-fDR to normalization function.
problem Empirical risk minimization with f-divergence regularization.
method Dual formulation, Legendre-Fenchel transform, implicit function theorem, nonlinear ODE.
result Computational method to calculate normalization function efficiently.
New optimization method corrects data-driven optimizer's curse.
problem Over-optimistic evaluation in data-driven optimization.
method Smoothed f-Divergence Distributionally Robust Optimization (DRO). result Statistical bound on out-of-sample performance nearly tightest.
We study the geometry of probability distributions with respect to a generalized family of Csiszár f-divergences. A member of this family is the relative α-entropy which is also a Rényi analog of relative entropy in information theory and known as logarithmic or projective power divergence in statistics. We apply E…
The problem of f-divergence estimation is important in the fields of machine learning, information theory, and statistics. While several nonparametric divergence estimators exist, relatively few have known convergence properties. In particular, even for those estimators whose MSE convergence rates are known, the asympt…
New algorithms improve robust estimation in contaminated Gaussian models.
problem Simultaneous estimation of location and variance matrix in contaminated Gaussian models.
method Tractable adversarial algorithms with spline discriminators for robust estimation.
result Achieve minimax optimal rates or near-optimal rates under Huber's contamination model.