This work generalizes Log-Determinant divergences to infinite-dimensional settings.
problem Generalizing Log-Determinant divergences to infinite-dimensional spaces.
method Introducing a parametrized family of divergences, Alpha-Beta Log-Determinant divergences, for positive definite unitized trace class operators.
result The Alpha-Beta Log-Det divergences encompass various divergences and metrics, including the affine-invariant Riemannian distance and symmetric Stein divergence.
Study on geometric Jensen-Shannon divergence for Gaussian measures in Hilbert space.
problem Computing divergence between Gaussian measures in infinite-dimensional Hilbert space.
method Closed form expression and regularization for divergence calculation.
result Closed form expression and regularization for Geometric Jensen-Shannon divergence.
Optimizes machine learning by approximating log determinants efficiently.
problem Computational expense in calculating log determinants for large data sets.
method Demonstrates the optimality of Maximum Entropy methods in approximating log determinants.
result Reduction of KL divergence between proposal and true eigenvalue distribution by adding more moments.
The paper develops divergences for Gaussian processes and RKHS settings.
problem Estimating divergences in infinite-dimensional spaces.
method Formulations of Alpha Log-Det divergences, continuity in norm, laws of large numbers, consistent estimation from finite samples.
result Infinite-dimensional divergences can be estimated from finite-dimensional versions with dimension-independent sample complexities.
Estimates Markov chains from samples, solving two related prediction and estimation problems.
problem Estimating an unknown Markov chain from its samples.
method Considered two problems: predicting conditional distribution and estimating transition matrix, using KL-divergence and various f f f -divergences. result Resolved estimation problem for all sufficiently smooth f f f -divergences, including KL-, L 2 L_2 L 2 , Chi-squared, Hellinger, and Alpha-divergences. The paper improves support recovery in high-dimensional precision matrix estimation using meta learning.
problem Support recovery in high-dimensional precision matrix estimation with reduced sample complexity.
method Pooling samples from different tasks and using an improper ℓ 1 \ell_1 ℓ 1 -regularized log-determinant Bregman divergence to estimate a single precision matrix. result The support of the improperly estimated single precision matrix is equal to the true support union with high probability.
Unified framework for data-free sampling using Wasserstein gradient flows.
problem Efficient sampling from unnormalized distributions without data.
method Unified theoretical framework based on Wasserstein gradient flows.
result Unified form of velocity field for various f-divergences.
An algorithm learns a kernel matrix from relative-distance constraints for semi-supervised clustering.
problem Learning metrics from relative-distance constraints to capture finer structures.
method Log determinant divergence for kernel matrix learning with relative-distance constraints.
result Kernels learned from relative-distance constraints yield better clusterings than existing methods.
New method estimates log-determinant using trace powers, avoiding classical limitations.
problem Estimating log-determinant of large matrices efficiently and accurately.
method Interpolating moment-generating function and its derivative at zero using trace powers.
result No continuous estimator using finite moments can be uniformly accurate over unbounded conditioning.
VI struggles to fully quantify uncertainty when distributions don't factorize.
problem Uncertainty quantification in non-factorizable distributions.
method Analysis of variational inference trade-offs and divergence choices.
result Different divergences yield different measures of uncertainty in VI.
The paper develops inequalities for log-concave functions and related surface areas.
problem Understanding log-concave functions and their inequalities.
method Establishing new inequalities through f-divergences and functional affine surface areas.
result New inequalities on functional affine surface area and bounds for Kullback-Leibler divergence.
RHMC accelerates sampling from log-concave distributions.
problem Sampling from log-concave probability distributions efficiently.
method RHMC uses simulated Hamiltonian dynamics with random integration times.
result RHMC converges exponentially fast in KL divergence for log-concave distributions.
Improved GANs estimate convergence rate for density estimation.
problem Improving the accuracy of density estimation with GANs.
method Proved an oracle inequality for JS divergence between GAN estimate and true density.
result JS-divergence rate of convergence is ( log n / n ) 2 β / ( 2 β + d ) (\log{n}/n)^{2β/(2β+ d)} ( log n / n ) 2 β / ( 2 β + d ) . A new VIS approach improves log-likelihood estimation in latent variable models.
problem Challenges in achieving high log-likelihood with VI for complex posterior distributions.
method Uses forward χ 2 χ^2 χ 2 divergence to optimize proposal distribution for better log-likelihood estimation. result Consistently outperforms state-of-the-art baselines in log-likelihood and parameter estimation.
This paper explores policy improvement using various f-divergences, enhancing stability in reinforcement learning.
problem Ensuring stability in reinforcement learning algorithms through policy improvement with trust regions.
method The paper considers a general class of f-divergences and derives policy update rules, including the KL divergence as a special case.
result The study reveals different policy updates and evaluations for various f-divergences, including Pearson χ 2 χ^2 χ 2 -divergence and KL divergence. ULA rapidly converges to target distribution without convexity assumptions.
problem Sampling from complex probability distributions efficiently.
method Unadjusted Langevin Algorithm with KL and Rényi divergence guarantees.
result ULA converges in KL divergence under log-Sobolev inequality.
We develop fast bounds for mixture model divergences.
problem Bounding the Kullback-Leibler divergence of mixtures.
method Piecewise log-sum-exp inequalities for closed-form bounds.
result Fast and generic method for entropy, cross-entropy, and KL divergence bounds.
PLA improves sampling from distributions under isoperimetry with faster KL divergence convergence.
problem Sampling from distributions with KL divergence under isoperimetry.
method Proximal Langevin Algorithm (PLA) with KL and Rényi divergence convergence guarantees.
result PLA achieves faster KL divergence convergence rates than ULA under log-Sobolev inequality.
SRFE clarifies KL divergences without unifying learning frameworks.
problem Inductive biases of KL divergences and their limitations.
method Introducing SRFE, a log-moment-based functional of the likelihood ratio.
result SRFE recovers KL divergences as limits and reveals a mean-variance tradeoff.
We consider a stochastic volatility model which captures relevant stylized facts of financial series, including the multi-scaling of moments. The volatility evolves according to a generalized Ornstein-Uhlenbeck processes with super-linear mean reversion. Using large deviations techniques, we determine the asymptotic sh…
Zigzag sampling algorithm efficiently samples from strongly log-concave distributions with low computational cost.
problem Sampling from strongly log-concave distributions efficiently and with low computational complexity.
method Zigzag sampling algorithm with warm start assumption, focusing on gradient evaluations.
result Achieves ε error in chi-square divergence with computational cost of O(κ²d^(1/2)(log(1/ε))^(3/2)) gradient evaluations.
New framework for logarithmically divergent integrals on manifolds with corners.
problem Logarithmically divergent integrals on manifolds with corners.
method Introduces new geometric framework and morphisms in logarithmic geometry.
result Functorial characterization of regularized integration.
Paper shows how sparse inversion speeds up log determinant derivatives.
problem Deriving log determinant derivatives for sparse matrices.
method Sparse inversion, selected inversion, accelerates computation.
result Derivative of log determinant can be computed faster with sparse inversion.
Rényi Neural Processes replace KL divergence with Rényi divergence to improve NP performance.
problem Parameterization coupling in Neural Processes leads to prior misspecification.
method Propose Rényi Neural Processes (RNP) by replacing KL divergence with Rényi divergence.
result Significant performance improvements in real-world problems, including better log-likelihoods.
Paper proves Shapley value convergence in Bayesian learning games.
problem Measuring contributions in cooperative games using Bayesian inference.
method Established convergence of Shapley value in parametric Bayesian learning games.
result Shapley value differences converge in probability to a limiting game.
Unified analysis of KL divergence using shifted composition for sampling.
problem Sampling from target distributions with KL divergence guarantees.
method Shifted composition rule applied to KL divergence, combining local error analysis and Girsanov's theorem.
result Unified KL guarantees for strongly log-concave, weakly log-concave, and log-Sobolev distributions.
The paper analyzes convergence rates of Langevin dynamics and Proximal Sampler using Φ Φ Φ -divergence.
problem Analyzing convergence rates of Langevin dynamics and Proximal Sampler.
method Extending mixing time analyses to Φ Φ Φ -divergence, using strong data processing inequalities. result Convergence of Φ Φ Φ -divergence to 0 exponentially fast along Unadjusted Langevin Algorithm and Proximal Sampler. Dynamic Vocabulary Pruning stabilizes LLM training by removing low-probability tokens.
problem Training Large Language Models (LLMs) with Reinforcement Learning (RL) causes numerical divergence between inference and training.
method Dynamic Vocabulary Pruning (DVP) constrains the RL objective to a safe vocabulary that excludes low-probability tokens.
result DVP stabilizes training by reducing systematic bias introduced by the extreme tail of the token distribution.
Bayesian approach estimates log-determinant with uncertainty quantification.
problem Intractable computation of log-determinant in large kernel matrices.
method Reinterpreting as Bayesian inference with prior bounds and evidence.
result Probabilistic estimates of log-determinant and uncertainty.
A new upper bound for variational inference improves the efficiency of Bayesian deep learning.
problem Improving variational inference in Bayesian deep learning.
method Presented a new upper bound (EUBO) for evidence, derived from KL-divergence and log marginal likelihood, and used SGD for optimization.
result The new upper bound (EUBO) is tighter than previous methods and outperforms state-of-the-art results in Bayesian neural networks.
Given i.i.d. observations of a random vector X ∈ R p X \in \mathbb{R}^p X ∈ R p , we study the problem of estimating both its covariance matrix Σ ∗ Σ^* Σ ∗ , and its inverse covariance or concentration matrix { Θ ∗ = ( Σ ∗ ) − 1 Θ^* = (Σ^*)^{-1} Θ ∗ = ( Σ ∗ ) − 1 .} We estimate Θ ∗ Θ^* Θ ∗ by minimizing an ℓ 1 \ell_1 ℓ 1 -penalized log-determinant Bregman divergence; in the multivariate G…
A new DC programming approach improves RBM training efficiency.
problem Improving the training efficiency of Restricted Boltzmann Machines (RBMs).
method Formulated a stochastic DC programming approach to minimize RBM log-likelihood.
result The new algorithm achieves higher log-likelihood more rapidly with the same computational budget.
LinMED is a new linear bandit algorithm with near-optimal regret bound.
problem Optimizing decision-making in linear bandit problems with sub-Gaussian distributions.
method LinMED is a randomized linear bandit algorithm with closed-form arm sampling probabilities.
result LinMED achieves a near-optimal regret bound of d n d\sqrt{n} d n up to logarithmic factors. New guarantees for VI in symmetric cases, extending previous results.
problem Symmetry in variational inference for complex distributions.
method Analysis of f f f -divergences and their stationary points under symmetry. result Symmetry-matching principles ensure recovery of mean and correlation matrix.
New method samples from non-log-concave distributions with weak dissipativity.
problem Sampling from distributions that are not log-concave and weakly dissipative.
method Taming scheme tailored to growth and decay properties of the target distribution.
result Explicit non-asymptotic guarantees for KL, TV, and Wasserstein distances.
A test assesses how well data fits a target density function.
problem Measuring how well data fits a target density function without assuming a specific form.
method Stein's method using Reproducing Kernel Hilbert Space functions to construct a divergence measure, estimated via V-statistic.
result The proposed test accurately assesses goodness of fit for various data types and contexts.
New divergence measures improve KL approximation.
problem Improving KL divergence approximation without AC condition.
method Introduced α α α -geodesical skew divergence. result Properties of α α α -geodesical skew divergence studied. Estimates log determinants using entropy for scalable machine learning.
problem Scalable calculation of matrix determinants is a bottleneck in machine learning.
method Maximum entropy framework with moment constraints for stochastic trace estimation.
result Significant improvement over state-of-the-art methods on various UFL sparse matrices.
New algorithm samples from log concave distributions efficiently.
problem Sampling from log concave distributions efficiently.
method Stochastic Proximal Langevin Algorithm (SPLA) with potential splitting.
result Established nonasymptotic convergence rates for SPLA.
New bounds for general unbounded loss functions, optimizing for various estimators.
problem Excess risk bounds for general unbounded loss functions, including log loss and squared loss.
method Optimized bounds for η η η -generalized Bayesian, MDL, and empirical risk minimization estimators, using v v v -GRIP and witness conditions. result Achieves i l d e O ( 1 / n ) ilde{O}(1/n) i l d e O ( 1/ n ) rates for certain loss functions under favorable v v v and small model complexity. EL_2O improves posterior inference without sampling noise.
problem Statistical inference of analytically non-tractable posteriors.
method Expectation optimization of L_2 distance squared between approximate and true posteriors.
result EL_2O provides a reliable estimate of posterior quality and converges rapidly.
A method to approximate posterior distributions using Monte Carlo and variational inference.
problem Lack of systematic understanding of how optimizing different objectives relates to approximating the posterior distribution.
method Divide and couple procedure to identify augmented proposal and target distributions.
result Maximizing the VI objective leads to an augmented variational distribution that approximates the posterior distribution.
This work broadens calibeating to various proper losses using Bregman divergence.
problem Calibration for a wide range of proper losses.
method Regret minimization and Bregman divergence approach.
result U-calibration results for a family of Tsallis losses with logarithmic regret and dimension independence.
New method estimates densities using Sobolev regularization, outperforming existing algorithms.
problem Non-parametric density estimation with clear inductive bias.
method Regularizes Sobolev norm of density, approximates kernel via sampling, uses natural gradients for optimization.
result Method ranks second best on ADBench anomaly detection benchmark.
This work generalizes calibeating for a broader range of proper losses using Bregman divergence.
problem Calibration for a wide range of proper losses beyond Brier and log loss.
method Regret minimization based on Bregman divergence for a family of proper losses.
result U-calibration results for a family of Tsallis losses with logarithmic regret and dimension independence.
New bounds improve training of VAEs on noisy data.
problem Training Variational AutoEncoders on noisy data.
method Proposes Kullback-Leibler and Rényi divergence bounds for log-likelihood.
result Numerically stable training without extra noise.
The paper analyzes the error in variational Bayesian NMF compared to Bayesian NMF.
problem Analyzing the variational approximation error in Bayesian NMF.
method Using algebraic geometrical methods, the paper derives an upper bound for the learning coefficient and a lower bound for the approximation error.
result The paper finds a lower bound for the approximation error, showing how well VBNMF approximates Bayesian NMF.
A robust multiclass logistic regression method using Tsallis divergence.
problem Noise robustness in multiclass logistic regression.
method Two-temperature logistic regression with Tsallis divergence.
result Significant robustness to outliers and noise.