The paper analyzes the bias-variance tradeoff for Bregman divergences.
problem Understanding the bias-variance tradeoff for Bregman divergences.
method Analyzes the bias-variance tradeoff through operations in dual space.
result Derives several results including a generalized law of total variance and ensembling operations.
Improved KL divergence estimators for normalizing flows lead to faster convergence and better approximations.
problem Estimating KL divergences for normalizing flows efficiently and accurately.
method Path-gradient estimators for reverse and forward KL divergences.
result Path-gradient estimators lead to faster convergence and better approximation results.
POP3D is a new reinforcement learning algorithm that improves upon PPO.
problem The shortcomings of existing reinforcement learning algorithms.
method Policy Optimization with Penalized Point Probability Distance (POP3D) as a lower bound to the square of total variance divergence.
result POP3D is highly competitive compared to PPO in various benchmarks.
Deep learning models can have low bias and variance, contrary to classical theory.
problem Understanding the performance of deep learning models at high complexity.
method Developed a fine-grained bias-variance decomposition for random feature kernel regression, analyzing the effects of sampling, initialization, and labels.
result The variance terms exhibit non-monotonic behavior and can diverge at the interpolation boundary, even in the absence of label noise.
Generalizes bias-variance decomposition for Bregman divergences.
problem No specific problem stated; generalization of bias-variance for Bregman divergences.
method Provided a generalization of the bias-variance decomposition for Bregman divergences.
result A clear, standalone derivation of the bias-variance decomposition for Bregman divergences.
We analyze total, asymmetric and frequency connectedness between oil and forex markets using high-frequency, intra-day data over the period 2007 -- 2017. By employing variance decompositions and their spectral representation in combination with realized semivariances to account for asymmetric and frequency connectednes…
This paper generalizes beta divergence beyond its classical form associated with power variance functions of Tweedie models. Generalized form is represented by a compact definite integral as a function of variance function of the exponential dispersion model. This compact integral form simplifies derivations of many pr…
Contrastive Divergence (CD) and Persistent Contrastive Divergence (PCD) are popular methods for training the weights of Restricted Boltzmann Machines. However, both methods use an approximate method for sampling from the model distribution. As a side effect, these approximations yield significantly different biases and…
LMC algorithm converges to target in Chi-squared and Renyi divergence.
problem Sampling from target distribution using LMC with strong dissipativity and smoothness conditions.
method LMC algorithm with strong dissipativity and first-order smoothness, initialized with Gaussian.
result LMC reaches ε-neighborhood of target in Chi-squared and Renyi divergence in O(λ²dε⁻¹) steps.
SRFE clarifies KL divergences without unifying learning frameworks.
problem Inductive biases of KL divergences and their limitations.
method Introducing SRFE, a log-moment-based functional of the likelihood ratio.
result SRFE recovers KL divergences as limits and reveals a mean-variance tradeoff.
Black box variational inference (BBVI) with reparameterization gradients triggered the exploration of divergence measures other than the Kullback-Leibler (KL) divergence, such as alpha divergences. In this paper, we view BBVI with generalized divergences as a form of estimating the marginal likelihood via biased import…
Paper derives a new lower bound for KL-divergence using HCRB.
problem Estimating KL-divergence between distributions.
method Using Hammersley-Chapman-Robbins bound and information geometry.
result New lower bound for KL-divergence derived from HCRB.
New method for fast, accurate Gaussian process regression with data guarantees.
problem Slow and unreliable GP inference methods for nonparametric regression.
method Developed a novel objective function (preconditioned Fisher divergence) for scalable approximate GP regression with finite-data guarantees.
result Minimizing the pF divergence provides pointwise mean and variance estimates with tight 2-Wasserstein distance bounds and comparable empirical performance to variational sparse GPs.
The paper decomposes unsupervised learning's generalization error into model, data, and variance components.
problem Understanding the components of unsupervised learning's generalization error.
method Information-geometric decomposition of the Kullback-Leibler generalization error.
result The optimal rank in ε ε ε -PCA is the noise floor, balancing model-error gain and data-bias cost. Sample variance decay is shown in deep ReLU networks, impacting training dynamics.
problem Sample variance decay in deep ReLU networks during training.
method Decomposed total variance into sample variance and network-averaged sum of sample mean and variance.
result Sample variance decays in later layers of deep ReLU networks, impacting training dynamics.
Improves diffusion models by controlling total variance and signal-to-noise-ratio.
problem Long sampling time in diffusion models.
method Total-Variance/Signal-to-Noise-Ratio (TV/SNR) disentangled framework.
result Improves generation performance by controlling TV and SNR independently.
We describe the underlying probabilistic interpretation of alpha and beta divergences. We first show that beta divergences are inherently tied to Tweedie distributions, a particular type of exponential family, known as exponential dispersion models. Starting from the variance function of a Tweedie model, we outline how…
Large batch sizes reduce gradient variance in DP-SGD, improving privacy.
problem Understanding why large batch sizes work in DP-SGD.
method Decomposed total gradient variance into subsampling and noise-induced variances, proving batch size independence in the limit.
result Large batch sizes reduce effective total gradient variance, improving privacy in DP-SGD.
New divergences improve estimation and GAN training performance.
problem Improving estimation and training in machine learning models.
method Function-space regularized Rényi divergences.
result New divergences reduce variance and improve training performance.
On a compact n n n -dimensional manifold M M M , it is well known that a critical metric of the total scalar curvature, restricted to the space of metrics with unit volume, is Einstein. It has been conjectured that a critical metric of the total scalar curvature, restricted to the space of metrics with constant scalar curvat…
Flow matching KL divergence bound derived for smooth distributions.
problem Estimating smooth distributions efficiently.
method Deterministic upper bound on KL divergence derived from flow-matching loss.
result Flow matching achieves nearly minimax-optimal efficiency under TV distance.
Optimized α \alpha α -posteriors reduce KL divergence from true posterior in parametric misspecification.
problem Reduction of KL divergence from true posterior in parametric model misspecification.
method Derivation of Bernstein-von Mises theorem and optimization of α \alpha α -posteriors. result Optimized α \alpha α -posteriors minimize KL divergence from true posterior, especially in severe misspecification. Paper introduces Wasserstein total correlation for disentangled representation learning.
problem Learning disentangled representations from data.
method Adversarial training of a critic to estimate Wasserstein total correlation in variational and Wasserstein autoencoders.
result Proposed method achieves comparable disentanglement performance with less reconstruction loss.
Estimates f-divergences with strong structural assumptions in high dimensions.
problem Estimating f-divergences under weak assumptions is hard.
method Proposes an easy-to-implement estimator under stronger structural assumptions.
result Estimator works well in high dimensions and converges faster.
New Monte Carlo method outperforms existing strategy for estimating Sobol' indices.
problem Estimating first-and total-orders Sobol' indices accurately.
method Comparing two Monte Carlo estimators for Sobol' indices.
result New method outperforms current approach in accuracy.
New framework estimates staged tree models using hierarchical clustering on the probability simplex.
problem Estimating staged tree models with context-specific dependencies.
method Hierarchical clustering on the probability simplex, using simplex-based divergences and linkage methods.
result Total Variation divergence with Ward.D2 linkage produces staged trees with better model fit, structure recovery, and computational efficiency.
The paper proves gap properties for critical metrics under specific conditions.
problem Proving gap properties for critical metrics under divergence-free Bach tensor condition.
method Analyzing critical point equation of total scalar curvature with divergence-free Bach tensor.
result Proves gap properties for n ≥ 5 n \geq 5 n ≥ 5 and a similar condition for n = 4 n=4 n = 4 . New dispersion indices based on inaccuracy and divergence introduced for information measures.
problem Measuring variability in uncertainty measures.
method Introducing new dispersion indices based on Kerridge inaccuracy and Kullback-Leibler divergence.
result Properties, bounds, and examples of new dispersion indices presented.
Paper proposes a method to stabilize estimation of KL divergence using a discriminator in RKHS.
problem High variance and instability in estimating KL divergence using neural network discriminators.
method Developed a novel construction of the discriminator in RKHS, controlled its complexity, and proved the consistency of the estimator.
result Reduced variance and stabilized training of KL divergence estimates.
f f f -divergences are a general class of divergences between probability measures which include as special cases many commonly used divergences in probability, mathematical statistics and information theory such as Kullback-Leibler divergence, chi-squared divergence, squared Hellinger distance, total variation distance e…
Paper analyzes kNN estimator for KL divergence, proving its optimality.
problem Estimating KL divergence from identical samples.
method kNN estimator based on nearest neighbor distances.
result kNN method is asymptotically rate optimal for KL divergence estimation.
New methods minimize GFlowNet training divergences for better sampling.
problem Training GFlowNets with KL divergence leads to biased and high-variance estimators.
method Design and implement efficient estimators for four divergence measures.
result Properly minimizing these divergences yields a provably correct and effective training scheme.
A new model corrects inhomogeneity in Optimal Transport with Boundary.
problem Inhomogeneity in UROT models for Optimal Transport with Boundary.
method Proposed a modified entropic regularization term to make UROT models homogeneous.
result Homogeneous UROT model preserves properties of standard UROT while correcting inhomogeneity.
The paper develops estimators for variance in graph structures using fused lasso.
problem Variance estimation in graph-structured problems.
method Developed linear time estimator for homoscedastic case and total variation regularization estimator for heteroscedastic case.
result Minimax rates and consistency for variance estimation in various graph structures.
Privacy amplification improved through contraction coefficients and E γ E_γ E γ -divergence.
problem Improving privacy guarantees in iterative algorithms.
method Using contraction coefficients derived from E γ E_γ E γ -divergence to determine differential privacy parameters. result Tighter bounds on differential privacy parameters of iterative algorithms.
Proves Sard conjecture for specific distributions, controlling divergence of vector fields.
problem Proving the Sard conjecture for certain types of distributions.
method Constructs a singular distribution capturing essential abnormal lifts, proving the conjecture for rank 3 distributions in dimension 4 and generic corank 1 distributions.
result Proves the Sard conjecture for generic co-rank one distributions.
This work analyzes the statistical properties of adaptive gradient methods.
problem Lack of understanding of the statistical properties of adaptive gradient methods.
method Theoretical analyses and experiments on the variance of update magnitudes.
result The variance of update magnitudes is an increasing and bounded function of time, not diverging.
Two approaches integrate qualitative views into portfolio optimization, showing aggregation methods outperform robust optimization.
problem Incorporating qualitative views into portfolio optimization models.
method Robust optimization and order aggregation methods.
result Aggregation methods outperform robust optimization in portfolio performance analysis.
Optimizes deep neural network initialization variance for better performance.
problem Improving deep neural network performance through optimal initialization variance.
method Using SGD dynamics and Fokker-Planck equations, we study the relationship between initialization and expected loss function.
result An optimal condition for initialization variance that leads to lower training loss and higher test accuracy.
We quantify predictive uncertainty using the posterior predictive variance.
problem Quantifying uncertainty in predictive models.
method Using the law of total variance, we generate expansions for the posterior predictive variance.
result Identify the main contributors to prediction intervals and quantify term-wise uncertainty.
A new model relaxes constraints on exponential dispersion models.
problem Tight conditions on cumulant function limit the class of exponential dispersion models.
method Introduces K-LED model with Legendre cumulant function and Bregman divergence guidance.
result The model allows for easier computation of mean parameter and includes various distributions.
This work broadens calibeating to various proper losses using Bregman divergence.
problem Calibration for a wide range of proper losses.
method Regret minimization and Bregman divergence approach.
result U-calibration results for a family of Tsallis losses with logarithmic regret and dimension independence.
New method for optimizing complex composite functions with reduced variance.
problem Optimizing multi-level composite functions with nested random and smooth mappings.
method Normalized proximal approximate gradient (NPAG) method with nested stochastic variance reduction.
result Total sample complexity of O ( ε − 3 ) O(ε^{-3}) O ( ε − 3 ) in expectation and O ( N + N ε − 2 ) O(N+\sqrt{N}ε^{-2}) O ( N + N ε − 2 ) in finite-sum cases. We focus on the maximum regularization parameter for anisotropic total-variation denoising. It corresponds to the minimum value of the regularization parameter above which the solution remains constant. While this value is well know for the Lasso, such a critical value has not been investigated in details for the total…
This work generalizes calibeating for a broader range of proper losses using Bregman divergence.
problem Calibration for a wide range of proper losses beyond Brier and log loss.
method Regret minimization based on Bregman divergence for a family of proper losses.
result U-calibration results for a family of Tsallis losses with logarithmic regret and dimension independence.
QP improves Gaussian process inference by minimizing Wasserstein distance.
problem Approximate inference in Gaussian processes using KL divergence is inadequate.
method Quantile Propagation (QP) minimizes Wasserstein distance instead of KL divergence.
result QP outperforms EP and variational Bayes in classification and Poisson regression.
The paper improves generalization bounds using interpolation between various divergences.
problem Improving generalization bounds in machine learning.
method Derives new PAC-Bayes generalization bounds based on ( f , Γ ) (f, Γ) ( f , Γ ) -divergence and interpolates between various divergences. result Connects derived bounds to earlier statistical learning results and provides practical training objectives.
RHMC accelerates sampling from log-concave distributions.
problem Sampling from log-concave probability distributions efficiently.
method RHMC uses simulated Hamiltonian dynamics with random integration times.
result RHMC converges exponentially fast in KL divergence for log-concave distributions.