Theory for RLHF generalization under reward shift and clipped KL.
problem Theoretical understanding of RLHF generalization, especially with reward shift and clipped KL.
method Developed generalization theory for RLHF, accounting for reward shift and clipped KL.
result Presented generalization bounds for RLHF, suggesting generalization error from sampling, reward shift, and KL clipping.
New bounds tighten the generalization error of Gibbs algorithm.
problem Bounding the generalization error of Gibbs algorithm.
method Characterization of generalization error in terms of symmetrized KL information.
result Exact characterization of Gibbs algorithm's expected generalization error.
New bounds for sequential tests under power-one error levels.
problem Determining stopping times for sequential tests with power-one error levels.
method Proved two lower bounds for stopping times under specific conditions.
result Upper and lower bounds for sequential tests are shown to be tight.
KL regularization helps RL algorithms by implicitly averaging q-values.
problem Understanding why KL regularization improves RL performance.
method An approximate value iteration scheme, studying KL and entropy regularization.
result Strong performance bound combining linear horizon dependency and averaging effect of estimation errors.
Unified analysis of KL divergence using shifted composition for sampling.
problem Sampling from target distributions with KL divergence guarantees.
method Shifted composition rule applied to KL divergence, combining local error analysis and Girsanov's theorem.
result Unified KL guarantees for strongly log-concave, weakly log-concave, and log-Sobolev distributions.
The paper analyzes error bounds and KL properties for noisy matrix recovery problems.
problem Noisy low-rank matrix recovery problems.
method Squared F-norm regularization, accelerated alternating minimization method.
result Established error bounds and KL properties for critical points and global minimizers.
Improved analysis for diffusion models reduces KL divergence error dependence on data dimension and discretization step size.
problem Analyze the convergence of diffusion-based generative models under minimal assumptions.
method Model the generation process as a composition of reverse ODE and noising steps, leveraging Wasserstein-type error control and noise addition.
result Achieved a linear dependence on data dimension and improved dependence on discretization step size for KL divergence error.
Private KL distribution estimation improved with instance-optimality.
problem Minimizing KL divergence between true and estimated distributions.
method Construct minimax optimal private estimators, then focus on instance-optimality.
result Achieved instance-optimality up to constant factors for KL estimation.
New algorithm minimizes inclusive KL for VI, improving accuracy.
problem Improving variational inference accuracy with KL(p||q).
method Markovian score climbing (MSC) using stochastic gradients.
result MSC converges to local optimum of inclusive KL without bias.
Improved KL bounds and Wasserstein guarantees for diffusion flow matching under minimal conditions.
problem Theoretical convergence properties of Brownian motion based diffusion flow matching.
method Refined analysis under Kullback-Leibler and 2-Wasserstein distances.
result State-of-the-art scaling in KL convergence bounds under minimal conditions.
KL annealing helps VAEs avoid posterior collapse and overfitting.
problem Posterior collapse and overfitting in VAEs.
method Theoretical analysis of learning dynamics with KL annealing.
result Posterior collapse is inevitable when β β β exceeds a threshold. Conditional diffusion models can approximate target distributions well with Gaussian-mixture reverse kernels.
problem Approximating target distributions in conditional diffusion models.
method Using finite Gaussian mixtures with ReLU-network logits as reverse kernels, reducing the problem to static conditional density approximation.
result The resulting neural reverse-kernel class is dense in conditional KL divergence under exact terminal matching.
Paper analyzes kNN estimator for KL divergence, proving its optimality.
problem Estimating KL divergence from identical samples.
method kNN estimator based on nearest neighbor distances.
result kNN method is asymptotically rate optimal for KL divergence estimation.
In this paper, we study the Kurdyka-Łojasiewicz (KL) exponent, an important quantity for analyzing the convergence rate of first-order methods. Specifically, we develop various calculus rules to deduce the KL exponent of new (possibly nonconvex and nonsmooth) functions formed from functions with known KL exponents. In …
Bayesian sequence prediction is a simple technique for predicting future symbols sampled from an unknown measure on infinite sequences over a countable alphabet. While strong bounds on the expected cumulative error are known, there are only limited results on the distribution of this error. We prove tight high-probabil…
Study shows how neural networks generalize with minimal training data.
problem Understanding how neural networks generalize with limited data.
method Mean-field analysis of KL-regularized empirical risk minimization.
result Generalization error rate is O ( 1 / n ) \mathcal{O}(1/n) O ( 1/ n ) for large n n n . RHMC accelerates sampling from log-concave distributions.
problem Sampling from log-concave probability distributions efficiently.
method RHMC uses simulated Hamiltonian dynamics with random integration times.
result RHMC converges exponentially fast in KL divergence for log-concave distributions.
New schemes improve error estimates for sampling from non-log-concave distributions.
problem Improving sampling from non-log-concave distributions with super-linear drift growth.
method Developed tamed Euler and randomized Euler schemes with error estimates.
result Near-optimal error bounds for sampling and optimization problems.
Researchers establish bounds for SGMs' KL and Wasserstein divergences under various noise schedules.
problem Estimating the error between target and estimated distributions in SGMs.
method Established upper bounds for KL divergence and Wasserstein distance, incorporating target distribution properties and SGM hyperparameters.
result Optimal noise schedules identified for SGMs, improving generative quality.
Mutual information bounds generalization error in variational classifiers.
problem Controlling overfitting in variational classifiers.
method Derive bounds on generalization error using mutual information.
result Mutual information bounds the generalization error in variational classifiers.
Study compares chi-squared divergence and KL-divergence posteriors for PAC-Bayesian bounds.
problem Investigates optimal posteriors for PAC-Bayesian bounds using chi-squared divergence.
method Analyzes bounds for three distance functions, derives FP equations for computation.
result Chi-squared divergence based posteriors have weaker bounds and worse test errors.
Paper improves PAC-Bayes bounds using a better-than-KL divergence.
problem Estimating the generalization error of stochastic algorithms.
method Developed new PAC-Bayes bounds with a novel divergence.
result Achieved strictly tighter bounds than the KL divergence.
New method proves dimension-free convergence for ULD in KL divergence.
problem Polynomial scaling of existing convergence guarantees in high dimensions.
method Refined KL local error framework, focusing on tr(H) instead of d.
result First dimension-free KL divergence bounds for discretized ULD.
Improved error estimate for SGLD sampling algorithm.
problem Establishing a precise error bound for SGLD.
method Sharp uniform-in-time error estimate for SGLD under mild assumptions.
result Uniform-in-time O ( η 2 ) O(η^2) O ( η 2 ) bound for KL-divergence between SGLD and Langevin diffusion. Paper proposes a method to stabilize estimation of KL divergence using a discriminator in RKHS.
problem High variance and instability in estimating KL divergence using neural network discriminators.
method Developed a novel construction of the discriminator in RKHS, controlled its complexity, and proved the consistency of the estimator.
result Reduced variance and stabilized training of KL divergence estimates.
A1GM method improves efficiency in reconstructing missing data using KL divergence.
problem Efficiently reconstructing missing data in matrices.
method Fast non-gradient-based rank-1 NMF using KL divergence.
result A1GM outperforms gradient methods in efficiency with competitive reconstruction errors.
Optimized α \alpha α -posteriors reduce KL divergence from true posterior in parametric misspecification.
problem Reduction of KL divergence from true posterior in parametric model misspecification.
method Derivation of Bernstein-von Mises theorem and optimization of α \alpha α -posteriors. result Optimized α \alpha α -posteriors minimize KL divergence from true posterior, especially in severe misspecification. FORE evaluates occupancy ratios without requiring Bellman completeness.
problem Offline reinforcement learning occupancy ratio estimation.
method Fitted occupancy-ratio evaluation (FORE) using adjoint Bellman recursion.
result FORE achieves convergence in KL without Bellman completeness.
New bounds study class-specific generalization error in machine learning.
problem Existing generalization theories assume uniform class performance, but in practice, classes vary significantly.
method Developed novel information-theoretic bounds using KL divergence and CMI.
result Theoretical bounds accurately capture complex class-generalization error behavior.
New method decomposes KL error using refined information and mode interactions.
problem Learning probability distributions over discrete variables with higher-order interactions.
method Using information geometry, refined mode interactions, and a novel Monte-Carlo sampling technique.
result Complete decomposition of KL error and efficient data use.
Study improves density estimation for compact domains using h h h -lifted KL divergence.
problem Estimating probability density functions on compact domains.
method Introduced h h h -lifted Kullback--Leibler (KL) divergence for risk minimization. result Proved O ( 1 / n ) \mathcal{O}(1/{\sqrt{n}}) O ( 1/ n ) bound on estimation error. Decentralized Bayesian learning reduces KL-divergence exponentially.
problem Efficiently learning posterior distributions in a decentralized setting.
method Decentralized Langevin dynamics in a non-convex setting.
result The algorithm converges to the target posterior distribution with exponential decrease in KL-divergence and polynomial decrease in error contributions.
Flow-based models generate data with improved theoretical guarantees.
problem Theoretical analysis of flow-based generative models.
method Proximal gradient descent in Wasserstein space for JKO flow model.
result KL guarantee of data generation by JKO flow model is O ( ε 2 ) O(\varepsilon^2) O ( ε 2 ) . New algorithm clusters trajectories from multiple Markov chains with near-optimal error.
problem Clustering trajectories from multiple unknown Markov chains.
method Two-stage algorithm: spectral clustering followed by likelihood-based refinement.
result Achieves near-optimal clustering error with high probability.
We speed up marginal inference by ignoring factors that do not significantly contribute to overall accuracy. In order to pick a suitable subset of factors to ignore, we propose three schemes: minimizing the number of model factors under a bound on the KL divergence between pruned and full models; minimizing the KL dive…
Accelerates convergence in global non-convex optimization with reversible diffusion.
problem Global non-convex optimization challenges.
method Utilizes reversible diffusion processes with adaptive diffusion coefficients.
result Accelerated convergence with reduced discretization error.
This work analyzes discrete diffusion models using stochastic integrals, providing error bounds and insights.
problem Error analysis for discrete diffusion models remains less understood.
method Proposes a comprehensive framework based on Lévy-type stochastic integrals.
result Obtains the first error bound for the τ τ τ -leaping scheme in KL divergence. We define On-Average KL-Privacy and present its properties and connections to differential privacy, generalization and information-theoretic quantities including max-information and mutual information. The new definition significantly weakens differential privacy, while preserving its minimalistic design features such …
Convolutional neural networks improve KL grade prediction from Indian knee radiographs.
problem Improving accuracy of knee osteoarthritis grading from Indian radiographs.
method Two-stage approach: object detection followed by regression.
result Fine-tuning model on private hospital data reduces mean absolute error from 1.09 to 0.28.
The paper addresses instability in KL divergence estimation using a neural network discriminator.
problem Unstable estimation of KL divergence due to discriminator complexity.
method Using a Reproducing Kernel Hilbert Space (RKHS) to control discriminator complexity.
result Theoretical bound on error probability of KL estimates based on discriminator complexity in RKHS.
Paper introduces a diagnostic for approximate inference methods.
problem Estimating errors in probabilistic inference algorithms, especially for approximate methods.
method Repeatedly simulate datasets from the prior and perform inference on each, estimating a symmetric KL-divergence.
result A diagnostic for approximate inference methods can be estimated using symmetric KL-divergence.
The study analyzes transfer learning using information theory.
problem Transfer learning in different distributions.
method Information-theoretic analysis, focusing on KL divergence.
result Upper bounds for general transfer learning algorithms and specific ERM.
New bound matches exact generalization error for quadratic Gaussian problem.
problem Understanding generalization error in quadratic Gaussian problems.
method Information-theoretic approach with new ingredients.
result Exact tight bound for generalization error.
Eye Movement analysis with Hidden Markov Models (EMHMM) is a method for modeling eye fixation sequences using hidden Markov models (HMMs). In this report, we run a simulation study to investigate the estimation error for learning HMMs with variational Bayesian inference, with respect to the number of sequences and the …
Paper analyzes Langevin dynamics for multimodal Gaussian mixtures, controlling errors across dimensions.
problem Challenges in obtaining stable diffusion-based samplers in high- and infinite-dimensional settings.
method Study of preconditioned Annealed Langevin Dynamics (ALD) for Gaussian mixtures, focusing on Euler-Maruyama (EM) and exponential-integrator schemes.
result Proves dimension-uniform KL bounds for the exponential-integrator scheme, allowing arbitrarily small divergence with dimension.
We investigate the use of alternative divergences to Kullback-Leibler (KL) in variational inference(VI), based on the Variational Dropout \cite{kingma2015}. Stochastic gradient variational Bayes (SGVB) \cite{aevb} is a general framework for estimating the evidence lower bound (ELBO) in Variational Bayes. In this work, …
The paper analyzes the reward improvement of aligned policies in large language models.
problem Optimizing policies in large language models while staying close to a reference policy.
method Information-theoretic analysis and reduction to exponential order statistics.
result Information-theoretic upper bounds on reward improvement are derived.
This paper provides guarantees for DFM models using KL divergence.
problem Ensuring generative models match target distributions efficiently.
method Using KL divergence and Brownian motion bridge for generative models.
result Non-asymptotic guarantees for DFM models under specific conditions.