Optimal pre-processing reduces disparate impact by minimizing total variation distance.
problem Achieving fairness in data outputs based on protected attributes.
method Using pre-processing to enforce fairness, minimizing total variation distance between pre-processed and original data distributions.
result The problem of fairness can be formulated as a linear program, efficiently solvable.
Paper proposes a method to estimate total variation distance for synthetic data fidelity.
problem Assessing the fidelity of synthetic data generated by AI.
method Discriminative approach to estimate total variation distance between two distributions.
result Estimation of total variation distance reduces to quantifying Bayes risk in classification.
Sharp inequality between TV and Hellinger distances for Gaussian mixtures.
problem Understanding the relationship between total variation and Hellinger distances for Gaussian mixtures.
method Established a general upper bound on Hellinger distance in terms of TV distance raised to a power, demonstrating sharpness with specific examples.
result The Hellinger distance between two Gaussian mixtures is bounded by the TV distance raised to a power 1 − o ( 1 ) 1-o(1) 1 − o ( 1 ) , where o ( 1 ) o(1) o ( 1 ) is of order 1 / log log ( 1 / T V ) 1/\log\log(1/\mathrm{TV}) 1/ log log ( 1/ TV ) . The total variation distance is a core statistical distance between probability measures that satisfies the metric axioms, with value always falling in [ 0 , 1 ] [0,1] [ 0 , 1 ] . This distance plays a fundamental role in machine learning and signal processing: It is a member of the broader class of f f f -divergences, and it is related to …
New method relaxes TV distance for two-sample testing without distributional assumptions.
problem Challenges in certifying equality or providing tight bounds on TV distance for two distributions.
method Examined blurred total variation distance, a relaxation of TV distance.
result Provided theoretical guarantees for upper and lower bounds on blurred TV distance.
Error estimates found between SGD with momentum and Langevin diffusion.
problem Quantifying the difference between SGD with momentum and Langevin diffusion.
method Established error estimates using 1-Wasserstein and total variation distances.
result Quantitative error estimates between SGD with momentum and underdamped Langevin diffusion.
Algorithm learns affine transformations robustly from corrupted samples.
problem Learning affine transformations from corrupted samples.
method New geometric certificate and iterative improvement method.
result Total variation distance of O ( ε ) O(ε) O ( ε ) between learned and original distributions. Example shows learnable distributions not privately learnable.
problem Learnable distributions under non-private conditions not transferable to differential privacy.
method Example of a distribution class learnable up to constant error in total variation distance but not under differential privacy.
result Contradicts conjecture of Ashtiani on learnability under differential privacy.
We develop a variational theory of geodesics for the canonical variation of the metric of a totally geodesic foliation. As a consequence, we obtain comparison theorems for the horizontal and vertical Laplacians. In the case of Sasakian foliations, we show that sharp horizontal and vertical comparison theorems for the s…
The study improves PAC-Bayesian bounds for adversarial generative models.
problem Improving generalization bounds for adversarial generative models.
method Extending PAC-Bayesian theory to generative models, developing bounds for Wasserstein and total variation distances.
result New training objectives for Wasserstein and Energy-Based GANs.
New bounds on neural network convergence using information theory.
problem Quantifying convergence rates of neural networks to Gaussian distributions.
method Entropic inequalities and Gaussian approximations.
result Improved convergence rates in various distances for neural networks.
Efficiently estimates binary product distributions with privacy.
problem Estimating means of binary product distributions privately and accurately.
method Polynomial time, pure differential privacy approach.
result Optimal sample complexity with polylogarithmic factors.
Robust Bayesian inference improves model performance on discrete data.
problem Misspecification of discrete-valued models leads to poor inference and prediction.
method Total Variation Distance (TVD) for discrepancy, efficient estimator and inference method.
result Our approach significantly improves predictive performance on various data.
Reinforcement learning mimics expert behavior.
problem Learning from expert demonstrations in reinforcement learning.
method Reduction to reinforcement learning with a stationary reward.
result Expert reward can be recovered and imitation learning is bounded.
Improved error estimate for SGLD sampling algorithm.
problem Establishing a precise error bound for SGLD.
method Sharp uniform-in-time error estimate for SGLD under mild assumptions.
result Uniform-in-time O ( η 2 ) O(η^2) O ( η 2 ) bound for KL-divergence between SGLD and Langevin diffusion. Diffusion models achieve nearly optimal distribution estimation in various spaces.
problem Theoretical limitations of diffusion modeling for distribution estimation.
method Analysis of approximation and generalization abilities of diffusion models in Besov spaces.
result Diffusion models achieve nearly minimax optimal estimation rates in total variation and Wasserstein distances.
Parallel sampling for smooth distributions with fast convergence.
problem Efficiently sampling from distributions with smooth densities.
method Parallelization of Langevin algorithms under log-Sobolev inequalities.
result Samples close to target distribution with low KL divergence or TV distance.
Study shows private learning of mixtures of Gaussians is possible with polynomial samples.
problem Estimating mixtures of Gaussians under differential privacy constraints.
method Developed a new framework for privately learning mixtures of Gaussians without structural assumptions.
result Polynomial number of samples (poly(k,d,1/α,1/ε,log(1/δ))) sufficient for estimation up to total variation distance α with (ε, δ)-DP.
Optimized α \alpha α -posteriors reduce KL divergence from true posterior in parametric misspecification.
problem Reduction of KL divergence from true posterior in parametric model misspecification.
method Derivation of Bernstein-von Mises theorem and optimization of α \alpha α -posteriors. result Optimized α \alpha α -posteriors minimize KL divergence from true posterior, especially in severe misspecification. We study density estimation for classes of shift-invariant distributions over R d \mathbb{R}^d R d . A multidimensional distribution is "shift-invariant" if, roughly speaking, it is close in total variation distance to a small shift of it in any direction. Shift-invariance relaxes smoothness assumptions commonly used in non-p…
uHMC achieves fast mixing in high dimensions with gradient evaluations.
problem Quantifying mixing time of uHMC in high dimensions.
method Construction of successful couplings for uHMC.
result uHMC mixes in total variation with logarithmic dependence on dimension.
The paper examines how to test if two learning algorithms produce similar outcomes.
problem Testing if two learning algorithms produce similar outcomes when trained on different data sets.
method Using Total Variation (TV) distance to measure similarity of posterior distributions.
result TV indistinguishable learning rules are equivalent to existing stability notions and can be statistically amplified.
The study bounds the stability of Gaussian mixtures under small perturbations.
problem Stability of Gaussian mixtures under small changes in distribution.
method Deriving an explicit bound on parameter stability of spherical Gaussian Mixture Models (sGMM) in a pre-defined model class.
result Upper bound on parameter distance of close sGMMs to the original sGMM, dependent only on the original model.
Two new deterministic offspring selection methods reduce statistical distance in SMC and pMCMC.
problem Improving the performance of resampling in SMC methods.
method Proposes two deterministic offspring selection methods to minimize KL divergence and TV distance.
result Our methods outperform or match state-of-the-art resampling schemes on benchmarks.
Graphs with non-negative Ollivier-Ricci curvature cannot be expanders.
problem Understanding the relationship between graph curvature and expansion properties.
method Proving an inequality linking isoperimetric profiles to total variation decay of random walks.
result Graphs with non-negative Ollivier-Ricci curvature cannot be expanders.
The study examines the limitations of bi-Lipschitz Normalizing Flows in approximating certain distributions.
problem The expressivity of bi-Lipschitz Normalizing Flows in approximating specific target distributions.
method Characterization of expressivity through lower bounds on Total Variation distance and discussion of potential remedies.
result Several target distributions are difficult to approximate using bi-Lipschitz Normalizing Flows, and lower bounds on their approximation are provided.
New sampling method improves efficiency for diffusion models.
problem Efficient sampling from arbitrary smooth distributions in polynomial time.
method Randomized midpoint method for log-concave sampling.
result Achieves best known dimension dependence ( O ~ ( d 5 / 12 ) \widetilde O(d^{5/12}) O ( d 5/12 ) ) for total variation distance. Paper explores robust estimators for kernel exponential families using smoothed total variation distances.
problem Outliers can severely impact classical estimators in statistical inference.
method Proposes smoothed total variation (STV) distance as a class of IPMs for robust estimation of kernel exponential families.
result STV-based estimators are robust against distribution contamination for kernel exponential families.
New method trains Markov kernels for efficient sampling.
problem Efficient sampling from complex probability distributions.
method Adversarial learning of involutive Metropolis-Hastings kernels.
result Minimizes total variation distance to empirical data.
Unsupervised learning of disentangled representations involves uncovering of different factors of variations that contribute to the data generation process. Total correlation penalization has been a key component in recent methods towards disentanglement. However, Kullback-Leibler (KL) divergence-based total correlatio…
The paper solves robust learning of Gaussian mixtures with nearly optimal guarantees.
problem Learning a high-dimensional Gaussian mixture model with corrupted samples.
method Introduces a new framework called strong observability to circumvent the challenge of learning individual components.
result Achieves optimal robustness guarantees of ε ε ε in total variation distance for any constant number of components. New framework estimates staged tree models using hierarchical clustering on the probability simplex.
problem Estimating staged tree models with context-specific dependencies.
method Hierarchical clustering on the probability simplex, using simplex-based divergences and linkage methods.
result Total Variation divergence with Ward.D2 linkage produces staged trees with better model fit, structure recovery, and computational efficiency.
Paper relaxes differential privacy for correlated features, improving privacy-utility trade-off.
problem Standard differential privacy ignores feature correlation, leading to suboptimal privacy-utility balance.
method Introduces CorrDP framework that accounts for feature correlation, using total variation distance for quantification.
result CorrDP algorithms outperform standard DP in synthetic and real-world datasets with insensitive features.
We consider the initial situation where a dataset has been over-partitioned into k k k clusters and seek a domain independent way to merge those initial clusters. We identify the total variation distance (TVD) as suitable for this goal. By exploiting the relation of the TVD to the Bayes accuracy we show how neural networ…
New schemes improve error estimates for sampling from non-log-concave distributions.
problem Improving sampling from non-log-concave distributions with super-linear drift growth.
method Developed tamed Euler and randomized Euler schemes with error estimates.
result Near-optimal error bounds for sampling and optimization problems.
Sharp bounds found on expert error in binary advice aggregation.
problem Aggregating binary advice from conditionally independent experts.
method Sharp upper and lower bounds on optimal error probability in asymmetric case.
result Sharp bounds recover and sharpen known results in symmetric case.
Combines MALA and Adam for efficient uncertainty quantification in deep learning.
problem Uncertainty estimation in deep neural networks.
method Integrates Metropolis Adjusted Langevin Algorithm (MALA) with momentum-based optimization (Adam) for efficient sampling from posterior distributions.
result The algorithm approximates the Gibbs posterior in total variation distance and efficiently quantifies epistemic uncertainty.
New algorithm samples from log-concave distributions with high accuracy in polynomial time.
problem Sampling from log-concave distributions with high accuracy in infinity distance.
method Directly converts continuous samples from K K K with total-variation bounds to samples with infinity bounds. result Output a point ε ε ε -close to π π π in infinity distance with runtime bounds that depend on polylogarithmic and polynomial factors of 1 / ε 1/ε 1/ ε . The paper shows diffusion models can converge faster to a target distribution with low-dimensional structure.
problem Improving the convergence rate of diffusion models to target distributions.
method Analyzing DDIM and DDPM samplers under low-dimensional structure assumptions.
result The iteration complexities of DDIM and DDPM are no greater than k / ε k/\varepsilon k / ε in total variation distance. Efficiently learns tree-structured Ising models with minimal samples.
problem Learning tree-structured Ising models efficiently and accurately.
method Plug-in estimator for mutual information using the Chow-Liu algorithm.
result Proper learning of tree-structured Ising models with O ( n ln n / ε 2 ) O(n \ln n/ε^2) O ( n ln n / ε 2 ) samples. We generalize to tree graphs obtained by connecting path graphs an oracle result obtained for the Fused Lasso over the path graph. Moreover we show that it is possible to substitute in the oracle inequality the minimum of the distances between jumps by their harmonic mean. In doing so we prove a lower bound on the comp…
Study clusters distributions with known or unknown clusters using distribution testing.
problem Cluster distributions that are ε \varepsilon ε -far in total variation. method Distribution testing approach to establish upper and lower bounds on sample complexity.
result Achieves tight sample complexity bounds for all regimes (up to a logarithmic factor).
New findings show score matching's accuracy doesn't ensure numerical stability in diffusion sampling.
problem Numerical stability issues in diffusion sampling despite small forward-marginal error.
method Constructing a smooth score field with arbitrarily small forward-marginal L 2 L^2 L 2 error, showing nonexplosive behavior and moments of every order. result Euler--Maruyama discretizations can converge in probability even when moments diverge, demonstrating failure of weak convergence.
This paper analyzes speculative decoding, a method to speed up large language model inferences.
problem Theoretical understanding of speculative decoding is lacking.
method Conceptualizes speculative decoding as a markov chain problem and studies its key properties.
result Reveals fundamental connections between LLM components and their impact on decoding efficiency.
Flow matching KL divergence bound derived for smooth distributions.
problem Estimating smooth distributions efficiently.
method Deterministic upper bound on KL divergence derived from flow-matching loss.
result Flow matching achieves nearly minimax-optimal efficiency under TV distance.
The Minimum Description Length (MDL) principle selects the model that has the shortest code for data plus model. We show that for a countable class of models, MDL predictions are close to the true distribution in a strong sense. The result is completely general. No independence, ergodicity, stationarity, identifiabilit…
Unified analysis of MPLE for Ising models with bounded operator norm or infinity norm.
problem Estimating Ising models in Total Variation distance with limited samples.
method Maximum Pseudo-Likelihood Estimator (MPLE) for two general classes of Ising models.
result Unified framework for polynomial-time estimation in TV distance for two general classes of Ising models.
HMC improves Gaussian sampling efficiency with long, random steps.
problem Efficiently sampling from high-dimensional Gaussian distributions.
method Hamiltonian Monte Carlo with long and random integration times.
result HMC achieves ε \varepsilon ε -closeness in total variation distance with O ~ ( κ d 1 / 4 log ( 1 / ε ) ) \widetilde{O}(\sqrt{\kappa} d^{1/4} \log(1/\varepsilon)) O ( κ d 1/4 log ( 1/ ε )) gradient queries.