Ideas from the image processing literature have recently motivated a new set of clustering algorithms that rely on the concept of total variation. While these algorithms perform well for bi-partitioning tasks, their recursive extensions yield unimpressive results for multiclass clustering tasks. This paper presents a g…
Optimal pre-processing reduces disparate impact by minimizing total variation distance.
problem Achieving fairness in data outputs based on protected attributes.
method Using pre-processing to enforce fairness, minimizing total variation distance between pre-processed and original data distributions.
result The problem of fairness can be formulated as a linear program, efficiently solvable.
Total variation denoising improves image quality adaptively.
problem Improving image quality from noisy data.
method Total variation regularization for image denoising.
result Denoised images converge to true images at a parametric rate.
Total variation minimization clusters partially labeled data points.
problem Clustering partially labeled data points in stochastic block models.
method Total variation minimization as a clustering method.
result Total variation minimization allows for accurate clustering under certain model parameters.
The paper studies curves in Riemannian manifolds using total variation flow.
problem Analyzing the evolution of curves in Riemannian manifolds using total variation.
method Defining and proving the existence of strong solutions to the flow equations, showing variational equality, and proving convergence.
result Strong solutions converge to a constant map in finite time for non-positive sectional curvature.
The paper derives oracle inequalities for estimators with fast and slow rates.
problem Developing fast and slow oracle inequalities for estimators.
method Direct study of analysis estimator and adaptation of Dalalyan, Hebiri and Lederer's arguments.
result Constant-friendly rates for (square root) total variation regularized estimators over graphs.
The total variation distance is a core statistical distance between probability measures that satisfies the metric axioms, with value always falling in [0,1]. This distance plays a fundamental role in machine learning and signal processing: It is a member of the broader class of f-divergences, and it is related to …
We consider the problem of estimating a function defined over n locations on a d-dimensional grid (having all side lengths equal to n1/d). When the function is constrained to have discrete total variation bounded by Cn, we derive the minimax optimal (squared) ℓ2 estimation error rate, parametrized by …
Paper proposes a new method to minimize submodular functions with fewer calls to simpler oracles.
problem Minimizing the sum of submodular set functions with limited information.
method Introduces a modified convex problem requiring constrained total variation oracles that can be solved with fewer calls to minimization oracles.
result Shows significant reduction in the number of calls to minimization oracles.
SaR-SVM-STV improves hyperspectral image classification with shape-adaptive reconstruction and denoising.
problem Classifying hyperspectral images with limited labeled data.
method Shape-adaptive Reconstruction (SaR) for pixel preprocessing, SVM for probability estimation, and Smoothed Total Variation (STV) for denoising.
result SaR-SVM-STV outperforms SVM-STV with fewer labeled data.
The paper establishes prediction bounds for trend filtering with higher order total variation penalties.
problem Estimating signals with jumps of varying orders using total variation regularization.
method Combining oracle inequalities and interpolating vectors to bound effective sparsity.
result The ℓ1-penalty on (k−1)extth order differences allows adaptive estimation for k∈{1,2,3,4}. Estimates parameters of interconnected linear systems using total variation penalization.
problem Joint estimation of parameters in interconnected linear dynamical systems.
method Total variation penalized least-squares estimator.
result The MSE goes to zero as the number of systems increases, even with constant trajectory length.
Sharp inequality between TV and Hellinger distances for Gaussian mixtures.
problem Understanding the relationship between total variation and Hellinger distances for Gaussian mixtures.
method Established a general upper bound on Hellinger distance in terms of TV distance raised to a power, demonstrating sharpness with specific examples.
result The Hellinger distance between two Gaussian mixtures is bounded by the TV distance raised to a power 1−o(1), where o(1) is of order 1/loglog(1/TV). Paper proposes a method to estimate total variation distance for synthetic data fidelity.
problem Assessing the fidelity of synthetic data generated by AI.
method Discriminative approach to estimate total variation distance between two distributions.
result Estimation of total variation distance reduces to quantifying Bayes risk in classification.
Proximal algorithms applied to current deformation into cycles.
problem Deformation of de Rham currents into cycles.
method Proximal algorithms, total variation denoising for differential forms.
result Calibrated cycles constructed in calibrated manifolds.
Example shows learnable distributions not privately learnable.
problem Learnable distributions under non-private conditions not transferable to differential privacy.
method Example of a distribution class learnable up to constant error in total variation distance but not under differential privacy.
result Contradicts conjecture of Ashtiani on learnability under differential privacy.
The paper shows diffusion models can converge faster to a target distribution with low-dimensional structure.
problem Improving the convergence rate of diffusion models to target distributions.
method Analyzing DDIM and DDPM samplers under low-dimensional structure assumptions.
result The iteration complexities of DDIM and DDPM are no greater than k/ε in total variation distance. Graphs with non-negative Ollivier-Ricci curvature cannot be expanders.
problem Understanding the relationship between graph curvature and expansion properties.
method Proving an inequality linking isoperimetric profiles to total variation decay of random walks.
result Graphs with non-negative Ollivier-Ricci curvature cannot be expanders.
Total variation and mean curvature flows on a Lie group quotient enhance and denoise crossing structures.
problem Preserving crossing curvilinear structures in image enhancement and denoising.
method Lifting images to the homogeneous space M=RdtimesSd−1, applying PDEs for TVF and MCF, and using locally optimal differential frames. result Better preservation of bundle boundaries and angular sharpness in fiber orientation densities at crossings compared to data-driven diffusions.
New method relaxes TV distance for two-sample testing without distributional assumptions.
problem Challenges in certifying equality or providing tight bounds on TV distance for two distributions.
method Examined blurred total variation distance, a relaxation of TV distance.
result Provided theoretical guarantees for upper and lower bounds on blurred TV distance.
Error estimates found between SGD with momentum and Langevin diffusion.
problem Quantifying the difference between SGD with momentum and Langevin diffusion.
method Established error estimates using 1-Wasserstein and total variation distances.
result Quantitative error estimates between SGD with momentum and underdamped Langevin diffusion.
We focus on the maximum regularization parameter for anisotropic total-variation denoising. It corresponds to the minimum value of the regularization parameter above which the solution remains constant. While this value is well know for the Lasso, such a critical value has not been investigated in details for the total…
Hypergraphs allow one to encode higher-order relationships in data and are thus a very flexible modeling tool. Current learning methods are either based on approximations of the hypergraphs via graphs or on tensor methods which are only applicable under special conditions. In this paper, we present a new learning frame…
New bounds on neural network convergence using information theory.
problem Quantifying convergence rates of neural networks to Gaussian distributions.
method Entropic inequalities and Gaussian approximations.
result Improved convergence rates in various distances for neural networks.
Algorithm learns affine transformations robustly from corrupted samples.
problem Learning affine transformations from corrupted samples.
method New geometric certificate and iterative improvement method.
result Total variation distance of O(ε) between learned and original distributions. New TVD estimator adapts to piecewise constant functions, improving performance.
problem Improving TVD estimator performance for piecewise constant functions.
method Investigates adaptivity of TVD estimator to piecewise constant functions and proposes a data-driven tuning parameter.
result The ideally tuned TVD estimator performs better than in the worst case for piecewise constant functions.
Efficiently estimates binary product distributions with privacy.
problem Estimating means of binary product distributions privately and accurately.
method Polynomial time, pure differential privacy approach.
result Optimal sample complexity with polylogarithmic factors.
Proves sufficiency of countable test plans for BV functions on metric spaces.
problem Recovering BV functions and their measures on arbitrary metric spaces.
method Proves sufficiency of countable test plans on arbitrary metric measure spaces and geodesics on CD(K,N) spaces. result Countable test plans are sufficient for BV functions and their measures on metric spaces.
Robust Bayesian inference improves model performance on discrete data.
problem Misspecification of discrete-valued models leads to poor inference and prediction.
method Total Variation Distance (TVD) for discrepancy, efficient estimator and inference method.
result Our approach significantly improves predictive performance on various data.
In recent years, total variation (TV) and Euler's elastica (EE) have been successfully applied to image processing tasks such as denoising and inpainting. This paper investigates how to extend TV and EE to the supervised learning settings on high dimensional data. The supervised learning problem can be formulated as an…
Efficiently estimates densities of multidimensional shift-invariant distributions.
problem Density estimation for shift-invariant multidimensional distributions.
method Efficient algorithms for learning any distribution in the class from samples, using total variation distance.
result Shift-invariant distributions can be learned efficiently with a number of samples and time proportional to 1/εd+2 and 1/ε2d+2 respectively. While it is believed that denoising is not always necessary in many big data applications, we show in this paper that denoising is helpful in urban traffic analysis by applying the method of bounded total variation denoising to the urban road traffic prediction and clustering problem. We propose two easy-to-implement m…
Parallel neural networks estimate TVD for merging over-clustered datasets.
problem Merging over-partitioned clusters in unsupervised learning.
method Use neural networks to estimate TVD between clusters in parallel.
result Neural network estimates of TVD lead to better merge decisions.
We consider point clouds obtained as random samples of a measure on a Euclidean domain. A graph representing the point cloud is obtained by assigning weights to edges based on the distance between the points they connect. Our goal is to develop mathematical tools needed to study the consistency, as the number of availa…
Estimates TV distance between autoregressive models under different access models.
problem Estimating the total variation distance between two autoregressive distributions.
method Three access models: sample access, logit access, and noisy logit access; provides query complexity for each.
result Improved query complexity for estimating TV distance in autoregressive models.
Reinforcement learning mimics expert behavior.
problem Learning from expert demonstrations in reinforcement learning.
method Reduction to reinforcement learning with a stationary reward.
result Expert reward can be recovered and imitation learning is bounded.
New characterization limits sampling with inexact scores.
problem Limiting sampling with inexact scores for unbiased results.
method Characterized types of inexact score oracle access.
result Weaker error assumptions rule out tractability of unbiased sampling.
Method estimates noise transition matrix from noisy labels without relying on unreliable class-posterior estimation.
problem Estimating noise transition matrix from noisy data.
method Total variation regularization to encourage distinguishable predicted probabilities.
result Consistent estimator of the noise transition matrix under mild assumptions.
Polynomial-time algorithm forecasts TV-bounded sequences with optimal error rate.
problem Online forecasting of sequences with bounded total variation under noisy observations.
method Designing an O(nlogn)-time algorithm leveraging Haar wavelet basis and adaptivity. result Achieves optimal O(n1/3) cumulative square error with high probability. Based on a study of the coupling by reflection of diffusion processes, a new monotonicity in time of a time-dependent transportation cost between heat distribution is shown under Bakry-Emery's curvature-dimension condition on a Riemannian manifold. The cost function comes from the total variation between heat distribut…
Paper introduces a new method to model epidemic dynamics with varying parameters.
problem Capturing discontinuous variations in epidemic model parameters.
method Total variation regularization with Iterated Nelder--Mead optimization.
result The method accurately models epidemic dynamics with instant changes.
New schemes improve error estimates for sampling from non-log-concave distributions.
problem Improving sampling from non-log-concave distributions with super-linear drift growth.
method Developed tamed Euler and randomized Euler schemes with error estimates.
result Near-optimal error bounds for sampling and optimization problems.
The paper bounds information losses in neural classifiers from sampling.
problem Information losses in neural classifiers from finite datasets.
method Proves a relationship between information losses and expected total variation of the estimated neural model, bounds this expected total variation as a function of dataset size.
result Obtains bounds on information losses that are less sensitive to input compression and much smaller than existing bounds.
uHMC achieves fast mixing in high dimensions with gradient evaluations.
problem Quantifying mixing time of uHMC in high dimensions.
method Construction of successful couplings for uHMC.
result uHMC mixes in total variation with logarithmic dependence on dimension.
New method for tensor completion using nonconvex dual total variation.
problem Tensor completion from partial measurements with exponential-family noise.
method Proposed dual-TV (DTV) regularizers for tensor completion under exponential-family noise.
result Theoretical upper bounds on recovery error for tensor completion.
Data clustering is a fundamental problem with a wide range of applications. Standard methods, eg the k-means method, usually require solving a non-convex optimization problem. Recently, total variation based convex relaxation to the k-means model has emerged as an attractive alternative for data clustering. However…
We generalize to tree graphs obtained by connecting path graphs an oracle result obtained for the Fused Lasso over the path graph. Moreover we show that it is possible to substitute in the oracle inequality the minimum of the distances between jumps by their harmonic mean. In doing so we prove a lower bound on the comp…
New framework estimates staged tree models using hierarchical clustering on the probability simplex.
problem Estimating staged tree models with context-specific dependencies.
method Hierarchical clustering on the probability simplex, using simplex-based divergences and linkage methods.
result Total Variation divergence with Ward.D2 linkage produces staged trees with better model fit, structure recovery, and computational efficiency.