New method relaxes TV distance for two-sample testing without distributional assumptions.
problem Challenges in certifying equality or providing tight bounds on TV distance for two distributions.
method Examined blurred total variation distance, a relaxation of TV distance.
result Provided theoretical guarantees for upper and lower bounds on blurred TV distance.
A new metric HCP distance for comparing distributions.
problem Comparing high-dimensional probability distributions efficiently.
method Hilbert curve projection to low-dimensional coupling, followed by transport distance calculation.
result HCP distance is a proper metric for probability measures with bounded supports.
Robust test for distributions under Hellinger distance, simpler than optimal tests.
problem Testing and estimating distributions robustly under Hellinger distance.
method Simple robust hypothesis test with optimal sample complexity, robust to Hellinger distance perturbations.
result Empirically demonstrated robustness and power of the test on canonical distributions.
New distances for comparing multivariate normal distributions.
problem Comparing multivariate normal distributions efficiently and accurately.
method Approximated Fisher-Rao distance and pullback SPD cone distances.
result Efficient computation of distances between normal distributions.
Paper proposes a new Wasserstein distance for mixtures of radially contoured distributions.
problem Generalization of Wasserstein distance to non-elliptically contoured distributions.
method Relaxed formulation for mixtures of radially contoured distributions without marginal consistency.
result The new distance yields more stable error and better color distribution in image transfer tasks.
Exact 1-Wasserstein distance between location-scale distributions derived, with privacy effects studied.
problem Calculating the 1-Wasserstein distance between location-scale distributions and its impact on differential privacy.
method Exact expressions and special functions for 1-Wasserstein distance, new upper bounds, and asymptotic analysis.
result New linear upper bound and detailed asymptotic bounds for Gaussian case, effect of differential privacy studied.
Survey on closed-form Fisher-Rao distance expressions.
problem Finding closed-form expressions for Fisher-Rao distance.
method Collect and present examples of closed-form expressions for Fisher-Rao distance of discrete and continuous distributions.
result Presentation of closed-form expressions for Fisher-Rao distance of various distributions.
Deep metric learning employs deep neural networks to embed instances into a metric space such that distances between instances of the same class are small and distances between instances from different classes are large. In most existing deep metric learning techniques, the embedding of an instance is given by a featur…
This note shows how independent elliptical distributions minimize the Wasserstein distance.
problem Minimizing the Wasserstein distance between elliptical distributions.
method Analyzing the Wasserstein distance between independent elliptical distributions with the same density generators.
result Independent elliptical distributions minimize their Wasserstein distance from other elliptical distributions with the same density generators.
A new method embeds distributions in a common space for optimal transport comparison.
problem Comparing distributions in different metric spaces.
method Sub-embedding robust Wasserstein (SERW) distance.
result SERW mimics GW distance properties and provides a cost relation.
Energy distance measures feature heterogeneity in federated learning.
problem Heterogeneity across data sources hinders model aggregation in federated learning.
method Introduced Taylor approximations of energy distance for efficient computation.
result Taylor approximations accurately capture feature discrepancies, improving convergence.
A method for fast estimation of Wasserstein distances using sliced Wasserstein distances.
problem Efficiently computing Wasserstein distances for multiple pairs of distributions.
method Regression on sliced Wasserstein distances to predict true Wasserstein distances.
result The proposed method provides a better approximation of Wasserstein distance than state-of-the-art models, especially in low-data regimes.
A new Wasserstein distance method for comparing incomparable distributions.
problem Comparing distributions that are not supported on the same metric space.
method Distributional slicing, embeddings, and closed-form computation of Wasserstein distance.
result HWD preserves properties like rotation-invariance and can be efficiently learned.
This paper introduces a noise-robust clustering method using distribution distances.
problem Reducing noise impact on clustering results.
method Introduces expectation distance (ED) for distribution clustering, extending K-means and K-medoids.
result Improved clustering accuracy and reduced computation time.
Understanding proper distance measures between distributions is at the core of several learning tasks such as generative models, domain adaptation, clustering, etc. In this work, we focus on mixture distributions that arise naturally in several application domains where the data contains different sub-populations. For …
A new robust metric compares distributions more accurately than existing methods.
problem Sensitivity to outliers and sampling discrepancy in Wasserstein distances.
method Introducing k-RPW, a partial p-Wasserstein distance.
result k-RPW converges faster to true distance and is more robust to outliers.
Constructs portfolios based on Hellinger distance to normal, finding market invariance.
problem Finding a market invariant for portfolio construction.
method Uses Hellinger distance to normal distribution for portfolio construction and analysis.
result Minimum Hellinger distance varies drastically between markets, suggesting market invariance.
Proposes Mutual Regression Distance for better distribution comparison.
problem Lack of effective distance measures for manifold data.
method Constrained mutual regression problem to exploit manifold properties.
result Mutual Regression Distance (MRD) is a pseudometric effective for manifold data.
Correctly estimating the discrepancy between two data distributions has always been an important task in Machine Learning. Recently, Cuturi proposed the Sinkhorn distance which makes use of an approximate Optimal Transport cost between two distributions as a distance to describe distribution discrepancy. Although it ha…
Paper introduces CWDAE for better synthetic data generation.
problem Measuring discrepancy between generative and ground-truth distributions.
method Introduces mixture Cramer-Wold distance for joint and marginal distributional learning.
result CWDAE shows remarkable performance in generating synthetic data.
Paper proposes a method to estimate total variation distance for synthetic data fidelity.
problem Assessing the fidelity of synthetic data generated by AI.
method Discriminative approach to estimate total variation distance between two distributions.
result Estimation of total variation distance reduces to quantifying Bayes risk in classification.
This paper studies clustering of data sequences using the k-medoids algorithm. All the data sequences are assumed to be generated from \emph{unknown} continuous distributions, which form clusters with each cluster containing a composite set of closely located distributions (based on a certain distance metric between di…
We provide a unifying framework linking two classes of statistics used in two-sample and independence testing: on the one hand, the energy distances and distance covariances from the statistics literature; on the other, distances between embeddings of distributions to reproducing kernel Hilbert spaces (RKHS), as establ…
Proposes an energy-based sliced Wasserstein distance for improved probability measure comparison.
problem Inefficiencies and limitations in existing sliced Wasserstein distance approaches.
method Introduces an energy-based slicing distribution for better performance and stability.
result Demonstrates superior performance of the EBSW distance in various applications.
Flexible classifier using Mahalanobis distances for non-elliptical distributions.
problem Classifying non-elliptical and multimodal distributions.
method Semiparametric classifier based on Mahalanobis distances and generalized additive models.
result The proposed classifiers outperform traditional methods in high-dimensional, low-sample-size scenarios.
New distances measure mixtures of Gaussians, useful in machine learning.
problem Comparing distributions with disjoint supports.
method Schoenberg-Rao distances based on concave Rao's entropy.
result Closed-form distances for mixtures of Gaussians.
Upper bound for max-sliced 2-Wasserstein distance between measures.
problem Estimating distance between probability measures and their empirical counterparts.
method Same technique as previous work, upper bound approach.
result Upper bound for expected max-sliced 2-Wasserstein distance.
Sharp bounds for max-sliced Wasserstein distances derived for empirical distributions.
problem Estimating the expected max-sliced Wasserstein distance between a probability measure and its empirical distribution.
method Banach space version and operator norm approach for upper bounds.
result Upper bounds for max-sliced Wasserstein distances are essentially matching and sharp up to a log factor.
This work connects Cramér distance to QR-DQN for DRL.
problem Improving performance in DRL by capturing full distribution of returns.
method Proves Cramér distance's equivalence to 1-Wasserstein distance and proposes a low-complexity algorithm to compute Cramér distance.
result Cramér distance and quantile regression losses yield collinear gradients under non-crossing constraints.
A new distance metric compares probability distributions using kernel covariance operators.
problem Comparing probability distributions in machine learning tasks.
method Introduces a novel distance metric based on Schatten norm of kernel covariance operators.
result The new distance metric is more discriminative and robust to hyperparameters.
The paper studies how different entropic regularizations affect GAN solutions.
problem Improving numerical convergence and sparsity in GAN solutions.
method Entropic regularization of Wasserstein distance and Sinkhorn divergence.
result Entropy regularization promotes sparsity, while Sinkhorn divergence recovers unregularized solution.
New algorithm samples from log-concave distributions with high accuracy in polynomial time.
problem Sampling from log-concave distributions with high accuracy in infinity distance.
method Directly converts continuous samples from K with total-variation bounds to samples with infinity bounds. result Output a point ε-close to π in infinity distance with runtime bounds that depend on polylogarithmic and polynomial factors of 1/ε. Study on Wasserstein distance for numerical approximations of stochastic differential equations.
problem Estimating the Wasserstein distance between stochastic differential equation distributions and their numerical approximations.
method Unified framework for analyzing different integrators and a novel splitting method for underdamped Langevin dynamics.
result A novel splitting method for underdamped Langevin dynamics with optimal complexity.
This paper presents a distance-based discriminative framework for learning with probability distributions. Instead of using kernel mean embeddings or generalized radial basis kernels, we introduce embeddings based on dissimilarity of distributions to some reference distributions denoted as templates. Our framework exte…
Study robust hypothesis testing under Hellinger distance, proving lower bounds and providing tests.
problem Testing close variants of specified distributions robustly to Hellinger distance.
method Lower bound on slack factor, testing with Hellinger balls, symmetric chi-squared distance analysis.
result Lower bound on slack factor quantifies robustness under misspecification.
Optimizes distributions robustly with Sinkhorn distance.
problem Distributionally robust optimization with Wasserstein distance.
method Convex programming dual reformulation, stochastic mirror descent algorithm.
result Demonstrates superior performance in synthetic and real data.
The scaled complex Wishart distribution is a widely used model for multilook full polarimetric SAR data whose adequacy has been attested in the literature. Classification, segmentation, and image analysis techniques which depend on this model have been devised, and many of them employ some type of dissimilarity measure…
Paper introduces S3W distance for spherical probability distributions.
problem Comparing spherical probability distributions efficiently and accurately.
method S3W distance using stereographic projection and generalized Radon transform.
result Extensive theoretical analysis and evaluation of S3W performance.
In this letter, we derive the optimal discriminant functions for modulation classification based on the sampled distribution distance. The proposed method classifies various candidate constellations using a low complexity approach based on the distribution distance at specific testpoints along the cumulative distributi…
Optimal pre-processing reduces disparate impact by minimizing total variation distance.
problem Achieving fairness in data outputs based on protected attributes.
method Using pre-processing to enforce fairness, minimizing total variation distance between pre-processed and original data distributions.
result The problem of fairness can be formulated as a linear program, efficiently solvable.
A new copula minimizes distance between distributions.
problem Arbitrariness in copula choice.
method Minimizes Wasserstein distance; linear programming estimation.
result Natural copula provides parsimonious estimation.
New PAC-Bayesian bounds improve Sliced-Wasserstein distances.
problem Improving statistical properties of Sliced-Wasserstein distances.
method Leveraging PAC-Bayesian theory to provide bounds and learning procedures.
result PAC-Bayesian generalization bounds for adaptive SW distances.
New smoothing technique improves Wasserstein distance estimation in high dimensions.
problem Estimating statistical distances between high-dimensional distributions.
method Gaussian smoothing of p-Wasserstein distance and analysis of its asymptotic behavior. result Gaussian-smoothed p-Wasserstein distance converges at rate n−1/2, improving over n−1/d for unsmoothed distances. CADM proposes a cluster-specific distance metric for categorical data clustering.
problem Inadequate distance metrics for categorical data, especially varying within clusters.
method Cluster-customized adaptive distance metric for categorical data.
result Achieved competitive performance in categorical data clustering.
DCMA uses generative models to analyze complex treatment effects on outcome distributions.
problem Analyzing complex and nonlinear causal mechanisms through outcome-level summary contrasts.
method Generative learning framework for identifying and estimating treatment effects on entire outcome distributions.
result Reconstructs interventional outcome distributions via Monte Carlo forward simulation, capturing both summary and distributional contrasts.
Distributional (or distribution-valued) data are a new type of data arising from several sources and are considered as realizations of distributional variables. A new set of fuzzy c-means algorithms for data described by distributional variables is proposed. The algorithms use the L2 Wasserstein distance between dist…
Proposes a Coulomb-like model for international trade flows, fitting real-world data.
problem Describing and predicting international trade flows between countries.
method Formulated a coulomb force model where GDP represents charge and distance is influenced by various factors.
result Developed a trade strength distribution equation that fits real-world data well.
Diffusion models achieve nearly optimal distribution estimation in various spaces.
problem Theoretical limitations of diffusion modeling for distribution estimation.
method Analysis of approximation and generalization abilities of diffusion models in Besov spaces.
result Diffusion models achieve nearly minimax optimal estimation rates in total variation and Wasserstein distances.