Improved robustness for deep neural networks with tighter bounds and attacks.
problem Loose upper bounds and prohibitive computation in existing adversarial robustness methods.
method Primal approach with exact Lipschitz certificates for ReLU networks and modern architectures, and novel Wasserstein Distributional Attacks.
result Tighter upper bounds and greater flexibility in attack points compared to existing methods.
Our attacks are stronger and faster under Wasserstein metric.
problem Vulnerability of deep models to adversarial attacks.
method Developed an exact yet efficient projection operator and used the Frank-Wolfe method.
result Generated much stronger attacks and improved model robustness.
In the last couple of years, several adversarial attack methods based on different threat models have been proposed for the image classification problem. Most existing defenses consider additive threat models in which sample perturbations have bounded L_p norms. These defenses, however, can be vulnerable against advers…
We use distributionally-robust optimization for machine learning to mitigate the effect of data poisoning attacks. We provide performance guarantees for the trained model on the original data (not including the poison records) by training the model for the worst-case distribution on a neighbourhood around the empirical…
We improve image perturbation defenses using a better-defined Wasserstein threat model.
problem Real-world image perturbations are not pixel-independent, unlike ℓ p \ell_p ℓ p threat models. method We rectify flaws in the Wasserstein threat model and explore stronger attacks and defenses.
result Current Wasserstein-robust models are ineffective against real-world perturbations.
Study extends DRO with IPMs, linking robustness to regularization and GANs.
problem Addressing robustness of deep neural networks to adversarial attacks.
method Distributionally Robust Optimization (DRO) with Integral Probability Metrics (IPMs).
result DRO under any IPM corresponds to a family of regularization penalties.
WOOD detects out-of-distribution samples using Wasserstein distance.
problem Detecting samples from different distributions in neural networks.
method WOOD defines a Wasserstein-distance-based score to evaluate dissimilarity and solves an optimization problem.
result WOOD consistently outperforms other OOD detection methods.
DRO-Augment framework enhances deep neural network robustness.
problem Robustness of deep neural networks against various perturbations and adversarial attacks.
method Integrates Wasserstein Distributionally Robust Optimization with data augmentation.
result Significantly improves robustness across various corruptions and adversarial attacks.
Efficiently learns distributions corrupted by both global and local adversarial modifications.
problem Learning distributions with both global and local adversarial corruptions.
method Develops an efficient algorithm to minimize Wasserstein distance with orthogonal projections.
result Achieves optimal risk bounds with error ε k + ρ + i l d e O ( d k n − 1 / ( k ∨ 2 ) ) \sqrt{\varepsilon k} + ρ+ ilde{O}(d\sqrt{k}n^{-1/(k \lor 2)}) ε k + ρ + i l d e O ( d k n − 1/ ( k ∨ 2 ) ) . A rapidly growing area of work has studied the existence of adversarial examples, datapoints which have been perturbed to fool a classifier, but the vast majority of these works have focused primarily on threat models defined by ℓ p \ell_p ℓ p norm-bounded perturbations. In this paper, we propose a new threat model for adver…
GroupSort neural networks can approximate Lipschitz continuous functions.
problem Understanding and improving the expressive power of neural networks with Lipschitz constraints.
method Introduced and studied GroupSort neural networks with constraints on weights, proving their ability to approximate Lipschitz continuous functions.
result GroupSort networks can represent any Lipschitz continuous piecewise linear functions and are well-suited for approximating general Lipschitz continuous functions.
Paper defends sensitive attributes in GNNs from inference attacks.
problem Protecting sensitive attributes in GNNs from inference attacks.
method Proposes adversarial training with TV and Wasserstein distance to locally filter sensitive attributes.
result Framework creates strong defense against inference attacks with minimal performance loss.
Paper tackles adapting multiple domains to a target domain using distillation and dictionary learning.
problem Adapting multiple heterogeneous labeled source domains to an unlabeled target domain.
method Combines Multi-Source Domain Adaptation and Dataset Distillation with Dataset Dictionary Learning.
result Achieves state-of-the-art adaptation performance even with minimal labeled data.
Neural networks are vulnerable to adversarial examples and researchers have proposed many heuristic attack and defense mechanisms. We address this problem through the principled lens of distributionally robust optimization, which guarantees performance under adversarial input perturbations. By considering a Lagrangian …
Develops a robust multiclass classification method for deep image classifiers.
problem Tackles data contamination and robustness to outliers in deep image classifiers.
method Uses Distributionally Robust Optimization (DRO) with Wasserstein metric ambiguity sets and regularized learning.
result Reduces test error rate by up to 83.5% and loss by up to 91.3% in image classification tasks.
Study of Gaussian distributions using entropic Gromov-Wasserstein and inner product Gromov-Wasserstein.
problem Optimal transportation between Gaussian distributions with different dimensions.
method Entropic Gromov-Wasserstein and inner product Gromov-Wasserstein, with closed-form expressions and von Neumann's trace inequality.
result Closed-form expressions for the entropic IGW and its unbalanced variant between Gaussian distributions.
Exact 1-Wasserstein distance between location-scale distributions derived, with privacy effects studied.
problem Calculating the 1-Wasserstein distance between location-scale distributions and its impact on differential privacy.
method Exact expressions and special functions for 1-Wasserstein distance, new upper bounds, and asymptotic analysis.
result New linear upper bound and detailed asymptotic bounds for Gaussian case, effect of differential privacy studied.
A method for fast estimation of Wasserstein distances using sliced Wasserstein distances.
problem Efficiently computing Wasserstein distances for multiple pairs of distributions.
method Regression on sliced Wasserstein distances to predict true Wasserstein distances.
result The proposed method provides a better approximation of Wasserstein distance than state-of-the-art models, especially in low-data regimes.
Generates samples conditioned on labels using optimal transport.
problem Estimating conditional distributions for specific labels.
method Wasserstein geodesic generator based on optimal transport theory.
result Learned conditional distributions and optimal transport maps.
Wasserstein gradient boosting predicts probability distributions for supervised learning.
problem Distribution-valued supervised learning where outputs are probability distributions.
method Fits a new weak learner to Wasserstein gradients of loss functionals of probability distributions.
result Superior performance in probabilistic prediction compared to existing methods.
This note shows how independent elliptical distributions minimize the Wasserstein distance.
problem Minimizing the Wasserstein distance between elliptical distributions.
method Analyzing the Wasserstein distance between independent elliptical distributions with the same density generators.
result Independent elliptical distributions minimize their Wasserstein distance from other elliptical distributions with the same density generators.
GeoECG augments ECG data to improve heart disease detection.
problem Insufficient labeled ECG data and vulnerability to adversarial attacks.
method Wasserstein geodesic perturbation for data augmentation.
result Improved accuracy and robustness in ECG-based heart disease detection.
A new Wasserstein distance method for comparing incomparable distributions.
problem Comparing distributions that are not supported on the same metric space.
method Distributional slicing, embeddings, and closed-form computation of Wasserstein distance.
result HWD preserves properties like rotation-invariance and can be efficiently learned.
Understanding proper distance measures between distributions is at the core of several learning tasks such as generative models, domain adaptation, clustering, etc. In this work, we focus on mixture distributions that arise naturally in several application domains where the data contains different sub-populations. For …
Paper proposes a new Wasserstein distance for mixtures of radially contoured distributions.
problem Generalization of Wasserstein distance to non-elliptically contoured distributions.
method Relaxed formulation for mixtures of radially contoured distributions without marginal consistency.
result The new distance yields more stable error and better color distribution in image transfer tasks.
This paper extends exponential smoothing to distributional time series using Wasserstein distance.
problem Forecasting distributional time series with exponential smoothing.
method Generalized exponential smoothing in Wasserstein space, with consistent parameter estimation.
result Wasserstein exponential smoothing outperforms traditional methods in high-frequency financial and electricity demand data.
A new robust metric compares distributions more accurately than existing methods.
problem Sensitivity to outliers and sampling discrepancy in Wasserstein distances.
method Introducing k-RPW, a partial p-Wasserstein distance.
result k-RPW converges faster to true distance and is more robust to outliers.
Wasserstein Dropout improves uncertainty estimation in neural networks.
problem Estimating neural uncertainties for safe machine learning.
method A purely non-parametric approach using dropout-based sub-network distributions and Wasserstein distance.
result Wasserstein Dropout outperforms state-of-the-art methods in uncertainty estimation.
Upper bound for max-sliced 2-Wasserstein distance between measures.
problem Estimating distance between probability measures and their empirical counterparts.
method Same technique as previous work, upper bound approach.
result Upper bound for expected max-sliced 2-Wasserstein distance.
In this paper we study generative modeling via autoencoders while using the elegant geometric properties of the optimal transport (OT) problem and the Wasserstein distances. We introduce Sliced-Wasserstein Autoencoders (SWAE), which are generative models that enable one to shape the distribution of the latent space int…
New KL-divergence for Gaussian distributions based on Wasserstein geometry.
problem Computing KL-divergence for Gaussian distributions efficiently.
method Introducing WKL-divergence based on Wasserstein geometry.
result WKL-divergence evaluates to squared distance between points for Dirac measures.
Neural Local Wasserstein Regression models distribution-on-distribution regression with flexible, localized transport maps.
problem Estimating distribution-on-distribution regression with global optimal transport maps or linearization limitations.
method Proposes Neural Local Wasserstein Regression, a flexible nonparametric framework using locally defined transport maps in Wasserstein space.
result Demonstrates effective capture of nonlinear and high-dimensional distributional relationships.
New adversarial attack method based on deep feature distributions.
problem Adversarial attacks on CNN classifiers using output layer information.
method Modeling and exploiting class-wise and layer-wise deep feature distributions.
result Achieves state-of-the-art transfer-based attack results for undefended ImageNet models.
Develops a method to efficiently compute Wasserstein barycenters with variational distributions.
problem High computational burden in computing Wasserstein barycenters for high-dimensional and continuous settings.
method Introduces a variational distribution to approximate the continuous Wasserstein barycenter, reformulating the problem as an optimization with c-cyclical monotonicity.
result The method provides a tractable dual formulation for efficient computation of Wasserstein barycenters, demonstrated on real applications.
WAPPO optimizes feature distributions for better visual transfer in RL.
problem Improving visual transfer in reinforcement learning.
method WAPPO uses Wasserstein Confusion to minimize feature distribution distance.
result WAPPO outperforms previous methods in visual transfer across different environments.
Sharp bounds for max-sliced Wasserstein distances derived for empirical distributions.
problem Estimating the expected max-sliced Wasserstein distance between a probability measure and its empirical distribution.
method Banach space version and operator norm approach for upper bounds.
result Upper bounds for max-sliced Wasserstein distances are essentially matching and sharp up to a log factor.
Adversarial examples are a hot topic due to their abilities to fool a classifier's prediction. There are two strategies to create such examples, one uses the attacked classifier's gradients, while the other only requires access to the clas-sifier's prediction. This is particularly appealing when the classifier is not f…
New PAC-Bayesian bounds improve Sliced-Wasserstein distances.
problem Improving statistical properties of Sliced-Wasserstein distances.
method Leveraging PAC-Bayesian theory to provide bounds and learning procedures.
result PAC-Bayesian generalization bounds for adaptive SW distances.
A new slicing method speeds up sliced Wasserstein estimation.
problem Efficiently estimating sliced Wasserstein distance.
method Random-Path Projecting Direction (RPD) for fast sampling.
result RPSW and IWRPSW show favorable performance in training generative models.
Paper proposes SinkhornDRL for distributional RL using Sinkhorn divergence and regularized Wasserstein loss.
problem Improving distributional reinforcement learning by minimizing Bellman return distribution differences.
method Introduces SinkhornDRL, a distributional RL algorithm using Sinkhorn divergence and regularized Wasserstein loss.
result SinkhornDRL consistently outperforms or matches existing algorithms on Atari games, especially in multi-dimensional reward settings.
Paper develops a new method for differential privacy sampling using Wasserstein distance.
problem Sampling from distributions under differential privacy constraints with geometric structure consideration.
method Develops a novel framework with Wasserstein Projection Mechanism (WPM) for minimax optimal mechanisms.
result Proposes efficient algorithms for approximate computation of the Wasserstein Projection Mechanism.
The Wasserstein metric is an important measure of distance between probability distributions, with applications in machine learning, statistics, probability theory, and data analysis. This paper provides upper and lower bounds on statistical minimax rates for the problem of estimating a probability distribution under W…
Proposes hinge-Wasserstein to improve uncertainty estimation in regression tasks.
problem Estimating multimodal aleatoric uncertainty in regression tasks from images.
method Regression-by-classification paradigm with hinge-Wasserstein loss.
result Hinge-Wasserstein loss improves uncertainty estimation on challenging tasks.
Study robust distribution estimation with Wasserstein distance, achieving optimal risk.
problem Robust distribution estimation under adversarial corruption.
method Combining partial OT and minimum distance estimation, proving structural properties and deriving a novel dual form.
result Achieves minimax-optimal robust estimation risk in many settings.
Modified Wasserstein metric for Gaussian distributions, invariant to isometries.
problem Distance measurement for latent Gaussian distributions invariant to isometries.
method Modified Benamou-Brenier approach leading to a Procrustes Wasserstein metric.
result For Gaussian distributions, the metric reduces to Euclidean distance between eigenvalues.
Wasserstein t-SNE embeds hierarchical datasets considering within-unit distributions.
problem Exploring hierarchical datasets where units are compared based on means of sample distributions.
method Uses Wasserstein distance metric for 2D embeddings of units, approximating Gaussian distributions for efficiency.
result Demonstrates effective embedding of hierarchical datasets, uncovering meaningful structure.
New method for reducing dimensions of distributional data.
problem Nonlinear sufficient dimension reduction for distribution-on-distribution regression.
method Building universal kernels on metric spaces to characterize conditional independence.
result Method outperforms competing methods in synthetic and real data applications.
A new portfolio model improves on Kelly's by accounting for estimation error.
problem Estimation error in Kelly portfolio optimization.
method Wasserstein distributionally robust optimization (DRO) to define a robust log-optimal portfolio.
result The Wasserstein-Kelly portfolio outperforms the Kelly portfolio in out-of-sample testing.