WGANs use optimal 1-Wasserstein distance to generate distributions.
problem Characterize geometrical properties of generated distributions.
method Analyze WGANs in finite and asymptotic regimes, focusing on univariate latent space.
result WGANs can approach target distribution with optimal 1-Wasserstein distance as sample size increases.
Exact 1-Wasserstein distance between location-scale distributions derived, with privacy effects studied.
problem Calculating the 1-Wasserstein distance between location-scale distributions and its impact on differential privacy.
method Exact expressions and special functions for 1-Wasserstein distance, new upper bounds, and asymptotic analysis.
result New linear upper bound and detailed asymptotic bounds for Gaussian case, effect of differential privacy studied.
This paper approximates 1-Wasserstein distance using tree-based embedding.
problem Computational inefficiency of estimating 1-Wasserstein distance.
method L1-regularized approach to learn tree weights, using shortest path distance as a linear model.
result Tree-Wasserstein distance (TWD) approximates 1-Wasserstein distance efficiently.
Model approximates continuous functions in 1-Wasserstein space.
problem Approximating continuous functions in 1-Wasserstein space.
method Probabilistic Transformer (PT) model with three phases: feature map, deep neural network, and probabilistic extension of attention mechanism.
result Can approximate any continuous function from R^d to P1(R^D) uniformly on compact sets.
Stability result for a popular algorithm in optimal transport.
problem Stability of the Iterative Proportional Fitting Procedure in time and metric.
method Uniform stability analysis in the 1-Wasserstein metric.
result Quantitative stability result for entropy-regularized Optimal Transport and Schrödinger bridges.
LSDM uses unpaired data to match latent space distributions for generative modeling.
problem Generating high-quality images with limited paired data.
method Two-stage approach: latent space learning from paired and unpaired data, followed by joint distribution matching.
result LSDM enhances geometric fidelity in generated outputs and provides theoretical insights into LDMs.
Paper analyzes statistical efficiency of TD learning in Hilbert spaces with Freedman's inequality.
problem Statistical efficiency of distributional TD learning in Hilbert spaces.
method Non-parametric distributional TD (NTD) and variance-reduced variants of NTD and CTD.
result Sharp statistical rates achieved through novel Freedman's inequality in Hilbert spaces.
pHMC converges on infinite-dimensional spaces with bounds.
problem Convergence of pHMC on Hilbert spaces.
method Coupling of two pHMC copies, adapted from arXiv:1805.00452.
result Proven convergence bounds in 1-Wasserstein distance.
Deep neural networks can approximate any target probability distribution given certain conditions.
problem Approximating complex probability distributions with deep neural networks.
method Proving the existence of a deep neural network mapping that approximates a target distribution under various integral probability metrics.
result Upper bounds on the size of the neural network in terms of dimension and approximation error for different metrics.
Learning algorithms for implicit generative models can optimize a variety of criteria that measure how the data distribution differs from the implicit model distribution, including the Wasserstein distance, the Energy distance, and the Maximum Mean Discrepancy criterion. A careful look at the geometries induced by thes…
Sharp bounds for max-sliced Wasserstein distances derived for empirical distributions.
problem Estimating the expected max-sliced Wasserstein distance between a probability measure and its empirical distribution.
method Banach space version and operator norm approach for upper bounds.
result Upper bounds for max-sliced Wasserstein distances are essentially matching and sharp up to a log factor.
This work studies the smooth 1-Wasserstein distance and its limit distribution in high dimensions.
problem Addressing the curse of dimensionality in empirical approximation.
method Conducts a statistical study including limit distribution, bootstrap consistency, and concentration inequalities.
result Derives a nondegenerate limit distribution for empirical SWD, contrasting with classic W 1 W_1 W 1 . MPE framework proves universal approximation for quantum data distribution.
problem Challenges in generating quantum data from underlying distributions.
method Many-body Projected Ensemble (MPE) framework for quantum state design.
result MPE can approximate any quantum distribution within 1-Wasserstein distance error.
This work connects Cramér distance to QR-DQN for DRL.
problem Improving performance in DRL by capturing full distribution of returns.
method Proves Cramér distance's equivalence to 1-Wasserstein distance and proposes a low-complexity algorithm to compute Cramér distance.
result Cramér distance and quantile regression losses yield collinear gradients under non-crossing constraints.
Error estimates found between SGD with momentum and Langevin diffusion.
problem Quantifying the difference between SGD with momentum and Langevin diffusion.
method Established error estimates using 1-Wasserstein and total variation distances.
result Quantitative error estimates between SGD with momentum and underdamped Langevin diffusion.
This work improves scalability of Wasserstein distances in high dimensions.
problem Scalability issues in computing Wasserstein distances in high dimensions.
method Empirical convergence rates, robustness to data contamination, and computational methods.
result Established fast rates and robust estimation risks for sliced Wasserstein distances.
Pathfinder uses quasi-Newton optimization for variational inference.
problem Approximating complex posterior distributions efficiently.
method Pathfinder combines quasi-Newton optimization with variational methods to approximate log densities.
result Pathfinder produces draws with lower KL divergence than ADVI and comparable to HMC, requiring fewer evaluations.
Diffusion models learn multi-modal distributions with optimal efficiency.
problem Learning high-dimensional distributions with low-dimensional multi-modal structures.
method Score-based diffusion models, focusing on subgaussian distributions within subspaces.
result Diffusion models require O ~ ( ε − k ∨ 2 ) \widetilde{O}(\varepsilon^{-k \vee 2}) O ( ε − k ∨ 2 ) samples for 1-Wasserstein ε \varepsilon ε error, improving over prior guarantees. Many attempts have been made in recent decades to integrate machine learning (ML) and topological data analysis. A prominent problem in applying persistent homology to ML tasks is finding a vector representation of a persistence diagram (PD), which is a summary diagram for representing topological features. From the pe…
We prove that for combinatorial graphs with non-negative Ollivier curvature, one has \[ \|P_t μ- P_t ν\|_1 \leq \frac{W_1(μ,ν)}{\sqrt{t}} \] for all probability measures μ , ν μ,ν μ , ν where P t P_t P t is the heat semigroup and W 1 W_1 W 1 is the ℓ 1 \ell_1 ℓ 1 -Wasserstein distance. This turns out to be an equivalent formulation of a version of…
We consider the problem of sampling from a target distribution, which is \emph {not necessarily logconcave}, in the context of empirical risk minimization and stochastic optimization as presented in Raginsky et al. (2017). Non-asymptotic analysis results are established in the L 1 L^1 L 1 -Wasserstein distance for the behavio…
Recent methods for generating novel molecules use graph representations of molecules and employ various forms of graph convolutional neural networks for inference. However, training requires solving an expensive graph isomorphism problem, which previous approaches do not address or solve only approximately. In this wor…
A new robust metric compares distributions more accurately than existing methods.
problem Sensitivity to outliers and sampling discrepancy in Wasserstein distances.
method Introducing k-RPW, a partial p-Wasserstein distance.
result k-RPW converges faster to true distance and is more robust to outliers.
Algorithm aligns 3D density maps using Wasserstein distance.
problem Aligning 3D density maps in cryogenic electron microscopy.
method Minimizing 1-Wasserstein distance after rigid transformation using Bayesian optimization.
result Improved accuracy and efficiency in protein molecule alignment.
Topological data analysis offers a rich source of valuable information to study vision problems. Yet, so far we lack a theoretically sound connection to popular kernel-based learning techniques, such as kernel SVMs or kernel PCA. In this work, we establish such a connection by designing a multi-scale kernel for persist…
Combinatorial approach to α α α -Ricci and Lin-Lu-Yau Ricci curvatures on graphs
problem Curvature formulas for α α α -Ricci and Lin-Lu-Yau Ricci curvatures on graphs method Combinatorial construction of optimal transport plans and exact formulas
result Combinatorial proof of known curvature formulas
This thesis explores Ollivier-Ricci curvature in graphs and manifolds, with applications to graph neural networks.
problem Understanding curvature in metric spaces and graphs.
method Combines optimal transport theory, Riemannian manifolds, and graph theory to define and analyze Ollivier-Ricci curvature.
result Extensions of Ollivier-Ricci curvature to directed graphs and applications in network science.
We study unsupervised generative modeling in terms of the optimal transport (OT) problem between true (but unknown) data distribution P X P_X P X and the latent variable model distribution P G P_G P G . We show that the OT problem can be equivalently written in terms of probabilistic encoders, which are constrained to match the pos…
We study the problem of sampling from a distribution p ∗ ( x ) ∝ exp ( − U ( x ) ) p^*(x) \propto \exp\left(-U(x)\right) p ∗ ( x ) ∝ exp ( − U ( x ) ) , where the function U U U is L L L -smooth everywhere and m m m -strongly convex outside a ball of radius R R R , but potentially nonconvex inside this ball. We study both overdamped and underdamped Langevin MCMC and establish upper bound…
This paper improves topic model estimation for sparse distributions and applies it to Wasserstein distances.
problem Estimating sparse topic distributions in topic models with high-dimensional data.
method MLE for topic weights when A A A is known, plug-in estimator for unknown A A A . result MLE can be exactly sparse and contain true zero pattern of topic weights.
The paper analyzes reflected diffusion models on hypercube data.
problem Challenges in modeling bounded domains with low-dimensional data.
method Employed an infinite series expansion of transition densities to bound the score function and its approximation.
result Established convergence rates for generative algorithm adapting to intrinsic dimensionality.
Persistent homology (PH) is a rigorous mathematical theory that provides a robust descriptor of data in the form of persistence diagrams (PDs) which are 2D multisets of points. Their variable size makes them, however, difficult to combine with typical machine learning workflows. In this paper we introduce persistence c…
The paper explores multidimensional critic output in GANs, improving convergence and diversity.
problem Underexplored in GANs literature, multidimensional critic output.
method Generalized Wasserstein GAN framework, SRVT block, maximal p-centrality discrepancy.
result High-dimensional critic output improves GAN performance in convergence and diversity.
STRAND: A single representation for hypothesis testing and vectorisation of persistence diagrams
problem Comparing persistence diagrams
method Survival topological representation analysis
result Non-parametric two-sample test with calibrated Type I error and high power
The paper tackles learning from non-irreducible Markov chains, proving learnability and generalization bounds.
problem Learning from temporal dependent data with non-irreducible Markov chains.
method Uniform convergence and generalization bounds for sample error under uniform ergodicity.
result Learnability and generalization bounds for approximate sample error minimization algorithm.
New error bounds for GANs with nonlinear objective functions derived.
problem Statistical consistency of GANs with nonlinear objective functions.
method Derivation of statistical error bounds for ( f , Γ ) (f,Γ) ( f , Γ ) -GANs using Rademacher complexity. result Proves the statistical consistency of ( f , Γ ) (f,Γ) ( f , Γ ) -GANs. Given a sample Y Y Y from an unknown manifold X X X embedded in Euclidean space, it is possible to recover the homology groups of X X X by building a Vietoris--Rips or Čech simplicial complex on top of the vertex set Y Y Y . However, these simplicial complexes need not inherit the metric structure of the manifold, in particular…
Particle-optimization-based sampling (POS) is a recently developed effective sampling technique that interactively updates a set of particles. A representative algorithm is the Stein variational gradient descent (SVGD). We prove, under certain conditions, SVGD experiences a theoretical pitfall, {\it i.e.}, particles te…
New method learns chaotic dynamics from single noisy trajectory.
problem Chaos in complex systems is hard to model accurately with machine learning.
method Adversarial optimal transport objectives to learn summary statistics and emulator from single noisy data.
result Emulators trained with proposed objectives have significantly improved long-term statistical fidelity.
Diverse projection ensembles improve distributional reinforcement learning.
problem Learning the distribution of returns in reinforcement learning.
method Combining multiple projection methods to improve model diversity and exploration.
result Diverse projection ensembles lead to significant performance improvements in exploration tasks.
Develops a new divergence framework that combines f f f -divergences and IPMs.
problem Comparing distributions that are not absolutely continuous.
method Introduces ( f , Γ ) (f,Γ) ( f , Γ ) -divergences as a two-stage mass-redistribution/mass-transport process. result Improves estimation, learning, and uncertainty quantification in GANs for heavy-tailed distributions.
New proof shows incremental flow models are essential for universal generation.
problem Understanding the universality of flow-based models in generating natural maps.
method Topological-dynamical argument and algebraic properties of flows.
result Incremental generation is necessary and sufficient for universal flow-based generation.
New method for high-dimensional linear regression using empirical Bayes.
problem Estimating prior in high-dimensional linear regression.
method Variational empirical Bayes approach with NPMLE and mean field approximation.
result Established asymptotic consistency and computational efficiency of the method.
Generative Adversial Networks (GANs) have made a major impact in computer vision and machine learning as generative models. Wasserstein GANs (WGANs) brought Optimal Transport (OT) theory into GANs, by minimizing the 1 1 1 -Wasserstein distance between model and data distributions as their objective function. Since then, W…
CDM models counterfactual outcomes in longitudinal data with improved accuracy.
problem Predicting counterfactual outcomes in longitudinal data with complex time-dependent confounding.
method Causal Diffusion Model (CDM) using denoising diffusion architecture with relational self-attention.
result CDM outperforms state-of-the-art methods in generating full probabilistic distributions of counterfactual outcomes.
Improved persistence spheres map measures to functions, stable under partial transport.
problem Representing and comparing measures in topological machine learning.
method Persistence spheres map measures to continuous functions on the sphere, stable under 1-Wasserstein partial transport.
result Persistence spheres provide a stable, parameter-free representation of measures, improving upon existing methods.
This study investigates self-supervised learning with Wasserstein distance on tree structures.
problem Improving self-supervised learning methods using Wasserstein distance.
method Utilized Tree-Wasserstein distance (TWD) and Jeffrey divergence regularization for training.
result A simple combination of softmax function and Tree-Wasserstein distance outperforms cosine similarity-based methods.
Personalized federated learning adapts models to each user's data.
problem Training models across users without data exchange leads to a common model, not personalized.
method Develops a personalized federated learning approach using Model-Agnostic Meta-Learning (MAML).
result Personalized models can be adapted by a few gradient descent steps on local data.