Study Gaussian approximation for deep neural networks with random weights.
problem Understanding the distribution of deep neural networks with random weights.
method Established Gaussian approximation bounds in Wasserstein-1 norm.
result Convergence rates of order n−(1/6)L−1+ε for deep networks with proportional layer widths. A new discrete formula connects vertex and edge distributions on graphs.
problem Optimal transport on graphs with mixed vertex and edge distributions.
method Discrete transport equation and Benamou-Brenier formulation.
result Classification of all Wasserstein-1 geodesics on graphs.
Generative models learn complex data from low-dimensional manifolds.
problem Theoretical justification for generative models on manifold structures.
method Prove statistical guarantees of generative networks under Wasserstein-1 loss, considering intrinsic dimensionality.
result Generative networks converge to zero at a fast rate depending on intrinsic dimensionality, not ambient data dimension.
The paper introduces a new ODE approach to improve Wasserstein GANs.
problem Improving Wasserstein GANs for better training results.
method Derives an ODE representing the gradient flow of Wasserstein-1 loss and proposes a new model W1-FE.
result W1-FE outperforms WGAN in training experiments across various dimensions.
The study derives generalization bounds for neural oscillators, improving their performance with regularization.
problem Quantifying the generalization capacities of neural oscillators.
method Using Rademacher complexity and squared Wasserstein-1 distances, the study derives theoretical upper PAC generalization bounds for neural oscillators.
result Theoretical bounds show polynomial growth in estimation errors with MLP size and time length, and regularization improves performance.
Study error bounds in evaluating distributional computational graphs.
problem Error analysis in evaluating graphs with inputs as probability distributions.
method Establish non-asymptotic error bounds using Wasserstein-1 distance.
result Non-asymptotic error bounds for discretization errors in distributional computational graphs.
Wasserstein Generative Adversarial Networks (WGANs) provide a versatile class of models, which have attracted great attention in various applications. However, this framework has two main drawbacks: (i) Wasserstein-1 (or Earth-Mover) distance is restrictive such that WGANs cannot always fit data geometry well; (ii) It …
Study the tradeoff between signal distortion and human perception over finite channels.
problem Characterize the distortion-perception tradeoff for finite channels with arbitrary metrics.
method Solve linear programming problems to compute the distortion-perception function and optimal reconstructions.
result DP function is piecewise linear in the perception index.
Flow Matching improves Wasserstein 1 distance convergence in high dimensions.
problem Improving Wasserstein 1 distance estimation for unbounded distributions.
method Flow Matching approach based on ODEs, controlling Lipschitz constant.
result Derives a convergence rate for Wasserstein 1 distance, improving previous results.
Neural network models accurately price assets in rough Bergomi model.
problem Accurately pricing assets in the rough Bergomi model with hidden parameters.
method Used a neural SDE to learn the forward variance curve, proposing a numerical scheme for simulation.
result The learned forward variance curve calibrates asset prices and option prices simultaneously.
Generative flows learn distributions on low-dimensional manifolds robustly via Wasserstein proximals.
problem Learning distributions supported on low-dimensional manifolds robustly.
method Combining Wasserstein-1 and Wasserstein-2 proximal operators to formulate well-posed continuous-time generative flows.
result The combination of Wasserstein-1 and Wasserstein-2 proximals ensures the well-posedness of generative flows, leading to unique and robust learning.
The paper improves GANs' theoretical guarantees for low-dimensional data.
problem Theoretical guarantees for GANs' statistical accuracy remain pessimistic.
method Analytical derivation of statistical guarantees on estimated densities.
result Theoretical rates of convergence for GANs and BiGANs are derived.
Reduced sample complexity for group-invariant distributions.
problem Improving sample complexity for estimating divergences of group-invariant distributions.
method Quantified reduction in sample complexity for Wasserstein-1 metric and Lipschitz-regularized α-divergences under finite and infinite groups.
result Sample complexity reduction proportional to group size for finite groups, and convergence rate depends on intrinsic dimension for infinite groups.
Quantum Earth Mover's distance improves stability and efficiency in quantum learning.
problem Quantum learning's loss landscapes often lead to poor local minima and gradients.
method Introduced the quantum Earth Mover's (EM) distance and proposed a quantum Wasserstein generative adversarial network (qWGAN).
result The quantum EM distance makes quantum learning more stable and efficient.
Wasserstein GANs with Gradient Penalty compute a different optimal transport problem called congested transport.
problem Training generative models to produce high-quality synthetic data.
method Wasserstein GANs with Gradient Penalty (WGAN-GP) approach to calculate the Wasserstein 1 distance.
result WGAN-GP computes the minimum of the congested transport problem, not the Wasserstein 1 distance.
Neural SDEs improve time series generation efficiency.
problem High memory and computational costs in GANs for time series.
method Conditional Neural Stochastic Differential Equations (SDEs).
result More memory efficient and faster than traditional methods.
We propose an approach to fair classification that enforces independence between the classifier outputs and sensitive information by minimizing Wasserstein-1 distances. The approach has desirable theoretical properties and is robust to specific choices of the threshold used to obtain class predictions from model output…
We introduce a new approximation of f-divergences for machine learning.
problem Variational representations of f-divergences for machine learning. method Definition and analysis of Moreau-Yosida approximation of f-divergences with the Wasserstein-1 metric. result Generalization and relaxation of hard Lipschitz constraints in f-divergences. Paper proposes a new framework to improve policy optimization by aligning real and simulated data distributions.
problem Inaccurate model estimation leads to performance degradation in model-based reinforcement learning.
method Introduces unsupervised model adaptation to minimize the IPM between real and simulated data distributions.
result Achieves state-of-the-art performance in sample efficiency on various continuous control tasks.
We study the minimax optimal rate for estimating the Wasserstein-1 metric between two unknown probability measures based on n i.i.d. empirical samples from them. We show that estimating the Wasserstein metric itself between probability measures, is not significantly easier than estimating the probability measures u…
New method calculates Ricci curvature from distances between weighted volumes.
problem Calculating Ricci curvature for weighted Riemannian manifolds.
method Asymptotic retrieval of generalized Ricci tensor from scaled metric derivatives of Wasserstein 1-distances.
result Limiting coarse curvature of random graphs converges to generalized Ricci tensor.
Paper improves MMD flow efficiency with Riesz kernels for image generation.
problem High computational costs in MMD flows for large scale computations.
method Introduces Riesz kernels and sliced MMD for efficient computation.
result Efficient computation of MMD gradients in one-dimensional setting.
New methods stabilize EEG classification performance across subjects.
problem Large performance drop in EEG classification models on unseen subjects.
method Regularization techniques using divergence estimation.
result Significant increase in balanced accuracy on test subjects.
Efficiently simulates and calibrates the rough Bergomi model using Wasserstein distance.
problem High computational complexity in pricing and calibration of the rough Bergomi model.
method Developed a modified-sum-of-exponentials Monte Carlo scheme and a calibration approach based on Wasserstein-1 distance.
result The method achieves high pricing accuracy and improved parameter recovery, optimization stability, and out-of-sample performance.
WAPPO optimizes feature distributions for better visual transfer in RL.
problem Improving visual transfer in reinforcement learning.
method WAPPO uses Wasserstein Confusion to minimize feature distribution distance.
result WAPPO outperforms previous methods in visual transfer across different environments.
New method improves sampling efficiency in complex stochastic systems.
problem Sampling efficiency in nonconvex stochastic gradient cases.
method Reflection coupling for unadjusted generalized Hamiltonian Monte Carlo.
result Quantitative Gaussian concentration bounds and convergence rates established.
In this paper, we are concerned with a non-asymptotic analysis of sampling algorithms used in nonconvex optimization. In particular, we obtain non-asymptotic estimates in Wasserstein-1 and Wasserstein-2 distances for a popular class of algorithms called Stochastic Gradient Langevin Dynamics (SGLD). In addition, the afo…
Novel coarse extrinsic curvature for Riemannian submanifolds.
problem Understanding extrinsic curvature of submanifolds.
method Derived from Wasserstein 1-distance between probability measures.
result New insights and approximation of mean curvature from data.
This paper proposes an efficient method for sampling from stochastic differential equations using PSD models.
problem Efficient sampling from stochastic differential equations with positive semi-definite models.
method The approach leverages a PSD model to sample from the Fokker-Planck equation or its fractional variant, with a complexity of m2dlog(1/ε). result The method produces i.i.d. samples with error ε in Wasserstein-1 distance, with a cost of O(dε−2(d+1)/β−2log(1/ε)2d+3) per sample. Generative Adversarial Networks (GANs) have achieved a great success in unsupervised learning. Despite its remarkable empirical performance, there are limited theoretical studies on the statistical properties of GANs. This paper provides approximation and statistical guarantees of GANs for the estimation of data distri…
Generative models improve inverse problems by providing tailored priors.
problem Analyzing the error in inverse problems solved with generative priors.
method Quantitative error bounds for minimum Wasserstein-2 generative models.
result The error in the posterior due to the generative prior is bounded by the prior's error in Wasserstein-1 distance.
Paper analyzes SGHMC for non-convex optimization with discontinuous gradients.
problem Training neural networks with ReLU activation.
method Non-asymptotic convergence analysis of SGHMC with discontinuous gradients.
result Explicit upper bounds for expected excess risk in non-convex optimization.
New algorithms improve sampling from complex distributions.
problem Sampling from high-dimensional target distributions with super-linearly growing potentials.
method Proposed aHOLA and aHOLLA algorithms with non-asymptotic convergence bounds.
result Achieved state-of-the-art rates of convergence in non-convex settings.
Robust VAE detects anomalies in corrupted data.
problem Detect anomalies in data with high corruption.
method Robust Variational Autoencoder (VAE) with four modifications.
result Establishes robustness to outliers and suitability to low-rank modeling.
LACD uses unlabeled data to improve conditional diffusion models.
problem Costly and time-consuming acquisition of labeled data.
method Label-augmented conditional diffusion (LACD) with joint denoising score matching.
result LACD converges faster in total variation and Wasserstein-1 distances with sufficient unlabeled data.
Abstract compares two norms in holomorphic quadratic differentials.
problem Comparing two norms in holomorphic quadratic differentials.
method Comparison between Avila-Gouëzel-Yoccoz norm and Teichmüller norm.
result Comparison of two norms in holomorphic quadratic differentials.
We introduce a new family of matrix norms, the "local max" norms, generalizing existing methods such as the max norm, the trace norm (nuclear norm), and the weighted or smoothed weighted trace norms, which have been extensively used in the literature as regularizers for matrix reconstruction problems. We show that this…
We introduce twisted Alexander norms of a compact connected orientable 3-manifold with first Betti number bigger than one generalizing norms of McMullen and Turaev. We show that twisted Alexander norms give lower bounds on the Thurston norm of a 3-manifold. Using these we completely determine the Thurston norm of many …
We study a regularizer which is defined as a parameterized infimum of quadratics, and which we call the box-norm. We show that the k-support norm, a regularizer proposed by [Argyriou et al, 2012] for sparse vector prediction problems, belongs to this family, and the box-norm can be generated as a perturbation of the fo…
New L0 norm added to TDA for market analysis.
problem Improving TDA tools for market prediction.
method Defined and applied L0 norm in TDA for four markets.
result Enhanced TDA tools for market analysis.
The k-support norm is a regularizer which has been successfully applied to sparse vector prediction problems. We show that it belongs to a general class of norms which can be formulated as a parameterized infimum over quadratics. We further extend the k-support norm to matrices, and we observe that it is a special …
CNN layers with large norms are still robust to adversarial attacks.
problem Understanding the relationship between layer norms and adversarial robustness in CNNs.
method Theoretical analysis of ℓ1 and ℓ∞ norms, norm decay method, adversarial training frameworks. result Adversarially robust CNNs can have comparable or larger layer norms than non-adversarially robust ones.
Study on minimal hypersurfaces in a special normed space.
problem Characterizing minimal hypersurfaces in a specific normed space.
method Investigate translation and separable minimal hypersurfaces.
result New insights into the properties of minimal hypersurfaces.
The paper defines minimal norm tensors for curvature and divergence tensors, explaining Weyl and Cotten tensors.
problem Understanding curvature tensors and their minimal norm.
method Analyzing minimal norm tensors for third and fourth covariant tensors, including Riemannian curvature and divergence.
result Weyl tensor and Cotten tensor are identified as minimal norm tensors of Riemannian curvature and divergence tensors, respectively.
A result of Bangert states that the stable norm associated to any Riemannian metric on the 2-torus T2 is strictly convex. We demonstrate that the space of stable norms associated to metrics on T2 forms a proper dense subset of the space of strictly convex norms on R2. In particular, given a strictly convex …
The spectral k-support norm enjoys good estimation properties in low rank matrix learning problems, empirically outperforming the trace norm. Its unit ball is the convex hull of rank k matrices with unit Frobenius norm. In this paper we generalize the norm to the spectral (k,p)-support norm, whose additional para…
The Schatten quasi-norm can be used to bridge the gap between the nuclear norm and rank function, and is the tighter approximation to matrix rank. However, most existing Schatten quasi-norm minimization (SQNM) algorithms, as well as for nuclear norm minimization, are too slow or even impractical for large-scale problem…
Optimization problems with rank constraints appear in many diverse fields such as control, machine learning and image analysis. Since the rank constraint is non-convex, these problems are often approximately solved via convex relaxations. Nuclear norm regularization is the prevailing convexifying technique for dealing …