Wasserstein Autoencoders improve model efficiency and interpretability for low-dimensional data.
problem Limited statistical guarantees for WAEs in low-dimensional data.
method Proper network architecture selection and analysis of expected excess risk convergence rates.
result WAEs can learn data distributions efficiently when intrinsic dimension is considered.
New framework robustly handles outliers in Wasserstein DRO for better decision-making.
problem Non-geometric perturbations like adversarial outliers distort Wasserstein distance.
method Proposes an outlier-robust WDRO framework using a robust Wasserstein ball.
result Derives minimax optimal excess risk bounds for robust WDRO.
Paper establishes a universal growth rate for smooth surrogate losses in classification.
problem Analyzing growth rates of consistency bounds for various surrogate losses.
method Proves square-root growth rate for smooth margin-based losses; extends to multi-class classification.
result Demonstrates a universal square-root growth rate for smooth comp-sum and constrained losses.
The paper explores the information-theoretic nature of excess risk in machine learning.
problem Understanding the excess risk in machine learning models.
method Formulates the minimax excess risk as a zero-sum game and modifies it to allow swapping of the order of play.
result Proves that under certain conditions, the duality gap is zero, allowing for the application of Bayesian results to provide bounds on minimax excess risk.
Audit shows risk claims from distributional reinforcement learning agents are often false.
problem Evaluating the risk claims made by distributional reinforcement learning agents.
method Combines a decision-relevant screening metric, ground truth from Monte Carlo, and statistical methods to audit risk claims.
result 40-95% of the strongest risk claims are refuted, indicating the learned risk reflects a training artifact rather than environment stochasticity.
This paper investigates WDRO for nonparametric regression, achieving robustness against distributional uncertainty.
problem Addressing model misspecification in nonparametric regression under distributional uncertainty.
method Wasserstein distributionally robust optimization (WDRO) with structural distinction based on Wasserstein distance order.
result Achieves a convergence rate of n−2β/(d+2β) up to logarithmic factors, showing minimax optimality. Mathematical study of excess growth rate connects info theory with finance.
problem Understanding the excess growth rate in portfolio theory.
method Axiomatic characterization theorems of excess growth rate in terms of relative entropy, Jensen's inequality gap, and logarithmic divergence.
result Established rich connections between information theory and finance.
WR-CP reduces prediction set size and coverage gap under distribution shift.
problem Guaranteed coverage under distribution shift not achievable with i.i.d. assumption.
method Wasserstein distance, probability measure pushforwards, importance weighting, regularized representation learning.
result Reduces coverage gap to 3.2% across different confidence levels.
Paper introduces a novel framework for supervised graph prediction using Optimal Transport.
problem Supervised labeled graph prediction.
method Fused Gromov-Wasserstein (FGW) loss and FGW barycenter with neural network weights and learned graphs.
result The method can interpolate in the labeled graph space and achieve good performance on difficult problems.
A new method for distribution regression using sliced Wasserstein distance.
problem Learning functions over spaces of probabilities.
method Proposes an OT-based estimator using the Sliced Wasserstein distance.
result Proves universal consistency and excess risk bounds for the proposed estimator.
Exact generalization guarantees for robust models using Wasserstein distance are established.
problem Capturing data uncertainty and distribution shifts in machine learning models.
method Establishes exact generalization guarantees for robust models based on the Wasserstein distance, covering various cases and transport costs.
result Exact generalization guarantees are provided for a wide range of cases, including deep learning objectives with nonsmooth activations.
Unified framework for ICL in causal and masked models.
problem Understanding ICL in masked language models and comparing it to causal models.
method Developed a statistical learning framework representing context by empirical measure and predicting using context and query.
result Upper bounds for masked and autoregressive objectives under Wasserstein-type regularity conditions.
This paper bridges variational inference and Wasserstein gradient flows.
problem Combining variational inference and Wasserstein gradient flows for more efficient approximations.
method Recasting Bures-Wasserstein gradient flow as a Euclidean gradient flow and using path-derivative gradient estimator.
result A new gradient estimator for f-divergences that can be implemented using machine learning libraries. In this paper, we are concerned with a non-asymptotic analysis of sampling algorithms used in nonconvex optimization. In particular, we obtain non-asymptotic estimates in Wasserstein-1 and Wasserstein-2 distances for a popular class of algorithms called Stochastic Gradient Langevin Dynamics (SGLD). In addition, the afo…
New Langevin algorithm works well even for rough distributions.
problem Sampling from non-smooth distributions.
method Simple Langevin algorithm without smoothness assumptions.
result Algorithm performs well even with discontinuous gradients.
The study provides statistical theory for WGANs in time series forecasting.
problem Statistical analysis of WGANs for time series forecasting.
method Statistical theory and upper bounds for excess Bayes risk, weak convergence, and confidence intervals.
result Developed confidence intervals for time series forecasting using WGANs.
Generative adversarial networks (GANs) have enjoyed much success in learning high-dimensional distributions. Learning objectives approximately minimize an f-divergence (f-GANs) or an integral probability metric (Wasserstein GANs) between the model and the data distribution using a discriminator. Wasserstein GANs en…
Paper analyzes SGHMC for non-convex optimization with discontinuous gradients.
problem Training neural networks with ReLU activation.
method Non-asymptotic convergence analysis of SGHMC with discontinuous gradients.
result Explicit upper bounds for expected excess risk in non-convex optimization.
It is proved by Brendle in [4] that the equatorial disk Dk has least area among k-dimensional free boundary minimal surfaces in the Euclidean ball Bn. By comparing the excess of free boundary minimal surfaces with the excess of the associated cones over the boundary, we prove the existence of a gap for the area…
New algorithms sample from log concave distributions without gradient Lipschitz continuity.
problem Sampling from log concave distributions without gradient Lipschitz continuity.
method Two algorithms based on monotone polygonal (tamed) Euler schemes.
result Non-asymptotic 2-Wasserstein distance bounds between the process and target measure.
The paper develops a theory for one-step Wasserstein-guided models for PDE-induced measures.
problem Theoretical understanding of generative models' accuracy in scientific computing.
method Regularity theory for optimal transport between doubling measures, excess-risk bounds.
result One-step Wasserstein-guided generative models can approximate PDE-induced measures with Hölder continuity.
Sharp stability of Alexandrov's theorem for C1 domains in the small-excess regime
problem Stability of Alexandrov's theorem for C1 domains in the small-excess regime method Combines a BV version of Fuglede's spectral-gap argument, a star-shaped rearrangement for sets of finite perimeter, quantitative estimates for the part of the boundary contained in the tentacles, and a polyhedral approximation argument for the non-graphical region result Sharp stability estimate in a genuinely non-parametric regime
A new tamed stochastic gradient Hamiltonian Monte Carlo algorithm for superlinearly growing stochastic gradients.
problem Sampling and stochastic optimization problems with superlinearly growing stochastic gradients.
method Tamed Stochastic Gradient Hamiltonian Monte Carlo (tSGHMC) algorithm.
result Established a non-asymptotic error bound in Wasserstein-2 distance with a convergence rate of 1/4. Machine learning's predictive power is limited by sample size, as shown by the Limits-to-Learning Gap.
problem The limitations of machine learning in approximating true data-generating processes.
method Characterization of a universal lower bound (LLG) quantifying the discrepancy between empirical fit and population benchmark.
result Standard ML approaches can substantially understate true predictability in financial data.
The paper studies scaling limits of Wasserstein metrics on Gaussian mixture models.
problem Understanding the scaling limits of Wasserstein metrics on Gaussian mixture models.
method Scaling limit approach on Gaussian mixture models, including inhomogeneous and extended models.
result Existence of the limit of the Wasserstein metric after renormalization for GMMs with zero variance.
New bounds close the score matching gap for diffusion models.
problem The difference between sample quality and score matching loss in diffusion models.
method Theoretical analysis of score matching gap, developing tighter bounds for KL divergence, reverse KL divergence, and Wasserstein distance.
result The quality of score approximation impacts closing the score matching gap for low noise scales.
Paper introduces a novel measure to analyze excess error in classification under covariate shift.
problem Analyzing excess error in classification under covariate shift.
method Utilizes vicinity information to characterize excess error.
result Faster or competitive convergence rates compared to previous techniques.
Vanilla GANs are connected to Wasserstein distance for better understanding.
problem Understanding the statistical properties of Vanilla GANs.
method Connecting Vanilla GANs to Wasserstein distance and proving an oracle inequality.
result An oracle inequality for Vanilla GANs in Wasserstein distance is obtained.
Paper proposes a robust method for inferring parameters in multiobjective optimization.
problem Uncertainty in hypothetical decision-making problem, data quality, and parameter space.
method Wasserstein distributionally robust approach for inverse multiobjective optimization.
result WRO-IMOP minimizes worst-case expected loss over a Wasserstein ball of distributions.
New framework analyzes deep learning optimization with finite width networks, revealing generalization gaps and excess risks.
problem Analyzing generalization error of deep learning with finite width networks.
method Formulating neural network training as transportation map estimation and analyzing via infinite dimensional Langevin dynamics.
result Achieves fast learning rate and minimax optimal rates for classification and regression problems.
Paper develops KMS Wasserstein for high-dimensional data reduction.
problem Optimal transport's curse of dimensionality in high-dimensional data.
method Kernel max-sliced (KMS) Wasserstein distance for dimensionality reduction.
result Sharp finite-sample guarantees for KMS p-Wasserstein distance. We study differentially private (DP) algorithms for stochastic convex optimization (SCO). In this problem the goal is to approximately minimize the population loss given i.i.d. samples from a distribution over convex and Lipschitz loss functions. A long line of existing work on private convex optimization focuses on th…
New bounds improve generalization in machine learning with high probability.
problem Erratic behavior of KL divergence limits practical applications.
method Replaced KL divergence with Wasserstein distance for better bounds.
result Proved high probability generalization bounds for i.i.d. and non-i.i.d. data.
Our attacks are stronger and faster under Wasserstein metric.
problem Vulnerability of deep models to adversarial attacks.
method Developed an exact yet efficient projection operator and used the Frank-Wolfe method.
result Generated much stronger attacks and improved model robustness.
Defines MER for Bayesian learning, a gap between achievable and optimal performance.
problem Analyzing the best performance of Bayesian learning under generative models.
method Two methods for deriving upper bounds for MER: conditional mutual information and minimum estimation error.
result Quantifies the rate at which MER decays to zero with more data and relates it to model richness.
Gradient descent methods for deep ReLU networks achieve optimal generalization rates.
problem Generalization of gradient descent methods for deep neural networks
method Establishing minimax-optimal rates for GD and SGD with deep ReLU networks
result Gradient descent methods for deep ReLU networks achieve optimal generalization rates
Develops a framework to test excessive influence of small data subsets.
problem Identifying when small data subsets significantly impact model conclusions.
method Formalizes the concept of most influential sets, deriving influence formulas and extreme value distributions.
result Allows rigorous hypothesis testing for excessive influence, resolving contested findings.
This study tightens bounds on how GD and SGD generalize in smooth convex optimization problems.
problem Understanding how GD and SGD generalize in smooth stochastic convex optimization problems.
method Provided tight excess risk lower bounds for GD and SGD under different conditions.
result Lower bounds suggest overfitting occurs and gaps remain in some cases.
ITSPACE improves covariance alignment faster than other methods.
problem Optimizing covariance matrices for machine learning tasks.
method Proximal majorization-minimization method that directly optimizes the Bures-Wasserstein objective.
result ITSPACE achieves lower BW gap solutions faster than other methods.
Partial Wasserstein Covering aims to identify missing patterns in datasets.
problem Identifying missing patterns in datasets compared to actual applications.
method Formulated as a discrete optimization problem with partial Wasserstein divergence. Proved submodular, allowing greedy approximation. Proposed quasi-greedy algorithms with acceleration techniques.
result Efficiently fills gaps and finds missing scenes in real driving scenes datasets.
Develops a method for fairness in multi-task learning using Wasserstein barycenters.
problem Extending fairness to multi-task learning with shared representations.
method Definition of Strong Demographic Parity extended to multi-task learning using multi-marginal Wasserstein barycenters. Closed form solution for optimal fair predictor.
result Empirical results show practical value of post-processing methodology in promoting fair decision-making.
Paper introduces a method to control early classification accuracy gaps.
problem Maintaining accuracy in early classification without full input processing.
method Statistical framework for a calibrated stopping rule.
result Reduces up to 94% of timesteps while controlling accuracy gaps.
Improves point-cloud reconstruction by optimizing projections with self-attention.
problem Inefficient and non-metric projection methods for sliced Wasserstein distances.
method Proposes distributional sliced Wasserstein distance with self-attention for permutation-invariant and metric optimization.
result Self-attention amortized distributional projection optimization achieves better performance in point-cloud reconstruction.
Paper provides a rigorous proof of the index theorem for economists.
problem Lack of a rigorous proof in textbooks for economists.
method Constructs a readable proof of the index theorem under specific assumptions.
result Provides a gap-free proof of the index theorem.
A new method for learning gradient flows from population dynamics.
problem Reconstructing population dynamics from limited data.
method Residual approach to enforce continuity equations, combining with data-fitting divergence.
result Demonstrated state-of-the-art performance across trajectory inference benchmarks.
We consider a rigidity problem for the spectral gap of the Laplacian on an RCD(K,∞)-space (a metric measure space satisfying the Riemannian curvature-dimension condition) for positive K. For a weighted Riemannian manifold, Cheng--Zhou showed that the sharp spectral gap is achieved only when a 1-dimensional G…
High-dimensional curved diffusions show abrupt convergence at a critical time.
problem Understanding abrupt convergence in high-dimensional curved diffusions.
method Functional inequalities and spectral rigidity.
result Abrupt convergence (cutoff) occurs in high dimensions, linked to spectral rigidity.
Improved reSGLD accelerates convergence in non-convex learning problems.
problem Inefficient swaps due to noisy energy estimators in reSGLD.
method Variance reduction for noisy energy estimators, theoretical analysis, and numerical experiments.
result Exponential acceleration in convergence for non-convex learning problems.