Paper addresses concentration of distances for fractional quasi p-norms, identifying conditions for concentration and anti-concentration.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study on volume of tubes and concentration in Riemannian geometry.
In this paper, we study the Lévy-Milman concentration phenomenon of 1-Lipschitz maps into infinite dimensional metric spaces. Our main theorem asserts that the concentration to an infinite dimensional -ball with the -distance function for is equivalent to the concentration to the…
Sharp stability result for maps near infinitely concentrated minimisers.
We define generalized distance-squared mappings, and we concentrate on the plane to plane case. We classify generalized distance-squared mappings of the plane into the plane in a recognizable way.
A new method approximates the Sliced-Wasserstein distance without random projections.
Study on estimating distances between covariance operators and Gaussian processes.
Improved estimation of concentration using half-spaces for adversarial vulnerability.
The paper analyzes how noise affects distances in high-dimensional data and when they remain useful.
New robust method for optimal transportation improves statistical inference.
This paper analyzes minibatch optimal transport distances and their applications.
This paper shows how to estimate distances in latent space of random graphs using entropic OT.
Constructs flows on manifolds with small curvature, proving Euclidean topology.
A new distance metric for vMF distributions simplifies spherical data analysis.
We provide upper bounds of the expected Wasserstein distance between a probability measure and its empirical version, generalizing recent results for finite dimensional Euclidean spaces and bounded functional spaces. Such a generalization can cover Euclidean spaces with large dimensionality, with the optimal dependence…
This paper presents a unified approach based on Wasserstein distance to derive concentration bounds for empirical estimates for two broad classes of risk measures defined in the paper. The classes of risk measures introduced include as special cases well known risk measures from the finance literature such as condition…
We study the non-asymptotic behavior of a Coulomb gas on a compact Riemannian manifold. This gas is a symmetric n-particle Gibbs measure associated to the two-body interaction energy given by the Green function. We encode such a particle system by using an empirical measure. Our main result is a concentration inequalit…
Many recent works have shown that adversarial examples that fool classifiers can be found by minimally perturbing a normal input. Recent theoretical results, starting with Gilmer et al. (2018b), show that if the inputs are drawn from a concentrated metric probability space, then adversarial examples with small perturba…
We present a novel notion of outlier, called the Concentration Free Outlier Factor, or CFOF. As a main contribution, we formalize the notion of concentration of outlier scores and theoretically prove that CFOF does not concentrate in the Euclidean space for any arbitrary large dimensionality. To the best of our knowled…
A robust conformal method for set estimation using non-conformity scores.
Study on computing and estimating calibration distance, showing hardness and efficiency.
SGD converges to an invariant distribution with sub-Gaussian or sub-exponential properties.
Develops robust MDPs for unknown disturbances with performance guarantees.
Many modern machine learning classifiers are shown to be vulnerable to adversarial perturbations of the instances. Despite a massive amount of work focusing on making classifiers robust, the task seems quite challenging. In this work, through a theoretical study, we investigate the adversarial risk and robustness of cl…
The paper analyzes sparse high-dimensional linear regression with random design and unknown error variance, providing adaptiveness and concentration rates.
The study proves inequalities and curvature properties for Markov chains.
New methods improve accuracy in detecting concentric objects.
We propose fast approximations for the generalized sliced-Wasserstein distance.
Optimal transport theory has recently found many applications in machine learning thanks to its capacity for comparing various machine learning objects considered as distributions. The Kantorovitch formulation, leading to the Wasserstein distance, focuses on the features of the elements of the objects but treat them in…
Kernel methods are successful approaches for different machine learning problems. This success is mainly rooted in using feature maps and kernel matrices. Some methods rely on the eigenvalues/eigenvectors of the kernel matrix, while for other methods the spectral information can be used to estimate the excess risk. An …
Develops a two-sample test using projected Wasserstein distance to handle high-dimensional data.
It is well known that isoperimetric inequalities imply in a very general measure-metric-space setting appropriate concentration inequalities. The former bound the boundary measure of sets as a function of their measure, whereas the latter bound the measure of sets separated from sets having half the total measure, as a…
New method improves sampling efficiency in complex stochastic systems.
The study uses heat flow to analyze properties of Laplace eigenfunctions on manifolds and domains.
Method detects effects of synthesis parameters on plutonium oxide microstructure.
Many statistical and machine learning approaches rely on pairwise distances between data points. The choice of distance metric has a fundamental impact on performance of these procedures, raising questions about how to appropriately calculate distances. When data points are real-valued vectors, by far the most common c…
We introduce an universum of the Polish (=complete separable metric) space - the convex cone of distance matrices and study its geometry. It happened that the generic Polish spaces in this sense of this universum is so called Urysohn spaces defined by P.S.Urysohn in 20-th, and generic metric triple (= metric space with…
Random feature matrices' singular values concentrate near their full expectation in high dimensions.
Robustly clusters mixtures of Gaussians even with outliers.
A new framework tightens risk measure confidence bounds.
For incomplete sub-Riemannian manifolds, and for an associated second-order hypoelliptic operator, which need not be symmetric, we identify two alternative conditions for the validity of Gaussian-type upper bounds on heat kernels and transition probabilities, with optimal constant in the exponent. Under similar conditi…
New approach to quantify posterior concentration rates using Wasserstein dynamics.
Optimal transport distances are powerful tools to compare probability distributions and have found many applications in machine learning. Yet their algorithmic complexity prevents their direct use on large scale datasets. To overcome this challenge, practitioners compute these distances on minibatches {\em i.e.} they a…
Construction of ambiguity set in robust optimization relies on the choice of divergences between probability distributions. In distribution learning, choosing appropriate probability distributions based on observed data is critical for approximating the true distribution. To improve the performance of machine learning …
Adaptive sampling theory has shown that, with proper assumptions on the signal class, algorithms exist to reconstruct a signal in with an optimal number of samples. We generalize this problem to the case of spatial signals, where the sampling cost is a function of both the number of samples taken and t…
Estimating entropy and mutual information consistently is important for many machine learning applications. The Kozachenko-Leonenko (KL) estimator (Kozachenko & Leonenko, 1987) is a widely used nonparametric estimator for the entropy of multivariate continuous random variables, as well as the basis of the mutual inform…
Max-sliced Wasserstein metric reduces high-dimensional data to 1D for better estimation.
This work aims to study the Portuguese regional agglomeration process, using the linear form the New Economic Geography models that emphasize the importance of spatial factors (distance, costs of transport and communication) in explaining of the concentration of economic activity in certain locations. In a theoretical …