Kernel embeddings separate distinct probability distributions, simplifying testing.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A Hilbert space embedding for probability measures has recently been proposed, wherein any probability measure is represented as a mean element in a reproducing kernel Hilbert space (RKHS). Such an embedding has found applications in homogeneity testing, independence testing, dimensionality reduction, etc., with the re…
A Hilbert space embedding for probability measures has recently been proposed, with applications including dimensionality reduction, homogeneity testing, and independence testing. This embedding represents any probability measure as a mean element in a reproducing kernel Hilbert space (RKHS). A pseudometric on the spac…
The paper studies PCA of probability measures with varying sample sizes and finds optimal convergence rates.
This note optimizes distributions using kernel mean embeddings with a new parameterization.
This work studies an explicit embedding of the set of probability measures into a Hilbert space, defined using optimal transport maps from a reference probability density. This embedding linearizes to some extent the 2-Wasserstein space, and enables the direct use of generic supervised and unsupervised learning algorit…
The Skorokhod embedding problem aims to represent a given probability measure on the real line as the distribution of Brownian motion stopped at a chosen stopping time. In this paper, we consider an extension of the optimal Skorokhod embedding problem to the case of finitely-many marginal constraints. Using the classic…
The study shows that ergodic measures are not generic on non-positively curved manifolds.
The paper extends optimal transport for linear separability of sheared distributions in supervised learning.
A new method optimizes slicing directions for SW distances to improve high-dimensional probability measure comparison.
A novel kernel-based test detects equality versus singularity of two probability measures.
Due to Čencov's theorem, there exists a unique family of invariant symmetric -tensor fields on the space of positive probability measures on a set of -points indexed by under Markov embeddings. We deform Markov embeddings keeping sufficiency, and prove existence and uniqueness of invariant f…
Kernel embeddings help estimate causal effects from observational data.
Measures time-delay embedding for noisy, sparse data.
A new metric for comparing probability measures on graphs, scalable and negative definite.
Paper proposes a new method to learn distribution kernels via entropy maximization.
A new kernel for probability measures based on optimal transport.
Handlebody groups are rigid under measure equivalence.
Wassmap reduces image complexity while preserving key features.
Expected centre of mass for random embeddings is constant.
Paper generalizes kernel mean embedding to von Neumann-algebra-valued measures.
Embedding methods which enforce a partial order or lattice structure over the concept space, such as Order Embeddings (OE) (Vendrov et al., 2016), are a natural way to model transitive relational data (e.g. entailment graphs). However, OE learns a deterministic knowledge base, limiting expressiveness of queries and the…
If M is a smooth compact Riemannian manifold, let P(M) denote the Wasserstein space of probability measures on M. If S is an embedded submanifold of M, and is an absolutely continuous measure on S, then we compute the tangent cone of P(M) at .
A new method estimates multi-dimensional value distributions using Hilbert space embeddings.
In this note we prove certain necessary and sufficient conditions for the existence of an embedding of statistical manifolds. In particular, we prove that any compact smooth ( resp.) statistical manifold can be embedded into the space of probability measures on a finite set. As a result, we get an answer to the La…
Kernel methods are one of the mainstays of machine learning, but the problem of kernel learning remains challenging, with only a few heuristics and very little theory. This is of particular importance in methods based on estimation of kernel mean embeddings of probability measures. For characteristic kernels, which inc…
We propose to formulate multi-label learning as a estimation of class distribution in a non-linear embedding space, where for each label, its positive data embeddings and negative data embeddings distribute compactly to form a positive component and negative component respectively, while the positive component and nega…
We introduce a weak notion of barycenter of a probability measure on a metric measure space , with the metric and reference measure . Under the assumption that optimal transport plans are given by mappings, we prove that our barycenter is well defined; it is a probability measur…
Kernel mean embeddings have recently attracted the attention of the machine learning community. They map measures from some set to functions in a reproducing kernel Hilbert space (RKHS) with kernel . The RKHS distance of two mapped measures is a semi-metric over . We study three questions. (I) For a…
We study the problem of estimating, in the sense of optimal transport metrics, a measure which is assumed supported on a manifold embedded in a Hilbert space. By establishing a precise connection between optimal transport metrics, optimal quantization, and learning theory, we derive new probabilistic bounds for the per…
This paper extends the proof of density of neural networks in the space of continuous (or even measurable) functions on Euclidean spaces to functions on compact sets of probability measures. By doing so the work parallels a more then a decade old results on mean-map embedding of probability measures in reproducing kern…
We propose a novel node embedding of directed graphs to statistical manifolds, which is based on a global minimization of pairwise relative entropy and graph geodesics in a non-linear way. Each node is encoded with a probability density function over a measurable space. Furthermore, we analyze the connection between th…
In this note we discuss a common misconception, namely that embeddings are always used to reduce the dimensionality of the item space. We show that when we measure dimensionality in terms of information entropy then the embedding of sparse probability distributions, that can be used to represent sparse features or data…
New graph embedding method improves link prediction and node classification.
This paper presents a distance-based discriminative framework for learning with probability distributions. Instead of using kernel mean embeddings or generalized radial basis kernels, we introduce embeddings based on dissimilarity of distributions to some reference distributions denoted as templates. Our framework exte…
This paper presents a kernel-based discriminative learning framework on probability measures. Rather than relying on large collections of vectorial training examples, our framework learns using a collection of probability distributions that have been constructed to meaningfully represent training data. By representing …
For a positive integer , the collection of -sided polygons embedded in -space defines the space of geometric knots. We will consider the subspace of equilateral knots, consisting of embedded -sided polygons with unit length edges. Paths in this space determine isotopies of polygons, so path-components …
Motivated by the model- independent pricing of derivatives calibrated to the real market, we consider an optimization problem similar to the optimal Skorokhod embedding problem, where the embedded Brownian motion needs only to reproduce a finite number of prices of Vanilla options. We derive in this paper the correspon…
We address the problem of unsupervised domain adaptation (UDA) by learning a cross-domain agnostic embedding space, where the distance between the probability distributions of the two source and target visual domains is minimized. We use the output space of a shared cross-domain deep encoder to model the embedding spac…
Time-delayed embeddings avoid self-intersections for high enough delay.
Quantum probability metrics improve distribution comparison in high dimensions.
A Hilbert space embedding of a distribution---in short, a kernel mean embedding---has recently emerged as a powerful tool for machine learning and inference. The basic idea behind this framework is to map distributions into a reproducing kernel Hilbert space (RKHS) in which the whole arsenal of kernel methods can be ex…
With the rapid growth in fashion e-commerce and customer-friendly product return policies, the cost to handle returned products has become a significant challenge. E-tailers incur huge losses in terms of reverse logistics costs, liquidation cost due to damaged returns or fraudulent behavior. Accurate prediction of prod…
New insights into Markov chain geometry via positive transition measures.
Random hyperbolic surfaces are mostly tangle-free, with geometric implications.
Paper improves MMD estimation for analytical mean embeddings.
NNLMs optimize poorly for word probabilities due to embedding space structure.
Embedding complex objects as vectors in low dimensional spaces is a longstanding problem in machine learning. We propose in this work an extension of that approach, which consists in embedding objects as elliptical probability distributions, namely distributions whose densities have elliptical level sets. We endow thes…