This paper provides a dictionary of closed-form kernel mean embeddings.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This note optimizes distributions using kernel mean embeddings with a new parameterization.
Kernel methods are one of the mainstays of machine learning, but the problem of kernel learning remains challenging, with only a few heuristics and very little theory. This is of particular importance in methods based on estimation of kernel mean embeddings of probability measures. For characteristic kernels, which inc…
Conditional kernel mean embeddings form an attractive nonparametric framework for representing conditional means of functions, describing the observation processes for many complex models. However, the recovery of the original underlying function of interest whose conditional mean was observed is a challenging inferenc…
New KQEs improve probability metrics without mean function constraints.
Efficiently approximates kernel mean embeddings using Nyström method.
New recursive algorithm estimates conditional kernel mean embeddings in Hilbert space.
A new method estimates multi-dimensional value distributions using Hilbert space embeddings.
Faster convergence of kernel mean embeddings using variance information.
Kernel means are frequently used to represent probability distributions in machine learning problems. In particular, the well known kernel density estimator and the kernel mean embedding both have the form of a kernel mean. Unfortunately, kernel means are faced with scalability issues. A single point evaluation of the …
A Hilbert space embedding for probability measures has recently been proposed, wherein any probability measure is represented as a mean element in a reproducing kernel Hilbert space (RKHS). Such an embedding has found applications in homogeneity testing, independence testing, dimensionality reduction, etc., with the re…
IDK improves anomaly detection for points and groups without explicit learning.
Paper generalizes kernel mean embedding to von Neumann-algebra-valued measures.
We provide a theoretical foundation for non-parametric estimation of functions of random variables using kernel mean embeddings. We show that for any continuous function , consistent estimators of the mean embedding of a random variable lead to consistent estimators of the mean embedding of . For Matérn ke…
Conditional kernel mean embeddings are nonparametric models that encode conditional expectations in a reproducing kernel Hilbert space. While they provide a flexible and powerful framework for probabilistic inference, their performance is highly dependent on the choice of kernel and regularization hyperparameters. Neve…
Paper introduces RKHM and KME for richer data analysis.
Paper develops a unified framework for measuring differences between conditional distributions.
Kernelized Taylor diagram visualizes data populations with fewer assumptions.
We present an operator-free, measure-theoretic approach to the conditional mean embedding (CME) as a random variable taking values in a reproducing kernel Hilbert space. While the kernel mean embedding of unconditional distributions has been defined rigorously, the existing operator-based approach of the conditional ve…
Hermite polynomials improve private data generation by reducing feature count.
Paper improves MMD estimation for analytical mean embeddings.
New tests for binary classification regression functions without distribution assumptions.
New method uses Fokker-Planck equation for sampling and inference.
The paper introduces new KMEs to capture stochastic process filtrations.
CPME embeds counterfactual outcomes in RKHS for flexible policy evaluation.
Neural-Kernel CME tackles scalability and expressiveness challenges in conditional distribution representation.
We lay theoretical foundations for new database release mechanisms that allow third-parties to construct consistent estimators of population statistics, while ensuring that the privacy of each individual contributing to the database is protected. The proposed framework rests on two main ideas. First, releasing (an esti…
A mean function in reproducing kernel Hilbert space, or a kernel mean, is an important part of many applications ranging from kernel principal component analysis to Hilbert-space embedding of distributions. Given finite samples, an empirical average is the standard estimate for the true kernel mean. We show that this e…
A Hilbert space embedding of a distribution---in short, a kernel mean embedding---has recently emerged as a powerful tool for machine learning and inference. The basic idea behind this framework is to map distributions into a reproducing kernel Hilbert space (RKHS) in which the whole arsenal of kernel methods can be ex…
The objective in statistical Optimal Transport (OT) is to consistently estimate the optimal transport plan/map solely using samples from the given source and target marginal distributions. This work takes the novel approach of posing statistical OT as that of learning the transport plan's kernel mean embedding from sam…
Improved multi-task averaging reduces mean squared error in high-dimensional data.
A new method for distribution regression using sliced Wasserstein distance.
Mean embeddings provide an extremely flexible and powerful tool in machine learning and statistics to represent probability distributions and define a semi-metric (MMD, maximum mean discrepancy; also called N-distance or energy distance), with numerous successful applications. The representation is constructed as the e…
A mean function in a reproducing kernel Hilbert space (RKHS), or a kernel mean, is central to kernel methods in that it is used by many classical algorithms such as kernel principal component analysis, and it also forms the core inference step of modern kernel methods that rely on embedding probability distributions in…
This paper connects functional data analysis with machine learning techniques.
We propose a differentially private data generation paradigm using random feature representations of kernel mean embeddings when comparing the distribution of true data with that of synthetic data. We exploit the random feature representations for two important benefits. First, we require a minimal privacy cost for tra…
Conditional mean embeddings (CMEs) have proven themselves to be a powerful tool in many machine learning applications. They allow the efficient conditioning of probability distributions within the corresponding reproducing kernel Hilbert spaces (RKHSs) by providing a linear-algebraic relation for the kernel mean embedd…
Study optimizes learning rates for conditional mean embedding estimates.
This paper presents a distance-based discriminative framework for learning with probability distributions. Instead of using kernel mean embeddings or generalized radial basis kernels, we introduce embeddings based on dissimilarity of distributions to some reference distributions denoted as templates. Our framework exte…
Kernel K-means clusters probability distributions.
In likelihood-free settings where likelihood evaluations are intractable, approximate Bayesian computation (ABC) addresses the formidable inference task to discover plausible parameters of simulation programs that explain the observations. However, they demand large quantities of simulation calls. Critically, hyperpara…
Kernel embeddings separate distinct probability distributions, simplifying testing.
We introduce two kernels that extend the mean map, which embeds probability measures in Hilbert spaces. The generative mean map kernel (GMMK) is a smooth similarity measure between probabilistic models. The latent mean map kernel (LMMK) generalizes the non-iid formulation of Hilbert space embeddings of empirical distri…
We propose a new method to model multi-way similarities into hypergraphs for clustering.
Kernel mean embeddings have recently attracted the attention of the machine learning community. They map measures from some set to functions in a reproducing kernel Hilbert space (RKHS) with kernel . The RKHS distance of two mapped measures is a semi-metric over . We study three questions. (I) For a…
Proposes a new method to analyze the distributional effects of treatments.
Proposes CCME framework for estimating heterogeneous treatment effects.
New tests for distributional causal effects using improved kernel estimators.