This note optimizes distributions using kernel mean embeddings with a new parameterization.
problem Optimizing distributions using kernel mean embeddings is challenging due to the difficulty of characterizing probability distribution vectors.
method Proposes a new parameterization of positive functions using kernel sums-of-squares to fit distributions in the MMD geometry.
result Distributions with kernel sum-of-squares densities are dense in the MMD geometry, allowing optimization in the finite-sample setting.
A new method estimates multi-dimensional value distributions using Hilbert space embeddings.
problem Estimating value distributions in complex, multi-dimensional reinforcement learning settings.
method Hilbert space mappings and kernel mean embeddings to estimate the kernel mean embedding of multi-dimensional value distributions.
result Uniform convergence guarantees and robust off-policy evaluation demonstrated in simulations.
Paper proposes a new method to learn distribution kernels via entropy maximization.
problem Challenges in applying kernel methods to distribution regression tasks.
method Proposes a novel objective for unsupervised learning of data-dependent distribution kernels based on entropy maximization.
result Demonstrates the effectiveness of the learned kernel across different modalities.
Kernel embeddings help estimate causal effects from observational data.
problem Estimating causal effects from observational data with confounding variables.
method Kernel embeddings in reproducing kernel Hilbert spaces (RKHS).
result Robust nonparametric framework for causal inference.
Kernel mean embedding maps distributions into RKHS for machine learning.
problem Efficiently representing and comparing probability distributions.
method Mapping distributions into RKHS using kernel methods.
result Kernel mean embedding enables non-parametric operations on distributions.
Kernel embeddings map measures to functions in RKHS, addressing embedding and metric properties.
problem Characterizing sets of measures that can be embedded and the conditions for embedding to be injective.
method Study of kernel mean embeddings, focusing on universal, characteristic, and strictly positive definite kernels.
result Unified and extended results on embedding and metric properties of measures.
Kernel embeddings separate distinct probability distributions, simplifying testing.
problem Testing equality of non-atomic probability distributions.
method Kernel covariance embeddings and Gaussian measures in reproducing kernel Hilbert spaces.
result Testing for singularity between Gaussian measures is equivalent to testing for equality of non-atomic probability distributions.
Improved learning theory for kernel distribution regression with two-stage sampling.
problem Distribution regression problem and two-stage sampling setting.
method Kernel methods, near-unbiased condition, new error bounds, convergence rates.
result Strictly improved convergence rates for three important classes of kernels.
IDK improves anomaly detection for points and groups without explicit learning.
problem Anomaly detection for points and groups using kernel methods.
method Isolation Distributional Kernel (IDK) addresses data independence and intractable dimensionality issues.
result IDK outperforms existing methods for both point and group anomaly detection.
Test for joint distributions equivalence using kernel embedding.
problem Determining statistical difference between joint distributions.
method Joint kernel distribution embedding to extend kernel two-sample test.
result Can verify dataset-shift in learning frameworks.
New test for conditional independence using kernel embeddings.
problem Testing conditional independence in high-dimensional settings.
method Analytic kernel embeddings, asymptotic distribution.
result New test outperforms existing methods in high-dimensional settings.
Faster convergence of kernel mean embeddings using variance information.
problem Speeding up the convergence rate of kernel mean embeddings.
method Leveraging variance information in reproducing kernel Hilbert space and estimating variance from data.
result Efficiently estimate variance information from data to achieve distribution-agnostic convergence bounds.
New recursive algorithm estimates conditional kernel mean embeddings in Hilbert space.
problem Estimating conditional distributions in RKHS for supervised learning.
method Recursive algorithm in L2 space for conditional kernel mean map. result Strong L2 consistency of recursive estimator proved. Paper studies t-SNE convergence with generalized kernels.
problem Understanding convergence of t-SNE with generalized kernels.
method Concrete formulation of generalized kernels, proving convergence to an equilibrium distribution.
result t-SNE converges to an equilibrium distribution under certain conditions for generalized kernels.
New KQEs improve probability metrics without mean function constraints.
problem Improving probability metrics without relying on mean function representations.
method Kernel quantile embeddings (KQEs) to construct new distances.
result KQEs offer a competitive alternative to MMD with near-linear cost.
This paper provides a dictionary of closed-form kernel mean embeddings.
problem Challenges in deriving closed-form kernel mean embeddings.
method Comprehensive dictionary and practical tools for deriving new embeddings.
result Provides a Python library with minimal implementations of embeddings.
Paper develops a unified framework for measuring differences between conditional distributions.
problem Comparing conditional distributions in a unified and theoretically sound manner.
method Kernel embeddings and conditional maximum mean discrepancy (CMMD) framework.
result Established a coherent framework for measuring divergence between conditional distributions.
Kernel discriminant analysis uses nonlinear embeddings to improve classification.
problem Limited effectiveness of linear discriminant analysis in capturing nonlinear features.
method Study of nonlinear embeddings in kernel discriminant analysis using polynomial and Gaussian kernels, solving generalized eigenvalue problems.
result Polynomial and Gaussian discriminants capture class differences through population moments and randomized projections.
Kernel mean estimation for functions of random variables provides consistent estimators.
problem Estimating functions of random variables using kernel mean embeddings.
method Kernel mean embeddings for continuous functions of random variables.
result Consistent estimators of mean embeddings of functions of random variables.
New tools for uncertainty in dynamical systems without distribution assumptions.
problem Uncertainty representation in dynamical systems without distributional assumptions.
method Kernel mean embedding and kernel probabilistic programming.
result Distribution-free representation, comparison, and propagation of uncertainties.
Paper introduces a new method for learning with distributions using dissimilarity measures.
problem Learning with probability distributions using dissimilarity measures.
method Introduces embeddings based on dissimilarity of distributions to templates, extending similarity theory to population distributions.
result Proves that dissimilarity theory holds for empirical distributions and shows better performance of Wasserstein distance embedding.
Develops a rigorous theory for conditional mean embeddings.
problem Efficient conditioning of probability distributions in RKHSs.
method Mathematical theory for both centred and uncentred covariance operators.
result Significantly weakens conditions for applicability of CMEs.
A Hilbert space embedding for probability measures has recently been proposed, with applications including dimensionality reduction, homogeneity testing, and independence testing. This embedding represents any probability measure as a mean element in a reproducing kernel Hilbert space (RKHS). A pseudometric on the spac…
CPME embeds counterfactual outcomes in RKHS for flexible policy evaluation.
problem Estimating counterfactual policy outcomes for decision-making.
method Counterfactual Policy Mean Embedding (CPME) framework in RKHS, plug-in and doubly robust estimators, kernel test statistic.
result Doubly robust estimator improves convergence rates and asymptotic normality.
A statistical test of independence may be constructed using the Hilbert-Schmidt Independence Criterion (HSIC) as a test statistic. The HSIC is defined as the distance between the embedding of the joint distribution, and the embedding of the product of the marginals, in a Reproducing Kernel Hilbert Space (RKHS). It has …
Paper proposes FKR-F2E to make kernel regression fair.
problem Mitigating demographic biases in kernel methods.
method Fair feature embedding in kernel space.
result FKR-F2E achieves lower prediction disparity. New tests for binary classification regression functions without distribution assumptions.
problem Testing regression functions in binary classification without distributional assumptions.
method Conditional kernel mean embeddings and resampling-based framework.
result Distribution-free hypothesis tests with exact type I error control.
Neural-Kernel CME tackles scalability and expressiveness challenges in conditional distribution representation.
problem Scalability and expressiveness challenges in kernel conditional mean embeddings.
method Combines deep learning with CMEs using a neural network optimization framework.
result Achieves competitive and often superior performance in conditional density estimation and RL.
A new method for distribution regression using sliced Wasserstein distance.
problem Learning functions over spaces of probabilities.
method Proposes an OT-based estimator using the Sliced Wasserstein distance.
result Proves universal consistency and excess risk bounds for the proposed estimator.
We offer a new, rigorous approach to conditional mean embeddings without operator constraints.
problem Lack of rigorous, operator-free approach to conditional mean embeddings.
method Measure-theoretic approach to conditional mean embeddings.
result Natural regression interpretation and universal consistency of empirical estimates.
Estimates class prior for unlabeled data using kernel embedding.
problem Estimating class prior in PU learning scenario where only positive and full population samples are available.
method Direct estimator based on distribution matching and kernel embedding in Reproducing Kernel Hilbert Space.
result Asymptotic consistency and explicit deviation bound for the estimator.
Estimates exponential family distributions using a novel doubly dual embedding technique.
problem Estimating exponential family distributions with smoothness and efficiency.
method Doubly dual embedding for avoiding partition function computation and flexible sampling.
result Improves memory and time efficiency while offering stronger statistical properties.
A nonparametric approach for policy learning for POMDPs is proposed. The approach represents distributions over the states, observations, and actions as embeddings in feature spaces, which are reproducing kernel Hilbert spaces. Distributions over states given the observations are obtained by applying the kernel Bayes' …
New tests for distributional causal effects using improved kernel estimators.
problem Testing for higher-order moments and multidimensional outcomes affected by treatment.
method Improved kernel estimators based on doubly robust mean embeddings.
result New permutation-based tests for distributional causal effects with improved convergence rates.
Kernel-embedding tests can be suboptimal, but a simple modification improves their performance.
problem Optimizing goodness-of-fit tests using kernel embeddings.
method Analyzing and modifying kernel-embedding based goodness-of-fit tests within a minimax framework.
result A moderated kernel-embedding approach provides optimal tests for various deviations and is adaptive over a wide range of spaces.
Hermite polynomials improve private data generation by reducing feature count.
problem Infinite-dimensional features in kernel mean embedding are impractical for private data generation.
method Replace random features with Hermite polynomial features, leveraging their ordered nature.
result Hermite polynomial features yield a more accurate approximation of kernel mean embedding with fewer features.
Meta-learning for estimating complex conditional distributions.
problem Estimating conditional densities in multimodal distributions.
method Noise contrastive estimation with kernel mean embeddings.
result Meta-learning can share representations across tasks for conditional density estimation.
We provide a unifying framework linking two classes of statistics used in two-sample and independence testing: on the one hand, the energy distances and distance covariances from the statistics literature; on the other, distances between embeddings of distributions to reproducing kernel Hilbert spaces (RKHS), as establ…
Estimates low-rank distributional matrices from incomplete samples.
problem Matrix completion for distributional entries with limited observed data.
method Kernel mean embeddings, Tucker rank, functional unfolding operators.
result Effective estimator for distributional matrix completion established.
Efficiently approximates kernel mean embeddings using Nyström method.
problem Computational cost of kernel mean embeddings in large-scale settings.
method Nyström method for approximating a small random subset of the dataset.
result Upper bound on approximation error with sufficient subsample size conditions.
New graph kernel scales well with graph size and number, achieving state-of-the-art performance.
problem Graph kernels lose structure information when representing graphs.
method Proposes a positive-definite global alignment graph kernel using random features and random graph embeddings.
result Achieves quasi-linear scalability with respect to graph size and number.
Proposes ENCI for inferring nonstationary causal models.
problem Nonstationary data in real-world applications.
method Kernel embedding-based approach for multiple domains.
result Identifies cause-effect pairs and causal graphs under nonstationary conditions.
Kernel means are frequently used to represent probability distributions in machine learning problems. In particular, the well known kernel density estimator and the kernel mean embedding both have the form of a kernel mean. Unfortunately, kernel means are faced with scalability issues. A single point evaluation of the …
A new method compresses conditional distributions of labelled data.
problem No existing method directly compresses the conditional distribution of labelled data.
method Introduce Average Maximum Conditional Mean Discrepancy (AMCMD), derive a closed form estimator, and extend Kernel Herding (KH) to Average Conditional Kernel Herding (ACKH).
result Directly compressing conditional distributions outperforms joint distribution compression and greedy selection.
Novel approach to OT using kernel mean embeddings controls overfitting and achieves dimension-free sample complexity.
problem Consistently estimate optimal transport plan from samples.
method Pose OT as learning kernel mean embedding, employ MMD regularization.
result ε-optimal recovery of transport plan and map with dimension-free sample complexity.
Method calibrates simulators under covariate shift using kernel techniques.
problem Dealing with covariate shift in simulator inputs.
method Bayesian inference with kernel mean embedding and importance-weighted reproducing kernel.
result The method effectively calibrates simulators and demonstrates sensitivity analysis.
Paper generalizes kernel mean embedding to von Neumann-algebra-valued measures.
problem Analyzing complex multivariate distributions and quantum mechanics.
method Generalizes kernel mean embedding to von Neumann-algebra-valued measures in reproducing kernel Hilbert modules.
result Injectivity and universality of the generalized KME are confirmed.
Kernel K-means clusters probability distributions.
problem Clustering a sample of probability distributions.
method Mapping distributions to kernel mean embeddings in RKHS, then applying K-means.
result Effective unsupervised classification of probability distributions.