Strictly proper kernel scores are well-known tool in probabilistic forecasting, while characteristic kernels have been extensively investigated in the machine learning literature. We first show that both notions coincide, so that insights from one part of the literature can be used in the other. We then show that the m…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We connect shift-invariant characteristic kernels to infinitely divisible distributions on . Characteristic kernels play an important role in machine learning applications with their kernel means to distinguish any two probability measures. The contribution of this paper is two-fold. First, we show, usi…
A statistical test of independence may be constructed using the Hilbert-Schmidt Independence Criterion (HSIC) as a test statistic. The HSIC is defined as the distance between the embedding of the joint distribution, and the embedding of the product of the marginals, in a Reproducing Kernel Hilbert Space (RKHS). It has …
New kernels defined for various spaces, including measures.
A Hilbert space embedding for probability measures has recently been proposed, wherein any probability measure is represented as a mean element in a reproducing kernel Hilbert space (RKHS). Such an embedding has found applications in homogeneity testing, independence testing, dimensionality reduction, etc., with the re…
Kernel methods have been widely applied to machine learning and other questions of approximating an unknown function from its finite sample data. To ensure arbitrary accuracy of such approximation, various denseness conditions are imposed on the selected kernel. This note contributes to the study of universal, characte…
Decision forests are widely used for classification and regression tasks. A lesser known property of tree-based methods is that one can construct a proximity matrix from the tree(s), and these proximity matrices are induced kernels. While there has been extensive research on the applications and properties of kernels, …
Unified framework for constructing kernels for transport equations and Koopman eigenfunctions.
Kernel mean embeddings have recently attracted the attention of the machine learning community. They map measures from some set to functions in a reproducing kernel Hilbert space (RKHS) with kernel . The RKHS distance of two mapped measures is a semi-metric over . We study three questions. (I) For a…
Maximum mean discrepancy (MMD), also called energy distance or N-distance in statistics and Hilbert-Schmidt independence criterion (HSIC), specifically distance covariance in statistics, are among the most popular and successful approaches to quantify the difference and independence of random variables, respectively. T…
This note optimizes distributions using kernel mean embeddings with a new parameterization.
We discuss the finding that cross-sectional characteristic based models have yielded portfolios with higher excess monthly returns but lower risk than their arbitrage pricing theory counterparts in an analysis of equity returns of stocks listed on the JSE. Under the assumption of general no-arbitrage conditions, we arg…
Advanced kernels improve Gaussian process accuracy by incorporating domain knowledge.
A Hilbert space embedding for probability measures has recently been proposed, with applications including dimensionality reduction, homogeneity testing, and independence testing. This embedding represents any probability measure as a mean element in a reproducing kernel Hilbert space (RKHS). A pseudometric on the spac…
We examine groups whose resonance varieties, characteristic varieties and Sigma-invariants have a natural arithmetic group symmetry, and we explore implications on various finiteness properties of subgroups. We compute resonance varieties, characteristic varieties and Alexander polynomials of Torelli groups, and we sho…
Analyzes semi-characteristics on specific manifolds, proving a vanishing theorem.
Unified kernel framework extends to stochastic systems, improving numerical stability.
We propose non-stationary spectral kernels for Gaussian process regression. We propose to model the spectral density of a non-stationary kernel function as a mixture of input-dependent Gaussian process frequency density surfaces. We solve the generalised Fourier transform with such a model, and present a family of non-…
We prove a graph theoretic closed formula for coefficients in the Tian-Yau-Zelditch asymptotic expansion of the Bergman kernel. The formula is expressed in terms of the characteristic polynomial of the directed graphs representing Weyl invariants. The proof relies on a combinatorial interpretation of a recursive formul…
Paper extends RPD for better handling multiple modalities and non-convexity.
Paper improves kernel approximations for better statistical learning.
A faster graph kernel using optical random features.
Survey of kernels, RKHS, and their applications in machine learning.
Choosing the most adequate kernel is crucial in many Machine Learning applications. Gaussian Process is a state-of-the-art technique for regression and classification that heavily relies on a kernel function. However, in the Gaussian Process literature, kernels have usually been either ad hoc designed, selected from a …
Kernels are powerful and versatile tools in machine learning and statistics. Although the notion of universal kernels and characteristic kernels has been studied, kernel selection still greatly influences the empirical performance. While learning the kernel in a data driven way has been investigated, in this paper we e…
Kernel-Gradient Drifting improves generative modeling for non-Euclidean data.
Kernel for Lévy rough paths derived from PDE system.
New kernel handles irregularly-spaced multivariate time series.
New sparse GP model learns compositional kernels efficiently.
Recently, non-stationary spectral kernels have drawn much attention, owing to its powerful feature representation ability in revealing long-range correlations and input-dependent characteristics. However, non-stationary spectral kernels are still shallow models, thus they are deficient to learn both hierarchical featur…
IDK improves anomaly detection for points and groups without explicit learning.
New research optimizes HSIC estimation rate for translation-invariant kernels.
Proposes a Gaussian process for graph signals using adaptive spectral kernels.
Kernel methods are one of the mainstays of machine learning, but the problem of kernel learning remains challenging, with only a few heuristics and very little theory. This is of particular importance in methods based on estimation of kernel mean embeddings of probability measures. For characteristic kernels, which inc…
In this paper we continue our study on the moduli spaces of flat G-bundles, for any semi-simple Lie group G, over a Riemann surface by using heat kernel and Reidemeister torsion. Formulas for intersection numbers on the moduli spaces over a Riemann surface with several boundary components, over non-orientable Riemann s…
We provide a unifying framework linking two classes of statistics used in two-sample and independence testing: on the one hand, the energy distances and distance covariances from the statistics literature; on the other, distances between embeddings of distributions to reproducing kernel Hilbert spaces (RKHS), as establ…
Paper introduces MRCs that minimize worst-case 0-1 loss, providing tight performance guarantees.
We study the construction of coresets for kernel density estimates. That is we show how to approximate the kernel density estimate described by a large point set with another kernel density estimate with a much smaller point set. For characteristic kernels (including Gaussian and Laplace kernels), our approximation pre…
Permutation-valued features arise in a variety of applications, either in a direct way when preferences are elicited over a collection of items, or an indirect way in which numerical ratings are converted to a ranking. To date, there has been relatively limited study of regression, classification, and testing problems …
This paper proposes a new method for automatically selecting the optimal kernel bandwidth in density estimation.
We consider the problem of learning low-dimensional representations for large-scale Markov chains. We formulate the task of representation learning as that of mapping the state space of the model to a low-dimensional state space, called the kernel space. The kernel space contains a set of meta states which are desired …
Clinical measurements collected over time are naturally represented as multivariate time series (MTS), which often contain missing data. An autoencoder can learn low dimensional vectorial representations of MTS that preserve important data characteristics, but cannot deal explicitly with missing data. In this work, we …
Given two sets of independent samples from unknown distributions and , a two-sample test decides whether to reject the null hypothesis that . Recent attention has focused on kernel two-sample tests as the test statistics are easy to compute, converge fast, and have low bias with their finite sample estimate…
The reproducing kernel Hilbert space (RKHS) embedding of distributions offers a general and flexible framework for testing problems in arbitrary domains and has attracted considerable amount of attention in recent years. To gain insights into their operating characteristics, we study here the statistical performance of…
Do two data samples come from different distributions? Recent studies of this fundamental problem focused on embedding probability distributions into sufficiently rich characteristic Reproducing Kernel Hilbert Spaces (RKHSs), to compare distributions by the distance between their embeddings. We show that Regularized Ma…
Regularizes -divergences with MMD to analyze Wasserstein flows.
Researchers develop flexible kernels for biological sequences with guaranteed reliability.
We introduce a data-driven order reduction method for nonlinear control systems, drawing on recent progress in machine learning and statistical dimensionality reduction. The method rests on the assumption that the nonlinear system behaves linearly when lifted into a high (or infinite) dimensional feature space where ba…