Inference in popular nonparametric Bayesian models typically relies on sampling or other approximations. This paper presents a general methodology for constructing novel tractable nonparametric Bayesian methods by applying the kernel trick to inference in a parametric Bayesian model. For example, Gaussian process regre…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper corrects the proof of the Theorem 2 from the Gower's paper \cite[page 5]{Gower:1982} as well as corrects the Theorem 7 from Gower's paper \cite{Gower:1986}. The first correction is needed in order to establish the existence of the kernel function used commonly in the kernel trick e.g. for -means clusterin…
When solving data analysis problems it is important to integrate prior knowledge and/or structural invariances. This paper contributes by a novel framework for incorporating algebraic invariance structure into kernels. In particular, we show that algebraic properties such as sign symmetries in data, phase independence,…
KMRCD detects outliers in non-elliptical data using kernel trick.
Kernel method is a very powerful tool in machine learning. The trick of kernel has been effectively and extensively applied in many areas of machine learning, such as support vector machine (SVM) and kernel principal component analysis (kernel PCA). Kernel trick is to define a kernel function which relies on the inner-…
A fast algorithm speeds up training of pairwise kernels.
Kernel methods form a powerful, versatile, and theoretically-grounded unifying framework to solve nonlinear problems in signal processing and machine learning. The standard approach relies on the kernel trick to perform pairwise evaluations of a kernel function, which leads to scalability issues for large datasets due …
Study proposes a new metric for comparing Gaussian mixtures in RKHS.
Non-linear kernel methods can be approximated by fast linear ones using suitable explicit feature maps allowing their application to large scale problems. We investigate how convolution kernels for structured data are composed from base kernels and construct corresponding feature maps. On this basis we propose exact an…
We introduce a family of pairwise stochastic gradient estimators for gradients of expectations, which are related to the log-derivative trick, but involve pairwise interactions between samples. The simplest example of our new estimator, dubbed the fundamental trick estimator, is shown to arise from either a) introducin…
This paper speeds up kernel methods using sparsified Gaussian sketches.
We extend the herding algorithm to continuous spaces by using the kernel trick. The resulting "kernel herding" algorithm is an infinite memory deterministic process that learns to approximate a PDF with a collection of samples. We show that kernel herding decreases the error of expectations of functions in the Hilbert …
Deep neural networks for structured prediction using kernel-induced losses.
Kernel SIVI improves variational inference by avoiding lower-level optimization.
Kronecker product kernel provides the standard approach in the kernel methods literature for learning from graph data, where edges are labeled and both start and end vertices have their own feature representations. The methods allow generalization to such new edges, whose start and end vertices do not appear in the tra…
Paper presents a fast and adaptive filter for SI suppression in full-duplex transceivers.
Paper presents a new way to estimate model changes without full model evaluation.
Kernel methods form a theoretically-grounded, powerful and versatile framework to solve nonlinear problems in signal processing and machine learning. The standard approach relies on the \emph{kernel trick} to perform pairwise evaluations of a kernel function, leading to scalability issues for large datasets due to its …
A new method quickly identifies key variables and interactions.
Maps embed manifolds using heat kernels of connection Laplacian.
In presence of sparse noise we propose kernel regression for predicting output vectors which are smooth over a given graph. Sparse noise models the training outputs being corrupted either with missing samples or large perturbations. The presence of sparse noise is handled using appropriate use of -norm along-wi…
Kernelized cumulants improve statistical analysis in high-dimensional spaces.
4-manifolds with nonnegative sectional curvature are area-extremal.
Unified view on random walk and Weisfeiler-Leman kernels, improving accuracy.
Recent studies show overparameterized neural networks behave like convex systems.
Derives a primal-dual MLSVD formulation for multilinear data.
Paper studies kernel hyperparameters for clustering, proposing an efficient search method.
Variational Auto-Encoders (VAEs) have become very popular techniques to perform inference and learning in latent variable models as they allow us to leverage the rich representational power of neural networks to obtain flexible approximations of the posterior of latent variables as well as tight evidence lower bounds (…
A new framework optimizes fMRI and behavioral data for better understanding of Autism.
This paper presents a robust matrix elastic net based canonical correlation analysis (RMEN-CCA) for multiple view unsupervised learning problems, which emphasizes the combination of CCA and the robust matrix elastic net (RMEN) used as coupled feature selection. The RMEN-CCA leverages the strength of the RMEN to distill…
We present a new method which generalizes subspace learning based on eigenvalue and generalized eigenvalue problems. This method, Roweis Discriminant Analysis (RDA), is named after Sam Roweis to whom the field of subspace learning owes significantly. RDA is a family of infinite number of algorithms where Principal Comp…
We present a new framework for online Least Squares algorithms for nonlinear modeling in RKH spaces (RKHS). Instead of implicitly mapping the data to a RKHS (e.g., kernel trick), we map the data to a finite dimensional Euclidean space, using random features of the kernel's Fourier transform. The advantage is that, the …
The Wasserstein distance is a powerful metric based on the theory of optimal transport. It gives a natural measure of the distance between two distributions with a wide range of applications. In contrast to a number of the common divergences on distributions such as Kullback-Leibler or Jensen-Shannon, it is (weakly) co…
Sketching accelerates structured prediction methods for large datasets.
A new generator uses kernel distance to avoid GAN weaknesses.
Enhances feature augmentation for high-dimensional learning.
Enhances GPLVM for multi-view data with scalable latent representation learning.
Galerkin method outperforms graph-based methods in spectral decompositions.
New method distinguishes data noise from GP uncertainty.
Paper bridges VAEs and KDEs for more flexible posterior estimation.
Study of skateboard flips as continuous curves in group.
Nash's theorem proved with Günther's trick
Explains Conway's tangle trick and its mathematical origins.
We connect shift-invariant characteristic kernels to infinitely divisible distributions on . Characteristic kernels play an important role in machine learning applications with their kernel means to distinguish any two probability measures. The contribution of this paper is two-fold. First, we show, usi…
Unified framework for gradient estimation in combinatorial spaces.
Paper shows SVMs can interpolate data in various settings.
In this contribution, we propose a new computationally efficient method to combine Variational Inference (VI) with Markov Chain Monte Carlo (MCMC). This approach can be used with generic MCMC kernels, but is especially well suited to \textit{MetFlow}, a novel family of MCMC algorithms we introduce, in which proposals a…
New method uses path signatures for efficient likelihood estimation in time-series data.