Inference in popular nonparametric Bayesian models typically relies on sampling or other approximations. This paper presents a general methodology for constructing novel tractable nonparametric Bayesian methods by applying the kernel trick to inference in a parametric Bayesian model. For example, Gaussian process regre…
No-trick kernel adaptive filtering uses deterministic features for scalability and robustness.
problem Scalability issues in kernel methods for large datasets.
method Deterministic feature-map construction using polynomial-exact solutions.
result Deterministic features outperform random Fourier features in performance and scalability.
This paper corrects the proof of the Theorem 2 from the Gower's paper \cite[page 5]{Gower:1982} as well as corrects the Theorem 7 from Gower's paper \cite{Gower:1986}. The first correction is needed in order to establish the existence of the kernel function used commonly in the kernel trick e.g. for k-means clusterin…
When solving data analysis problems it is important to integrate prior knowledge and/or structural invariances. This paper contributes by a novel framework for incorporating algebraic invariance structure into kernels. In particular, we show that algebraic properties such as sign symmetries in data, phase independence,…
KMRCD detects outliers in non-elliptical data using kernel trick.
problem Outlier detection in non-elliptical data.
method KMRCD estimator that uses kernel trick to compute robust covariance matrix in a feature space.
result KMRCD performs well in simulations and real-life data.
Kernel method is a very powerful tool in machine learning. The trick of kernel has been effectively and extensively applied in many areas of machine learning, such as support vector machine (SVM) and kernel principal component analysis (kernel PCA). Kernel trick is to define a kernel function which relies on the inner-…
A fast algorithm speeds up training of pairwise kernels.
problem Training pairwise kernels efficiently for large datasets.
method Generalized vec trick for Kronecker product kernels.
result Pairwise kernels can be expressed as sums of Kronecker products.
Study proposes a new metric for comparing Gaussian mixtures in RKHS.
problem Comparing complex multimodal densities in RKHS.
method Wasserstein-type metric for kernel Gaussian mixtures.
result Enhanced capability to model multimodal densities.
Non-linear kernel methods can be approximated by fast linear ones using suitable explicit feature maps allowing their application to large scale problems. We investigate how convolution kernels for structured data are composed from base kernels and construct corresponding feature maps. On this basis we propose exact an…
We introduce a family of pairwise stochastic gradient estimators for gradients of expectations, which are related to the log-derivative trick, but involve pairwise interactions between samples. The simplest example of our new estimator, dubbed the fundamental trick estimator, is shown to arise from either a) introducin…
This paper speeds up kernel methods using sparsified Gaussian sketches.
problem Kernel methods' computational limitations.
method Sparsified Gaussian sketches for kernel methods.
result Efficient time and space savings for kernel methods.
We extend the herding algorithm to continuous spaces by using the kernel trick. The resulting "kernel herding" algorithm is an infinite memory deterministic process that learns to approximate a PDF with a collection of samples. We show that kernel herding decreases the error of expectations of functions in the Hilbert …
Deep neural networks for structured prediction using kernel-induced losses.
problem Structured prediction tasks for images and texts.
method Designing a novel family of deep neural architectures that predict in a finite-dimensional subspace derived from the kernel-induced loss.
result Gradient descent algorithms can be used for structured prediction with deep neural networks.
Kernel SIVI improves variational inference by avoiding lower-level optimization.
problem Intractable densities in semi-implicit variational distributions.
method Kernel SIVI-SM uses a minimax formulation and kernel tricks to avoid lower-level optimization.
result Kernel Stein discrepancy (KSD) objective is computable and leads to convergence guarantees.
Kronecker product kernel provides the standard approach in the kernel methods literature for learning from graph data, where edges are labeled and both start and end vertices have their own feature representations. The methods allow generalization to such new edges, whose start and end vertices do not appear in the tra…
Paper presents a fast and adaptive filter for SI suppression in full-duplex transceivers.
problem Self-interference suppression in full-duplex transceivers with nonlinearity.
method Adaptive projected subgradient method (APSM) in a reproducing kernel Hilbert space (RKHS).
result The proposed method achieves favorable digital SIC performance compared to benchmarks.
Paper presents a new way to estimate model changes without full model evaluation.
problem Efficiently estimating changes in model parameters and outputs due to data point removal.
method Dual representation of influence functions for linearizable models, reducing computational complexity.
result The dual representation can be an efficient alternative to original influence functions, especially for large models.
A new method quickly identifies key variables and interactions.
problem Identifying key variables and interactions in high-dimensional data.
method Kernel trick for sparse orthogonal decomposition in O(# covariates) time.
result Outperforms existing methods for large, high-dimensional data sets.
Improved ITL descriptors using explicit inner product spaces for scalable systems.
problem Scalability issues in ITL due to high computational complexity.
method Explicit inner product space (EIPS) kernels for ITL, leveraging data-independent basis.
result Superior performance of EIPS-ITL estimators and combined NT-KAF using EIPS-ITL cost functions.
Maps embed manifolds using heat kernels of connection Laplacian.
problem Embedding manifolds in Euclidean space.
method Using heat kernels of the connection Laplacian and truncated heat kernels.
result Maps can be made arbitrarily close to isometries.
In presence of sparse noise we propose kernel regression for predicting output vectors which are smooth over a given graph. Sparse noise models the training outputs being corrupted either with missing samples or large perturbations. The presence of sparse noise is handled using appropriate use of ℓ1-norm along-wi…
Kernelized cumulants improve statistical analysis in high-dimensional spaces.
problem Statistical analysis in high-dimensional spaces with low variance estimators.
method Extending cumulants to RKHS using tensor algebra and kernel trick.
result Kernelized cumulants provide new all-purpose statistics with computational tractability.
4-manifolds with nonnegative sectional curvature are area-extremal.
problem Finding extremal properties of 4-manifolds with curvature constraints.
method Analyzing sections in the kernel of a twisted Dirac operator and using the Finsler--Thorpe trick.
result Large classes of compact 4-manifolds are area-extremal.
Unified view on random walk and Weisfeiler-Leman kernels, improving accuracy.
problem Improving graph kernel methods for better classification accuracy.
method Define and analyze walk-based node refinement methods, relate to Weisfeiler-Leman test, and introduce new walk-based kernels.
result Walk-based kernels are as expressive as Weisfeiler-Leman subtree kernel but support non-strict neighborhood comparison.
Recent studies show overparameterized neural networks behave like convex systems.
problem Understanding the behavior of overparameterized neural networks.
method Analysis of two-layer neural networks, focusing on restricted settings and neural tangent kernel space.
result Overparameterized neural networks behave like convex systems under certain conditions.
Derives a primal-dual MLSVD formulation for multilinear data.
problem Efficiently decompose multilinear data for signal analysis and deep learning.
method Kernelizable primal-dual formulation of MLSVD.
result Derives a new MLSVD formulation with computational advantages.
New model for high rank matrix completion with online and batch methods.
problem Matrix completion for high rank matrices with latent structure.
method Kernel trick to map data into a high dimensional feature space, explicit parametrization of low dimensional subspace, online fitting procedure.
result Online method can handle streaming data and adapt to non-stationary latent structure.
Paper studies kernel hyperparameters for clustering, proposing an efficient search method.
problem Challenges in tuning kernel parameters for clustering, especially for RBF kernels.
method Derives a lower bound for RBF kernel parameters, proposes an efficient hyperparameter search algorithm.
result Proposes an efficient algorithm for hyperparameter search in kernel clustering, improving upon grid search.
The paper recovers missing data entries of high-rank matrices using polynomial polynomials.
problem Recovering missing entries of high-rank matrices with low intrinsic dimension.
method Developed a new polynomial matrix completion method using the kernel trick and relaxation of rank objective.
result Identified complete matrix of minimum intrinsic dimension by minimizing rank in high-dimensional feature space.
Variational Auto-Encoders (VAEs) have become very popular techniques to perform inference and learning in latent variable models as they allow us to leverage the rich representational power of neural networks to obtain flexible approximations of the posterior of latent variables as well as tight evidence lower bounds (…
A new framework optimizes fMRI and behavioral data for better understanding of Autism.
problem Linking complex fMRI data to behavioral measures is challenging.
method Coupled manifold optimization framework projecting fMRI onto a shared manifold and mapping to behavioral measures.
result Framework outperforms traditional methods in predicting clinical severity of Autism.
This paper presents a robust matrix elastic net based canonical correlation analysis (RMEN-CCA) for multiple view unsupervised learning problems, which emphasizes the combination of CCA and the robust matrix elastic net (RMEN) used as coupled feature selection. The RMEN-CCA leverages the strength of the RMEN to distill…
We present a new method which generalizes subspace learning based on eigenvalue and generalized eigenvalue problems. This method, Roweis Discriminant Analysis (RDA), is named after Sam Roweis to whom the field of subspace learning owes significantly. RDA is a family of infinite number of algorithms where Principal Comp…
We present a new framework for online Least Squares algorithms for nonlinear modeling in RKH spaces (RKHS). Instead of implicitly mapping the data to a RKHS (e.g., kernel trick), we map the data to a finite dimensional Euclidean space, using random features of the kernel's Fourier transform. The advantage is that, the …
The Wasserstein distance is a powerful metric based on the theory of optimal transport. It gives a natural measure of the distance between two distributions with a wide range of applications. In contrast to a number of the common divergences on distributions such as Kullback-Leibler or Jensen-Shannon, it is (weakly) co…
Sketching accelerates structured prediction methods for large datasets.
problem Scaling surrogate kernel methods for large datasets.
method Sketching approximations applied to input and output feature maps.
result Achieves close-to-optimal rates with reduced sketch size.
A new generator uses kernel distance to avoid GAN weaknesses.
problem Stability and mode collapse in GANs and autoencoders.
method LCW generator (Latent Cramer-Wold generator) using kernel distance.
result Very competitive FID values.
Enhances feature augmentation for high-dimensional learning.
problem Correlated high-dimensional measurements require dimensionality reduction.
method Augment features with factors extracted from design matrices and their transformations.
result Significantly weakens correlations between input variables, improving interpretability and numerical stability.
Enhances GPLVM for multi-view data with scalable latent representation learning.
problem Limited kernel expressiveness and computational inefficiency in multi-view GPLVM.
method Introduces a new duality between spectral density and kernel function, uses NG-SM kernel, and applies random Fourier feature approximation for scalability.
result Consistently outperforms state-of-the-art models in learning meaningful latent representations across diverse datasets.
Galerkin method outperforms graph-based methods in spectral decompositions.
problem Improving spectral decomposition methods in machine learning.
method Restricting study to a small set of test functions using the Galerkin method.
result Statistical and computational superiority of Galerkin method over graph-based approaches.
New method distinguishes data noise from GP uncertainty.
problem Uncertainty in kernel regression with non-Gaussian noise.
method Wiener chaos expansions for non-Gaussian noise.
result Can distinguish aleatoric from epistemic uncertainty.
Paper bridges VAEs and KDEs for more flexible posterior estimation.
problem Limitations of Gaussian latent space in VAEs and challenges in KL-divergence estimation.
method Approximate posterior with KDEs and derive upper bound of KL-divergence in ELBO.
result Epanechnikov kernel minimizes KL-divergence upper bound asymptotically.
Study of skateboard flips as continuous curves in SO(3) group.
problem Characterize skateboard flip tricks as continuous motions.
method Model flips as curves in SO(3), analyze lifts to S3, derive formulas. result There are only four distinct flip tricks up to continuous deformation.
Kernelized PCovR reveals structure-property relations in chemistry and materials.
problem Understanding structure-property relations in complex systems.
method Kernel Principal Covariates Regression (kernel PCovR) with sparsification.
result Kernelized PCovR effectively reveals and predicts structure-property relations.
Nash's theorem proved with Günther's trick
problem Proving Nash's smooth embedding theorem
method Using Günther's trick
result Nash's theorem proved
Explains Conway's tangle trick and its mathematical origins.
problem Understanding the relationship between braids and elliptic curves.
method Discusses the tangle trick, its mathematical underpinnings, and historical context.
result Establishes the connection between braids and elliptic curves.
We connect shift-invariant characteristic kernels to infinitely divisible distributions on Rd. Characteristic kernels play an important role in machine learning applications with their kernel means to distinguish any two probability measures. The contribution of this paper is two-fold. First, we show, usi…
Unified framework for gradient estimation in combinatorial spaces.
problem Scaling relaxed gradient estimators to large combinatorial distributions.
method Introducing stochastic softmax tricks within the perturbation model framework.
result Stochastic softmax tricks improve model performance and discover more latent structure.