The study extends kernel universality to Riemannian symmetric spaces.
problem Understanding kernel universality in non-Euclidean domains.
method Harmonic analysis on Riemannian symmetric spaces.
result Proves universality of recent kernels on Riemannian symmetric spaces.
Kernel methods have been widely applied to machine learning and other questions of approximating an unknown function from its finite sample data. To ensure arbitrary accuracy of such approximation, various denseness conditions are imposed on the selected kernel. This note contributes to the study of universal, characte…
A Hilbert space embedding for probability measures has recently been proposed, wherein any probability measure is represented as a mean element in a reproducing kernel Hilbert space (RKHS). Such an embedding has found applications in homogeneity testing, independence testing, dimensionality reduction, etc., with the re…
Quantum kernels can be efficiently embedded into classical feature spaces.
problem Can all quantum kernels be efficiently embedded into classical feature spaces?
method Invoking computational universality and using techniques like random Fourier features, the authors show that certain classes of quantum kernels can be efficiently embedded.
result For shift-invariant and composition kernels, embedding quantum kernels are universal and efficient.
We investigate iterated compositions of weighted sums of Gaussian kernels and provide an interpretation of the construction that shows some similarities with the architectures of deep neural networks. On the theoretical side, we show that these kernels are universal and that SVMs using these kernels are universally con…
The accuracy and complexity of kernel learning algorithms is determined by the set of kernels over which it is able to optimize. An ideal set of kernels should: admit a linear parameterization (tractability); be dense in the set of all kernels (accuracy); and every member should be universal so that the hypothesis spac…
A universal collection of 4 invariants improves neural network accuracy for molecular dynamics.
problem Improving accuracy of neural networks in molecular dynamics.
method Developed a universal collection of 4 smooth scalar invariants on M(3) x M(3) and evaluated their effectiveness in a PONITA neural network architecture.
result Using a universal collection of invariants significantly improves neural network accuracy.
ULFS-KDPE estimates parameters efficiently without influence functions.
problem Estimating pathwise differentiable parameters in nonparametric models.
method Kernel debiased plug-in estimator based on universal least favorable submodel.
result Semiparametric efficiency achieved without influence function derivation.
New method for reducing dimensions of distributional data.
problem Nonlinear sufficient dimension reduction for distribution-on-distribution regression.
method Building universal kernels on metric spaces to characterize conditional independence.
result Method outperforms competing methods in synthetic and real data applications.
Statistical machine learning plays an important role in modern statistics and computer science. One main goal of statistical machine learning is to provide universally consistent algorithms, i.e., the estimator converges in probability or in some stronger sense to the Bayes risk or to the Bayes decision function. Kerne…
We prove the statistical consistency of kernel Partial Least Squares Regression applied to a bounded regression learning problem on a reproducing kernel Hilbert space. Partial Least Squares stands out of well-known classical approaches as e.g. Ridge Regression or Principal Components Regression, as it is not defined as…
Quantum kernels show no advantage in stock return prediction, but differ in stability metrics.
problem Determining if quantum kernels improve stock return prediction.
method Controlled horse race on Chinese A-share market with identical training subsamples and tuning budgets.
result Quantum kernels do not outperform classical RBF controls in cross-sectional stock return prediction.
Functional input neural networks approximate continuous functions on weighted spaces.
problem Approximating continuous functions on infinite-dimensional weighted spaces.
method Additive family mapping, non-linear activation, linear readouts, Stone-Weierstrass theorem.
result Global universal approximation of continuous functions on weighted spaces.
Modeling videos and image-sets as linear subspaces has proven beneficial for many visual recognition tasks. However, it also incurs challenges arising from the fact that linear subspaces do not obey Euclidean geometry, but lie on a special type of Riemannian manifolds known as Grassmannian. To leverage the techniques d…
We establish upper bounds for the minimal number of hidden units for which a binary stochastic feedforward network with sigmoid activation probabilities and a single hidden layer is a universal approximator of Markov kernels. We show that each possible probabilistic assignment of the states of n output units, given t…
This study approximates neural network features for modeling relations and attention mechanisms.
problem Approximating neural network features for modeling relations and attention mechanisms.
method Analyzes inner products of multi-layer perceptrons for universal approximation of symmetric and asymmetric relation functions.
result Universal approximation of relation functions and attention mechanisms using inner products of neural networks.
The massive amount of available data potentially used to discover patters in machine learning is a challenge for kernel based algorithms with respect to runtime and storage capacities. Local approaches might help to relieve these issues. From a statistical point of view local approaches allow additionally to deal with …
New insights into model robustness for random features and NTK models.
problem Understanding and distinguishing robustness in machine learning models.
method Analyzing empirical risk minimization in random features and NTK models.
result Random features models are not robust under any degree of over-parameterization, even when satisfying the universal law of robustness.
Kernel mean embeddings have recently attracted the attention of the machine learning community. They map measures μ from some set M to functions in a reproducing kernel Hilbert space (RKHS) with kernel k. The RKHS distance of two mapped measures is a semi-metric dk over M. We study three questions. (I) For a…
This paper shows universality in spectrum behavior for random inner-product kernel matrices in polynomial regime.
problem Understanding spectrum behavior of random inner-product kernel matrices in polynomial regime.
method Analyzing matrices formed by a nonlinear function applied entrywise to a sample-covariance matrix, considering i.i.d. entries with all finite moments.
result The spectrum of random inner-product kernel matrices is universally described by the free convolution of the semicircular and Marčenko-Pastur distributions, with relative weights given by expanding the nonlinear function in the Hermite basis.
New algorithm reduces kernel optimization complexity.
problem Efficiently learn semiseparable kernels for robust machine learning.
method Proposed a SVD-QCQP primal-dual algorithm for minimax optimization.
result Significant improvement in accuracy over non-convex approaches.
HyBO optimizes hybrid structures using diffusion kernels.
problem Optimizing complex interactions between discrete and continuous variables.
method HyBO uses diffusion kernels over hybrid spaces with additive kernel formulation.
result HyBO significantly outperforms state-of-the-art methods on real-world benchmarks.
From the uniformization theorem, we know that every Riemann surface has a simply-connected covering space. Moreover, there are only three simply-connected Riemann surfaces: the sphere, the Euclidean plane, and the hyperbolic plane. In this paper, we collect the known heat kernels, or Green's functions, for these three …
Transformers are explained as infinite-dimensional kernel machines.
problem Understanding the mechanics of Transformers in AI.
method Characterized Transformers' attention mechanism as a kernel learning method on Banach spaces.
result Transformer's kernel has infinite feature dimension and can learn any binary non-Mercer reproducing kernel Banach space pair.
We construct a canonical correspondence from a wide class of reproducing kernels on infinite-dimensional Hermitian vector bundles to linear connections on these bundles. The linear connection in question is obtained through a pull-back operation involving the tautological universal bundle and the classifying morphism o…
Paper introduces a new kernel model for PSD-valued functions with theoretical guarantees and applications.
problem Enforcing positive semi-definiteness (PSD) in function models with good performance and theoretical guarantees.
method Kernel sum-of-squares model for PSD-valued functions, extending previous models for non-negative scalar functions.
result The model constitutes a universal approximator of PSD functions and can represent any smooth and strongly convex function.
The paper introduces new KMEs to capture stochastic process filtrations.
problem Missing filtration information in stochastic processes.
method Higher order kernel mean embeddings (KMEs) conditioned on filtrations.
result Consistent estimators and tests for filtration-sensitive information.
Study the cost of overfitting in noisy KRR models.
problem Cost of overfitting in noisy kernel ridge regression.
method An agnostic view of overfitting cost as a function of sample size for any target function, using Gaussian universality ansatz and task eigenstructure.
result Characterization of benign, tempered, and catastrophic overfitting.
Estimates KRR risk from training data for various kernels and hyperparameters.
problem Predicting the generalization error of Kernel Ridge Regression.
method Introduces SCT and KARE to approximate KRR risk from training data.
result KARE provides an excellent approximation of KRR risk and helps select good kernels.
A new method for distribution regression using sliced Wasserstein distance.
problem Learning functions over spaces of probabilities.
method Proposes an OT-based estimator using the Sliced Wasserstein distance.
result Proves universal consistency and excess risk bounds for the proposed estimator.
NOs can learn any finite collection of classes in functional data.
problem Learning finite collections of classes in infinite-dimensional spaces.
method Proved sample-based neural operators can learn any finite collection of classes in an infinite-dimensional reproducing kernel Hilbert space.
result NOs can learn any finite collection of classes in an infinite-dimensional reproducing kernel Hilbert space, even when the classes are not convex or connected.
Improved Strichartz estimates for Schrödinger equation on manifolds with nonpositive curvature.
problem Improving Strichartz estimates for Schrödinger equation on compact manifolds with nonpositive sectional curvature.
method Improved global kernel estimates for microlocalized operators exploiting geometric assumptions.
result No-loss LtpLxq-estimates on intervals of length logλ⋅λ−1 for all admissible pairs (p,q). Maximum mean discrepancy (MMD), also called energy distance or N-distance in statistics and Hilbert-Schmidt independence criterion (HSIC), specifically distance covariance in statistics, are among the most popular and successful approaches to quantify the difference and independence of random variables, respectively. T…
The generalization properties of Gaussian processes depend heavily on the choice of kernel, and this choice remains a dark art. We present the Neural Kernel Network (NKN), a flexible family of kernels represented by a neural network. The NKN architecture is based on the composition rules for kernels, so that each unit …
Study shows convergence of Bergman kernels on covering spaces of Kähler manifolds.
problem Analyzing convergence of Bergman kernels on covering spaces of Kähler manifolds.
method Proving convergence of Bergman kernels and L2-Hodge numbers on a tower of coverings. result Sections of canonical line bundles give rise to immersions into projective spaces.
Conditional diffusion models can approximate target distributions well with Gaussian-mixture reverse kernels.
problem Approximating target distributions in conditional diffusion models.
method Using finite Gaussian mixtures with ReLU-network logits as reverse kernels, reducing the problem to static conditional density approximation.
result The resulting neural reverse-kernel class is dense in conditional KL divergence under exact terminal matching.
We present a universal algorithm for online trading in Stock Market which performs asymptotically at least as good as any stationary trading strategy that computes the investment at each step using a fixed function of the side information that belongs to a given RKHS (Reproducing Kernel Hilbert Space). Using a universa…
The paper analyzes deep ReLU CNNs' approximation properties in 2D space.
problem Establishing L2 approximation properties for deep ReLU CNNs. method Analysis based on decomposition theorem for convolutional kernels, properties of ReLU activation, and connections with one-hidden-layer ReLU NNs.
result Universal approximation theorem for deep ReLU CNNs with classic structure.
Kernel methods are ubiquitous tools in machine learning. However, there is often little reason for the common practice of selecting a kernel a priori. Even if a universal approximating kernel is selected, the quality of the finite sample estimator may be greatly affected by the choice of kernel. Furthermore, when direc…
New algorithm optimizes tessellated kernels for larger datasets and improved performance.
problem Limited accuracy and complexity in machine learning algorithms based on kernel optimization.
method 2-step algorithm for optimizing tessellated kernels, scaling to 10,000 data points and extending to regression.
result Significant improvement in performance over Neural Nets and SimpleMKL with similar computation time.
Survey of kernels, RKHS, and their applications in machine learning.
problem Understanding kernels and their applications in machine learning.
method Review of historical context, mathematical definitions, and practical applications of kernels.
result Comprehensive overview of kernels, RKHS, and their applications.
A new kernel for probability measures based on optimal transport.
problem Efficiently comparing and modeling distributions.
method Kernel over probability measures using regularized optimal transport and Hilbertian embedding.
result The proposed kernel enables Gaussian process modeling on distributions with theoretical and computational advantages.
Geometric theory connects machine learning classifiers to differential geometry.
problem Classifying data points in machine learning.
method Mapping binary classification to vector bundles and differential geometry.
result Harmonic interpolation solves RKHS interpolation problems.
Paper introduces RKHM and KME for richer data analysis.
problem Lack of rich data structures in kernel methods.
method Proposes RKHM and KME for functional data analysis.
result RKHM captures structural properties in functional data.
Kernels are powerful and versatile tools in machine learning and statistics. Although the notion of universal kernels and characteristic kernels has been studied, kernel selection still greatly influences the empirical performance. While learning the kernel in a data driven way has been investigated, in this paper we e…
Study shows that ridgeless Gaussian kernel regression overfits even with varying bandwidth or dimensionality.
problem Analyzing overfitting in Gaussian kernel ridgeless regression with varying bandwidth or dimensionality.
method Examined the behavior of minimum norm interpolating solutions for fixed and increasing dimensions under varying bandwidth and sample size.
result Ridgeless solutions are never consistent and can be worse than null predictor with large enough noise, even with varying bandwidth or dimensionality.
We establish certain Gaussian type upper bound for the heat kernel of the conjugate heat equation associated with 3 dimensional ancient κ solutions to the Ricci flow. As an application, using the W entropy associated with the heat kernel, we give a different and shorter proof of Perelman's classification of backwar…
Kernel-based Bayesian filter for nonlinear systems using infinite-dimensional operators.
problem Modeling and predicting nonlinear dynamical systems.
method Functional Bayesian perspective, reproducing kernel Hilbert space, Gaussian kernel.
result Effective approximation and accurate results for nonlinear systems.