Improved text classification performance through conformal transformations of kernels.
problem Text document categorization in high-dimensional spaces.
method Introduced new Gaussian Cosine kernel and two conformal transformations.
result Conformal transformations significantly improve kernel performance, especially for sub-optimal kernels.
Transformers are explained as infinite-dimensional kernel machines.
problem Understanding the mechanics of Transformers in AI.
method Characterized Transformers' attention mechanism as a kernel learning method on Banach spaces.
result Transformer's kernel has infinite feature dimension and can learn any binary non-Mercer reproducing kernel Banach space pair.
Proposes a new feature preprocessing method using kernel density integral transformation.
problem Feature preprocessing for tabular data in machine learning and statistics.
method Kernel density integral transformation as a drop-in replacement or improved alternative to min-max scaling and quantile transformation.
result Frequently outperforms min-max scaling and quantile transformation with hyperparameter tuning.
KITT uses transformers to quickly recommend kernels for GP models.
problem Kernel selection for high-dimensional GP regression models.
method Transformer-based architecture for generating kernel recommendations.
result KITT selects kernels that perform well on various regression benchmarks.
Learning the kernel functions used in kernel methods has been a vastly explored area in machine learning. It is now widely accepted that to obtain 'good' performance, learning a kernel function is the key challenge. In this work we focus on learning kernel representations for structured regression. We propose use of po…
Transformers improve with Fourier integral attentions.
problem Inefficiency of dot-product attention in capturing feature dependencies.
method Interpreted attention as kernel regression, proposed FourierFormer with generalized Fourier integral kernels.
result FourierFormer achieves better accuracy and reduces redundancy.
Estimates heat kernel gradients on fractal-like cable systems.
problem Bounding gradients of heat kernels on complex fractal structures.
method Pointwise upper estimates for heat kernel gradients.
result Derives Lp-boundedness of quasi-Riesz transforms. Transformer is a powerful architecture that achieves superior performance on various sequence learning tasks, including neural machine translation, language understanding, and sequence prediction. At the core of the Transformer is the attention mechanism, which concurrently processes all inputs in the streams. In this …
Distance-based tests, also called "energy statistics", are leading methods for two-sample and independence tests from the statistics community. Kernel-based tests, developed from "kernel mean embeddings", are leading methods for two-sample and independence tests from the machine learning community. A fixed-point transf…
New interpretation of attention in Transformers and Graph Attention Networks.
problem Understanding and improving attention mechanisms in deep learning models.
method Decomposed attention into a kernel and a normalization term; generalized the kernel function and norm.
result Generalized attention leads to better performance on various tasks.
New STKR estimators use unlabeled data for smoother function learning.
problem Leveraging unlabeled data for smoother function learning.
method Spectrally transformed kernel regression (STKR) with scalable implementations.
result STKR can learn any sufficiently smooth function.
New RFs reduce kernel approximation variance and improve Transformer performance.
problem Efficient approximation of Gaussian and softmax kernels for kernel methods and Transformers.
method Parameterized, positive, non-trigonometric RFs optimized for variance reduction.
result Significant variance reduction in practice, outperforming previous methods.
Kernel models learn low-dimensional predictive subspaces from input data.
problem Learning effective feature transformations in kernel models.
method Study of a compositional kernel ridge regression model.
result Global minimizers of the objective function identify the subspace with high probability.
Efficiently accelerates attention calculation for Transformers with relative positional encoding.
problem Quadratic complexity of attention in long sequences.
method Kernelized attention with Fast Fourier Transform (FFT) for RPE.
result Achieves O(n log n) time complexity, mitigates training instability, and outperforms other models.
Heat kernel estimates on manifolds with mixed boundary conditions.
problem Estimating heat kernels on manifolds with ends and mixed boundary conditions.
method Global harmonic function construction and h-transform technique. result Two-sided heat kernel estimates for Riemannian manifolds with mixed boundary conditions.
Transformers with linear space and time complexity for accurate attention estimation.
problem Efficiently estimating attention in large-scale tasks without relying on priors.
method Performers use Fast Attention Via positive Orthogonal Random features (FAVOR+) for linear approximation of softmax attention.
result Performers achieve competitive results on various tasks, demonstrating the effectiveness of their attention-learning approach.
Enhances Fourier estimator performance for asynchronous event-data.
problem Improving correlation and covariance estimation on event-data.
method Implement and test NUFFT methods with different averaging kernels.
result Demonstrates improved performance and relationship between averaging scales.
The X-ray transform on the periodic slab [0,1]×Tn, n≥0, has a non-trivial kernel due to the symmetry of the manifold and presence of trapped geodesics. For tensor fields gauge freedom increases the kernel further, and the X-ray transform is not solenoidally injective unless n=0. We characterize t…
Deep vanilla transformers trained without shortcuts achieve similar performance to standard models.
problem Training deep vanilla transformers without shortcuts and normalizations.
method Parameter initializations, bias matrices, and location-dependent rescaling.
result Deep vanilla transformers can train at similar speeds and performance to standard models.
Kernel estimator optimally recovers function from noisy exponential Radon transform.
problem Inverting noisy exponential Radon transform of a function.
method Proposed a kernel estimator to estimate the true function.
result The estimator converges to the true function at minimax optimal rate.
Quantum kernel improves solar irradiance forecasting.
problem Improving short-term solar irradiance forecasting accuracy.
method Quantum Fourier Transform kernel in KRR with feature mixing.
result Consistently improves R2 and nRMSE over classical kernels.
New unsupervised learning technique learns independent kernels for better machine learning tasks.
problem Improving unsupervised representation learning for machine learning tasks.
method Stacking convolutional transforms using alternating proximal minimization scheme.
result DCTL outperforms shallow version CTL on benchmark datasets.
New method simplifies tomographic reconstruction using RKHS.
problem Tomographic reconstruction challenges.
method RKHS framework for X-ray transform.
result Sharp stability results without Fourier transform.
One considers the class of complete non-compact Riemannian manifolds whose heat kernel satisfies Gaussian estimates from above and below. One shows that the Riesz transform is Lp bounded on such a manifold, for p ranging in an open interval above 2, if and only if the gradient of the heat kernel satisfies a certai…
TNP-KR improves scalability of NPs with Transformer blocks and attention mechanisms.
problem Scalability bottleneck in NPs and GPs, especially for large datasets.
method Introduces TNP-KR with KRBlock, kernel-based attention, and two attention mechanisms.
result TNP-KR with DKA outperforms Performer and achieves state-of-the-art results.
Canonical correlation analysis (CCA) is a multivariate statistical technique for finding the linear relationship between two sets of variables. The kernel generalization of CCA named kernel CCA has been proposed to find nonlinear relations between datasets. Despite their wide usage, they have one common limitation that…
Skyformer uses Gaussian kernel and Nyström method to speed up self-attention in transformers.
problem High computational cost of self-attention in transformers.
method Replaces softmax with Gaussian kernel and applies Nyström method for matrix approximation.
result Skyformer achieves comparable or better performance with fewer computation resources.
The study reveals how attention paths in Transformers influence learning outcomes.
problem Understanding the theoretical basis of Transformers' performance.
method Developed a statistical mechanics theory for a simplified attention network.
result The predictor statistics are influenced by the combination of attention paths.
Transformers can approximate Kalman Filtering in linear systems with small error.
problem Approximating Kalman Filtering using Transformers for linear dynamical systems.
method Two-step reduction: 1) Softmax self-attention block approximates Nadaraya-Watson kernel smoothing, 2) This estimator approximates Kalman Filter.
result Constructs a Transformer that implements the Kalman Filter with small additive error, uniformly bounded in time.
Positive definite kernels are an important tool in machine learning that enable efficient solutions to otherwise difficult or intractable problems by implicitly linearizing the problem geometry. In this paper we develop a set-theoretic interpretation of the Earth Mover's Distance (EMD) and propose Earth Mover's Interse…
This paper develops tools for nonreversible MCMC with convergence guarantees.
problem Designing nonreversible MCMC kernels with convergence guarantees.
method Develops tools for nonreversible Markov kernels using conditional invertible transforms.
result Ensures nonreversible kernels have the desired invariance property and lead to convergent algorithms.
Complex analysis techniques link Gaussian RBF kernels to quantum mechanics.
problem Understanding the Gaussian RBF kernel in machine learning and SVMs.
method Using Fock space and Segal-Bargmann theories in complex analysis.
result Proves connections between Gaussian RBF kernels and quantum mechanics operators.
We consider the kernel completion problem with the presence of multiple views in the data. In this context the data samples can be fully missing in some views, creating missing columns and rows to the kernel matrices that are calculated individually for each view. We propose to solve the problem of completing the kerne…
Kernel density estimation (KDE) is a popular statistical technique for estimating the underlying density distribution with minimal assumptions. Although they can be shown to achieve asymptotic estimation optimality for any input distribution, cross-validating for an optimal parameter requires significant computation do…
DGPFM uses deep Gaussian processes to map functions accurately and quantify uncertainty.
problem Learning mappings between functional spaces, especially when data are noisy, sparse, or irregularly sampled.
method Constructs a sequence of GP-based linear and nonlinear transformations directly in function space, leveraging kernel integral transforms, GP conditional means, and nonlinear activations sampled from Gaussian processes.
result Empirical results show DGPFM outperforms existing methods in predictive accuracy and uncertainty calibration.
Estimates path-valued data using signature metrics and local kernels.
problem Nonparametric regression and classification for path-valued data.
method Combines signature transform and local kernel regression.
result Establishes convergence bounds and demonstrates competitive accuracy.
Heat kernel resurgent structure from Picard-Lefschetz theory
problem Short-time heat kernel asymptotics
method Picard-Lefschetz theory
result 1-Gevrey small-time expansion
Scalable kernel methods for large datasets using Fourier representations and NUFFT.
problem Cubic complexity in kernel methods limits their use on large-scale datasets.
method Fourier representation of kernels combined with NUFFT for O(n log n) complexity.
result Achieves minimax convergence rates and processes up to tens of billions of samples.
Study finds the number of modes in Gaussian kernel density estimators scales with sqrt(β log β).
problem Determining the number of clusters in Transformers.
method Used Kac-Rice formula and Edgeworth expansion to prove scaling.
result The expected number of modes scales as Θ(√(β log β)).
Transformers can efficiently approximate nonparametric regression with minimal parameters and sequences.
problem Efficiently approximating nonparametric regression functions with transformers.
method Kernel-weighted polynomial basis and gradient descent.
result Achieves minimax optimal rate of convergence with fewer parameters and sequences.
Efficiently computes sparse signature coefficients using kernels.
problem Lack of efficient methods for sparse signature coefficients.
method Signature kernels and PDE-based methods.
result Sparse groups of signature coefficients can be isolated effectively.
The paper proves boundedness of a Riesz transform on weighted manifolds.
problem Establishing \(L^p\)-boundedness of the covariant Riesz transform on differential forms.
method Heat-kernel criterion, volume doubling, heat kernel estimates, curvature control, gradient bounds.
result The covariant Riesz transform is \(L^p\)-bounded for \(p>2\) on weighted Riemannian manifolds.
HRFs adaptively linearize kernels for accurate approximations.
problem Linearizing softmax and Gaussian kernels for machine learning applications.
method Generalizes Bochner's Theorem for kernels, uses random features for compositional kernels.
result Strong theoretical guarantees and unbiased approximation with smaller relative errors.
We characterize the kernel of the mixed ray transform on simple 2-dimensional Riemannian manifolds, that is, on simple surfaces for tensors of any order.
Paper converts deep networks to flat, equivalent kernel machines.
problem Capacity control and uniform convergence in deep learning.
method Push-forward transformation from deep networks to indefinite kernel machines.
result Flat network weights are Lp-norm regularized (0<p<1).
Let M be a Riemannian globally symmetric space of compact type, M′ its set of maximal flat totally geodesic tori, and ad(M) its adjoint space. We show that the kernel of the maximal flat Radon transform τ:L2(M)→L2(M′) is precisely the orthogonal complement of the image of the pullback map…
New method transforms complex stochastic equations into simpler ones for efficient simulation.
problem Efficient simulation of complex path-dependent stochastic processes.
method Transforms Volterra-type SDEs into standard diffusion processes using convolution kernels.
result Proposes a numerical simulation scheme with a strong convergence rate of 1/2.
Study shows how transformers classify symbols without naming them, proving a margin-versus-collision criterion.
problem How transformers classify symbols without naming them.
method Logistic classification analysis of transformer-kernel regime, colored collision graph.
result Decomposes learned predictor into ideal template-level classifier and finite-sample perturbation.