Paper introduces new neural network models and theories.
problem Understanding neural networks beyond over-parameterized regime.
method Develops two exact models and a novel representor theory.
result Provides insights into neural network training and kernel evolution.
Improved learning theory for kernel distribution regression with two-stage sampling.
problem Distribution regression problem and two-stage sampling setting.
method Kernel methods, near-unbiased condition, new error bounds, convergence rates.
result Strictly improved convergence rates for three important classes of kernels.
Analyzes feature learning in neural networks using a self-consistent dynamical field theory.
problem Feature learning in infinite-width neural networks.
method Constructs deterministic dynamical order parameters as inner-product kernels for hidden unit activations and gradients.
result Reveals the hidden layer activation distribution, neural tangent kernel evolution, and output predictions.
New approach to learning kernels from data using AIT principles.
problem Learning kernels from data in machine learning.
method Sparse Kernel Flows method based on AIT principles.
result Sparse Kernel Flows aligns with MDL principle and offers a robust theoretical foundation.
Constructs index for elliptic operators using rapidly decaying kernels.
problem Index of elliptic operators in Fréchet algebra.
method Uses heat operators and heat kernel asymptotics.
result Index can be represented by an idempotent involving heat operators.
The paper studies heat kernel behavior on symmetric spaces.
problem Large-time behavior of heat operator traces on symmetric spaces.
method Uses representation theory and Carmona's proof of Vogan's lambda map.
result Provides an asymptotic formula for heat kernel behavior.
In this paper we study the variational problem associated to support vector regression in Banach function spaces. Using the Fenchel-Rockafellar duality theory, we give explicit formulation of the dual problem as well as of the related optimality conditions. Moreover, we provide a new computational framework for solving…
Quantum Kerr learning shows enhancements in convergence and generalization for kernel-based methods.
problem Improving convergence and generalization in kernel-based methods for quantum computing.
method Combining quantum mechanics with neural tangent kernel theory and first-order perturbation theory.
result Quantum enhancements in terms of convergence time and generalization error.
Kernel-based methods improve policy evaluation in MRP models.
problem Estimating value functions in infinite-horizon discounted MRP models.
method Kernel-based temporal difference methods using reproducing kernel Hilbert spaces.
result Optimal error bounds derived for the kernel-based LSTD estimate.
Paper develops a new kernel approximation framework.
problem High time and space complexity of kernel methods for large datasets.
method Perturbation-based kernel approximation framework using classical perturbation theory.
result Framework generalizes and improves upon existing methods.
Kernel sparsity ("dying ReLUs") and lack of diversity are commonly observed in CNN kernels, which decreases model capacity. Drawing inspiration from information theory and wireless communications, we demonstrate the intersection of coding theory and deep learning through the Grassmannian subspace packing problem in CNN…
Bayesian neural networks explore rare fluctuations for better feature learning.
problem Understanding rare but dominant fluctuations in Bayesian neural networks.
method Large-deviation theory and joint optimization over predictors and internal kernels.
result Posterior rate function optimization reveals data-dependent kernel selection.
Generative models use kernel smoothing for conditioning on small example sets.
problem Improving generative models' performance with limited conditioning examples.
method Showed that cross-attention conditioning is equivalent to kernel smoothing, specifically a Nadaraya--Watson kernel smoother.
result The approach predicts and confirms three failure regimes for kernel-based conditioning.
New method interpolates high-dimensional scattered data using kernel theory.
problem Scattered data in high-dimensional spaces defy traditional distributional assumptions.
method Kernel interpolation framework based on integral operator theory.
result Spectra of kernel matrices predict performance of interpolation methods.
The study reveals how attention paths in Transformers influence learning outcomes.
problem Understanding the theoretical basis of Transformers' performance.
method Developed a statistical mechanics theory for a simplified attention network.
result The predictor statistics are influenced by the combination of attention paths.
The paper studies convergence of kernel autocovariance operators for stationary processes.
problem Estimating autocovariance operators of stationary processes on Polish spaces.
method Investigates convergence of empirical estimates of autocovariance operators under various conditions.
result Provides consistency results for kernel PCA and spectral analysis methods.
Heat kernel resurgent structure from Picard-Lefschetz theory
problem Short-time heat kernel asymptotics
method Picard-Lefschetz theory
result 1-Gevrey small-time expansion
Develops method to construct Lie algebra weight system kernel using Vogel algebra.
problem Detecting correlators and distinguishing knots in 3D Chern-Simons theory.
method Uses Vogel's Λ algebra and Jacobi diagrams.
result Explicitly provides Jacobi diagrams in the kernel of sl_N weight system.
This work develops a learning theory for inferring interaction kernels in complex agent systems.
problem Modeling complex interactions in systems of particles or agents.
method Nonparametric regression and approximation theory.
result Strong consistency and optimal convergence rates for estimators of interaction kernels.
The study explains how neural networks align their kernels to target functions during training.
problem Understanding how neural networks align their kernels to target functions during training.
method Theoretical analysis of kernel evolution in toy models and deep networks.
result Kernel alignment naturally emerges during training to accelerate convergence and improve generalization.
The paper examines when NTK theory applies to real finite-width neural networks.
problem Understanding when NTK theory accurately predicts the behavior of finite-width neural networks.
method Empirical study of fully-connected ReLU and sigmoid DNNs with various hyperparameters and depths.
result NTK theory does not always apply to sufficiently deep networks with exploding gradients, and the kernel changes significantly during training.
Kernel methods linked to feature subspaces and maximal correlation kernels.
problem Understanding kernel methods and their relationship to feature extraction.
method Established a correspondence between feature subspaces and kernels, introduced maximal correlation kernels, and demonstrated their optimality.
result Kernel SVM on maximal correlation kernel achieves minimum prediction error.
We propose a novel class of Gaussian processes (GPs) whose spectra have compact support, meaning that their sample trajectories are almost-surely band limited. As a complement to the growing literature on spectral design of covariance kernels, the core of our proposal is to model power spectral densities through a rect…
Gaussian process framework learns interaction kernels in multi-species particle systems.
problem Learning interaction kernels in multi-species interacting particle systems from trajectory data.
method Nonparametric Bayesian approach with Gaussian processes.
result Established rigorous statistical guarantees for recoverability and optimality of interaction kernels.
The book explores stochastic areas and heat kernels on manifolds.
problem Understanding stochastic area functionals and heat kernels on manifolds.
method Study of Brownian motions and heat kernels on Lie groups and Riemannian manifolds.
result Rich interactions between stochastic calculus, geometry, and random matrices.
In this paper, we propose a random projection approach to estimate variance in kernel ridge regression. Our approach leads to a consistent estimator of the true variance, while being computationally more efficient. Our variance estimator is optimal for a large family of kernels, including cubic splines and Gaussian ker…
Study of regularized least squares in RKKS with indefinite kernels.
problem Asymptotic properties of regularized least squares with indefinite kernels in RKKS.
method Introducing a bounded hyper-sphere constraint, theoretical demonstration of globally optimal solution, modified error decomposition techniques, matrix perturbation theory.
result Derivation of learning rates in RKKS, same as RKHS under certain conditions.
Entropy analysis via kernel methods for probabilistic inference.
problem Entropy analysis of probability distributions.
method Kernel methods and reproducing kernel Hilbert spaces for entropy estimation.
result New upper-bounds on log partition functions for probabilistic inference.
Positive definite kernels and their associated Reproducing Kernel Hilbert Spaces provide a mathematically compelling and practically competitive framework for learning from data. In this paper we take the approximation theory point of view to explore various aspects of smooth kernels related to their inferential proper…
Random matrix theory predicts neural representations generalize well.
problem Understanding why neural representations generalize well in practice.
method Applied random matrix theory to kernel regression and neural networks.
result GCV estimator accurately predicts generalization risk in overparameterized settings.
Kernel Dynamic Mode Decomposition reconstructs dynamical systems using Laplacian kernel.
problem Reconstructing spatial-temporal dynamics of complex systems.
method Kernel Dynamic Mode Decomposition with Laplacian kernel.
result Laplacian kernel allows for the closability of Koopman operators in RKHS, enabling reconstruction.
Study on kernel methods in large-scale machine learning problems.
problem Large-scale machine learning with many interacting variables.
method Mean field limit analysis of kernels and their Hilbert spaces.
result Mean field convergence of empirical and infinite-sample solutions.
The boundary-value problem for Laplace-type operators acting on smooth sections of a vector bundle over a compact Riemannian manifold with generalized local boundary conditions including both normal and tangential derivatives is studied. The condition of strong ellipticity of this boundary-value problem is formulated. …
We give a version of the comparison principle from pluripotential theory where the Monge-Ampère measure is replaced by the Bergman kernel and use it to derive a maximum principle
Neural networks outperform kernels by learning features better.
problem Current theories of feature learning do not adequately assess feature quality.
method Introduced feature quality metric and examined existing theories empirically.
result Current theories of feature learning do not provide a sufficient foundation for neural network generalization.
Geometric theory connects machine learning classifiers to differential geometry.
problem Classifying data points in machine learning.
method Mapping binary classification to vector bundles and differential geometry.
result Harmonic interpolation solves RKHS interpolation problems.
Unified theory for kernel regression generalizes well under realistic assumptions.
problem Analyzing kernel regression under realistic conditions.
method Unified theory providing rigorous bounds for various settings.
result Self-regularization phenomenon in kernel matrices enables good generalization.
We construct a canonical correspondence from a wide class of reproducing kernels on infinite-dimensional Hermitian vector bundles to linear connections on these bundles. The linear connection in question is obtained through a pull-back operation involving the tautological universal bundle and the classifying morphism o…
Kernel methods are studied in a mean field limit for high-dimensional data.
problem Analyzing kernel methods in high-dimensional data with many variables.
method Investigation of kernel methods in the mean field limit of interacting particle systems.
result Rigorous mean field limit of kernels and detailed analysis of the limiting reproducing kernel Hilbert space.
This paper reviews the functional aspects of statistical learning theory. The main point under consideration is the nature of the hypothesis set when no prior information is available but data. Within this framework we first discuss about the hypothesis set: it is a vectorial space, it is a set of pointwise defined fun…
The study uses response theory to understand RNNs processing input signals.
problem Understanding how RNNs process sequential data.
method Deriving a Volterra series representation for SRNNs output using response theory from nonequilibrium statistical mechanics.
result SRNNs can be viewed as kernel machines operating on a reproducing kernel Hilbert space associated with the response feature.
We analyze in this paper a random feature map based on a theory of invariance I-theory introduced recently. More specifically, a group invariant signal signature is obtained through cumulative distributions of group transformed random projections. Our analysis bridges invariant feature learning with kernel methods, as …
Uniform bounds for neural networks' generalization error in overparameterized settings.
problem Generalization error in overparameterized neural networks.
method Neural Tangent kernel theory and Mercer decomposition of the NT kernel in spherical harmonics.
result Uniform generalization bounds for overparameterized neural networks in RKHS.
Kernels for structured data are commonly obtained by decomposing objects into their parts and adding up the similarities between all pairs of parts measured by a base kernel. Assignment kernels are based on an optimal bijection between the parts and have proven to be an effective alternative to the established convolut…
In signal analysis and synthesis, linear approximation theory considers a linear decomposition of any given signal in a set of atoms, collected into a so-called dictionary. Relevant sparse representations are obtained by relaxing the orthogonality condition of the atoms, yielding overcomplete dictionaries with an exten…
The paper analyzes and improves the learning rates of distributed kernel ridge regression.
problem Generalization performance and learning rates of distributed kernel ridge regression.
method The paper derives optimal learning rates for DKRR in expectation and probability, proposes a communication strategy to improve learning performance, and evaluates these through theory and experiments.
result The communication strategy significantly improves the learning performance of DKRR, as demonstrated by both theoretical assessments and numerical experiments.
Kernel PCA helps analyze multivariate extremes and clusters them effectively.
problem Analyzing the dependence structure of multivariate extremes.
method Kernel PCA as a method for clustering and dimension reduction.
result Kernel PCA preimages effectively identify clusters in multivariate extremes.
This work closes the theory-practice gap for distributed optimization methods by introducing a new regularity condition.
problem Existing convergence conditions for distributed optimization methods are violated by nearly all kernels used in practice.
method Introduces Hessian relative uniform continuity (HRUC) to guarantee convergence under mild conditions.
result Derives convergence guarantees for mirror descent-based gradient tracking without restrictive assumptions.