The computational and storage complexity of kernel machines presents the primary barrier to their scaling to large, modern, datasets. A common way to tackle the scalability issue is to use the conjugate gradient algorithm, which relieves the constraints on both storage (the kernel matrix need not be stored) and computa…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study heavy-tailed weights' impact on neural network's spectral distribution.
Study eigenvalue distributions of neural kernels for linear-width networks.
The paper analyzes contraction rates for GP regression approximations.
Study of eigenvalues in nonlinear kernels for classification of separable data.
The paper studies neural networks with wide layers and finds a deformed semicircle law.
In this article, we obtain a strict inequality between the conjugate Hardy kernels and the Bergman kernels on planar regular regions with boundary components, which is a conjecture of Saitoh.
Conjugate gradient methods improve efficiency for high-dimensional GLMMs.
Study shows deterministic equivalent for neural network kernel convergence.
New method speeds up Gaussian process training and inference for large datasets.
Regularized least-squares (kernel-ridge / Gaussian process) regression is a fundamental algorithm of statistics and machine learning. Because generic algorithms for the exact solution have cubic complexity in the number of datapoints, large datasets require to resort to approximations. In this work, the computation of …
We characterize the conjugate linearized Ricci flow and the associated backward heat kernel on closed three--manifolds of bounded geometry. We discuss their properties, and introduce the notion of Ricci flow conjugated constraint sets which characterizes a way of Ricci flow averaging metric dependent geometrical data. …
We extend kernelized matrix factorization with a fully Bayesian treatment and with an ability to work with multiple side information sources expressed as different kernels. Kernel functions have been introduced to matrix factorization to integrate side information about the rows and columns (e.g., objects and users in …
GLSKF improves tensor completion by capturing both global and local variations.
New framework estimates eigenvalues of kernel matrices without full matrix construction.
Resolution of a compact group action in the sense described by Albin and Melrose is applied to the conjugation action by the unitary group on self-adjoint matrices. It is shown that the eigenvalues are smooth on the resolved space and that the trivial bundle smoothly decomposes into the direct sum of global one-dimensi…
In this paper we investigate the small time heat kernel asymptotics on the cut locus on a class of surfaces of revolution, which are the simplest 2-dimensional Riemannian manifolds different from the sphere with non trivial cut-conjugate locus. We determine the degeneracy of the exponential map near a cut-conjugate poi…
Investigates O(n)-invariant metrics on SPD matrices, extending kernel metrics.
We prove ultradifferentiable Chevelley restriction theorems for a wide range of ultradifferentiable classes. As a special case we find that isotropic functions, i.e., functions defined on the vector space of real symmetric matrices invariant under the action of the special orthogonal group by conjugation, possess some …
Sharp Li-Yau equality proven for shrinking Ricci solitons without curvature assumptions.
We study the geodesic X-ray transform on compact Riemannian surfaces with conjugate points. Regardless of the type of the conjugate points, we show that we cannot recover the singularities and therefore, this transform is always unstable (ill-posed). We describe the microlocal kernel of and relate it to the con…
In this article we derive Harnack estimates for conjugate heat kernel in an abstract geometric flow. Our calculation involves a correction term D. When D is nonnegative, we are able to obtain a Harnack inequality. Our abstract formulation provides a unified framework for some known results, in particular including corr…
Two methods solve kernel ridge regression problems efficiently.
We provide bounds for kernel matrices and new approximations for high-dimensional data.
Despite their successes, what makes kernel methods difficult to use in many large scale problems is the fact that storing and computing the decision function is typically expensive, especially at prediction time. In this paper, we overcome this difficulty by proposing Fastfood, an approximation that accelerates such co…
Study eigenvalues and eigenvectors in neural networks, focusing on signal propagation.
With the huge influx of various data nowadays, extracting knowledge from them has become an interesting but tedious task among data scientists, particularly when the data come in heterogeneous form and have missing information. Many data completion techniques had been introduced, especially in the advent of kernel meth…
A new method for efficiently computing derivatives of skew-symmetric matrix exponentials.
In this article we obtain a result about the uniqueness of factorization in terms of conjugates of the matrix $U=(\xymatrix{1 & 1 0 & 1})$, of some matrices representing the conjugacy classes of those elements of arising as the monodromy around a singular fiber in an elliptic fibration (i.e. those matrices th…
New method tackles high-dimensional SBL without covariance matrices.
We prove statistical rates of convergence for kernel-based least squares regression from i.i.d. data using a conjugate gradient algorithm, where regularization against overfitting is obtained by early stopping. This method is related to Kernel Partial Least Squares, a regression method that combines supervised dimensio…
We describe how cross-kernel matrices, that is, kernel matrices between the data and a custom chosen set of `feature spanning points' can be used for learning. The main potential of cross-kernels lies in the fact that (a) only one side of the matrix scales with the number of data points, and (b) cross-kernels, as oppos…
The study examines Fisher information matrices and neural tangent kernels for simple ReLU networks with random weights.
DEQs and explicit networks are nearly equivalent for Gaussian mixtures.
The matrix completion problem consists of finding or approximating a low-rank matrix based on a few samples of this matrix. We propose a new algorithm for matrix completion that minimizes the least-square distance on the sampling set over the Riemannian manifold of fixed-rank matrices. The algorithm is an adaptation of…
Kernel methods are successful approaches for different machine learning problems. This success is mainly rooted in using feature maps and kernel matrices. Some methods rely on the eigenvalues/eigenvectors of the kernel matrix, while for other methods the spectral information can be used to estimate the excess risk. An …
New method speeds up Bayesian optimization in high dimensions.
A new method for deep Wishart processes improves kernel-based models.
We discuss a natural form of Ricci--flow conjugation between two distinct general relativistic data sets given on a compact -dimensional manifold . We establish the existence of the relevant entropy functionals for the matter and geometrical variables, their monotonicity properties, and the associated conve…
Study reveals an equivalence principle for the spectrum of random inner-product kernel matrices in polynomial scaling.
In this paper, we study the spectrum and the eigenvectors of radial kernels for mixtures of distributions in . Our approach focuses on high dimensions and relies solely on the concentration properties of the components in the mixture. We give several results describing of the structure of kernel matrices …
Deep kernel processes unify various models using Gram matrices and kernel functions.
A non-singular sesquilinear form is constructed that is preserved by the Lawrence-Krammer representation. It is shown that if the polynomial variables q and t of the Lawrence-Krammer representation are chosen to be appropriate algebraically independant unit complex numbers, then the form is negative-definite Hermitian.…
We prove rates of convergence in the statistical sense for kernel-based least squares regression using a conjugate gradient algorithm, where regularization against overfitting is obtained by early stopping. This method is directly related to Kernel Partial Least Squares, a regression method that combines supervised dim…
We propose a scheme for recycling Gaussian random vectors into structured matrices to approximate various kernel functions in sublinear time via random embeddings. Our framework includes the Fastfood construction as a special case, but also extends to Circulant, Toeplitz and Hankel matrices, and the broader family of s…
Kernel methods are an extremely popular set of techniques used for many important machine learning and data analysis applications. In addition to having good practical performances, these methods are supported by a well-developed theory. Kernel methods use an implicit mapping of the input data into a high dimensional f…
Kernel matrices (e.g. Gram or similarity matrices) are essential for many state-of-the-art approaches to classification, clustering, and dimensionality reduction. For large datasets, the cost of forming and factoring such kernel matrices becomes intractable. To address this challenge, we introduce a new adaptive sampli…
We propose and study kernel conjugate gradient methods (KCGM) with random projections for least-squares regression over a separable Hilbert space. Considering two types of random projections generated by randomized sketches and Nyström subsampling, we prove optimal statistical results with respect to variants of norms …