Paper identifies key function spaces for ReLU networks based on Fisher information.
problem Understanding the structure of Fisher information matrices in ReLU networks.
method Spectral decomposition of Fisher information matrices, focusing on the first three eigenspaces.
result The first three eigenspaces account for 97.7% of the trace of the Fisher information matrix, corresponding to spherical harmonic functions of order ≤2.
Develops a framework for learning nonlinear operators using Mercer kernels.
problem Learning nonlinear operators between infinite-dimensional spaces.
method Stochastic approximation framework with Mercer operator-valued kernels.
result Establishes dimension-free polynomial convergence rates for nonlinear operator learning.
Uniform bounds for neural networks' generalization error in overparameterized settings.
problem Generalization error in overparameterized neural networks.
method Neural Tangent kernel theory and Mercer decomposition of the NT kernel in spherical harmonics.
result Uniform generalization bounds for overparameterized neural networks in RKHS.
Bayesian neural networks with Mercer priors for interpretable uncertainty quantification.
problem Uncertainty quantification in neural networks, especially for complex input-to-output mappings.
method Introducing Mercer priors for BNNs, which approximate a specified GP and are scalable.
result BNNs with Mercer priors can approximate the uncertainty of a specified GP, making them interpretable and scalable.
Survey of kernels, RKHS, and their applications in machine learning.
problem Understanding kernels and their applications in machine learning.
method Review of historical context, mathematical definitions, and practical applications of kernels.
result Comprehensive overview of kernels, RKHS, and their applications.
New method approximates MMD using pseudo-differential operators and singular values.
problem Approximating MMD with pseudo-differential operators and singular values.
method Corresponding pseudo-differential operators to Mercer kernels, approximating p(x,y) with its first r singular values. result The new MMD distance measures the difference of two distributions with respect to r∗ local moments, where r∗ depends on singular values decay rate. New asymmetric kernel methods improve feature learning.
problem Improving feature learning with asymmetric kernels.
method Coupled covariance eigenproblem and Nyström method.
result Empirical evaluations show benefits of KSVD.
Study bounds on kernel function entropy for finite measures.
problem Investigate bounds on the ε-entropy of kernel classes.
method Sharp upper and lower bounds for p in [1, +∞] derived from eigenvalue behavior and Mercer series convergence.
result Proves tighter bounds for general kernels compared to previous work.
The study assesses low-rank approximations in Gaussian Process regression.
problem Improving Gaussian Process regression efficiency with low-rank approximations.
method Analyzes two low-rank approximations: random Fourier features and Mercer expansion truncation.
result Bounds on the divergence and error between exact and approximate GP models.
The study assesses low-rank approximations in Gaussian Process regression.
problem Improving the efficiency of Gaussian Process regression while maintaining accuracy.
method Analyzes two low-rank approximations: random Fourier features and Mercer expansion truncation, and bounds the divergence and error between exact and approximate models.
result Theoretical bounds on the divergence and error between exact and approximate Gaussian Process models are provided.
Transformers are explained as infinite-dimensional kernel machines.
problem Understanding the mechanics of Transformers in AI.
method Characterized Transformers' attention mechanism as a kernel learning method on Banach spaces.
result Transformer's kernel has infinite feature dimension and can learn any binary non-Mercer reproducing kernel Banach space pair.
Lecture notes on kernel functions and Random Fourier Features.
problem Understanding and approximating kernel functions in machine learning.
method Mathematical background and proofs of concentration results.
result Estimation of error in Random Fourier Features approximation.
Kernel interpolation improved with continuous volume sampling.
problem Approximating functions from RKHS using weighted sums of kernel translates.
method Continuous volume sampling for choosing node locations.
result Proved almost optimal bounds for interpolation and quadrature under VS.
A new PCA method for analyzing point processes.
problem Analyzing variability in replicated point processes.
method Functional Principal Component Analysis (fPCA) on cumulative mass functions.
result Established convergence and introduced principal measures.
New method embeds time span into self-attention for better temporal pattern recognition.
problem Capturing temporal patterns in event sequences without recurrent networks.
method Functional time representation learning with Bochner's and Mercer's Theorems.
result Proposed methods outperform baseline models in various continuous-time event sequence prediction tasks.
Overlapping clustering problem is an important learning issue in which clusters are not mutually exclusive and each object may belongs simultaneously to several clusters. This paper presents a kernel based method that produces overlapping clusters on a high feature space using mercer kernel techniques to improve separa…
This paper presents a unified framework to tackle estimation problems in Digital Signal Processing (DSP) using Support Vector Machines (SVMs). The use of SVMs in estimation problems has been traditionally limited to its mere use as a black-box model. Noting such limitations in the literature, we take advantage of sever…
Devoted to multi-task learning and structured output learning, operator-valued kernels provide a flexible tool to build vector-valued functions in the context of Reproducing Kernel Hilbert Spaces. To scale up these methods, we extend the celebrated Random Fourier Feature methodology to get an approximation of operator-…
New learning rates derived for Tikhonov-regularized problems without kernel assumptions.
problem Learning rates for Tikhonov-regularized learning problems.
method Minimax adaptive rates derived using Fourier isocapacitary condition and interpolation theory.
result Derivation of minimax adaptive rates without requiring kernel assumptions.
We investigate a generic problem of learning pairwise exponential family graphical models with pairwise sufficient statistics defined by a global mapping function, e.g., Mercer kernels. This subclass of pairwise graphical models allow us to flexibly capture complex interactions among variables beyond pairwise product. …
Interest in multioutput kernel methods is increasing, whether under the guise of multitask learning, multisensor networks or structured output data. From the Gaussian process perspective a multioutput Mercer kernel is a covariance function over correlated output functions. One way of constructing such kernels is based …
As a robust nonlinear similarity measure in kernel space, correntropy has received increasing attention in domains of machine learning and signal processing. In particular, the maximum correntropy criterion (MCC) has recently been successfully applied in robust regression and filtering. The default kernel function in c…
A new approach to cost-sensitive multiclass classification prioritizes certain classes over others.
problem Cost-sensitive multiclass classification where some classes are more important than others.
method Apportioned margin framework that shifts the decision boundary to prioritize certain classes.
result The method improves the error rate for important classes while reducing overall error.
In machine learning or statistics, it is often desirable to reduce the dimensionality of a sample of data points in a high dimensional space Rd. This paper introduces a dimensionality reduction method where the embedding coordinates are the eigenvectors of a positive semi-definite kernel obtained as the sol…
Recently, there has been emerging interest in constructing reproducing kernel Banach spaces (RKBS) for applied and theoretical purposes such as machine learning, sampling reconstruction, sparse approximation and functional analysis. Existing constructions include the reflexive RKBS via a bilinear form, the semi-inner-p…
This work analyzes how different layers in deep neural networks contribute to generalization error.
problem Understanding the role of each layer in deep neural networks for generalization.
method Spectral analysis, Neural Tangent Kernel, Hermite polynomials, Spherical Harmonics.
result Initial layers in deep neural networks have a larger bias towards high-frequency functions.
A simple framework Probabilistic Multi-view Graph Embedding (PMvGE) is proposed for multi-view feature learning with many-to-many associations so that it generalizes various existing multi-view methods. PMvGE is a probabilistic model for predicting new associations via graph embedding of the nodes of data vectors with …
In data science, determining proximity between observations is critical to many downstream analyses such as clustering, information retrieval and classification. However, when the underlying structure of the data probability space is unclear, the function used to compute similarity between data points is often arbitrar…
We reformulate unsupervised dimension reduction problem (UDR) in the language of tempered distributions, i.e. as a problem of approximating an empirical probability density function by another tempered distribution, supported in a k-dimensional subspace. We show that this task is connected with another classical prob…
A new deep neural network tackles nonlinear functional regression with improved dimensionality reduction.
problem Nonlinear functional regression in infinite-dimensional functional data analysis.
method Functional deep neural network with adaptive kernel embedding and projection steps.
result Explicit rates of approximating nonlinear smooth functionals are derived, and the network is shown to be effective in both simulated and real datasets.
Researchers describe and compare decompositions of Poincaré duality pairs.
problem Understanding and comparing different decompositions of Poincaré duality pairs.
method Developed and described edge splittings of decompositions based on group properties.
result Compared decompositions with two other related decompositions.
The paper proposes and discusses semiorthogonal decompositions for moduli spaces of vector bundles.
problem Decompositions of moduli spaces of vector bundles with fixed determinant of odd degree.
method Semiorthogonal decompositions, Grothendieck ring of varieties, mirror symmetry, graph potentials, Fukaya category.
result Evidence for a conjectural semiorthogonal decomposition of moduli spaces of rank 2 bundles with odd determinant.
The paper classifies decompositions of 3-sphere and lens spaces with handlebodies.
problem Classifying decompositions of 3-manifolds with handlebodies.
method Studied decompositions of 3-sphere and lens spaces with three handlebodies, using stabilizations.
result Determined whether decompositions are stabilized.
Paper develops a new algorithm for distribution regression with optimal learning rates.
problem Distribution regression with limited second-stage samples.
method Multi-penalty regularization in a reproducing kernel Hilbert space.
result Derives optimal learning rates for distribution regression.
We combine aspects of the notions of finite decomposition complexity and asymptotic property C into a notion that we call finite APC-decomposition complexity. Any space with finite decomposition complexity has finite APC-decomposition complexity and any space with asymptotic property C has finite APC-decomposition comp…
This paper generalizes octahedral decomposition to links in thickened surfaces.
problem Understanding the geometry of links in thickened surfaces.
method Octahedral decomposition of links in thickened surfaces.
result Nonpositive curvature of the complement and essential-ness of edges proved.
Paper learns optimal kernels for Gaussian process regression in aerodynamics.
problem Approximating complex functions from limited data in aerodynamics.
method Two algorithms: Kernel Flow and Spectral Kernel Ridge Regression.
result Explicit construction of optimal kernels based on target function features.
Researchers compute Goeritz groups for all (1,1)-link decompositions.
problem Computing Goeritz groups for all (1,1)-link decompositions.
method Analyzing surface decompositions and isotopy classes of homeomorphisms.
result Computed Goeritz groups for all (1,1)-link decompositions.
Study concordance of decompositions from defining sequences in 3-sphere.
problem Understanding concordance and bordism of decompositions from defining sequences.
method Relate to invariants of toroidal decompositions and cobordism of homology manifolds.
result At least uncountably many concordance classes of decompositions in 3-sphere.
Given a Delaunay decomposition of a compact hyperbolic surface, one may record the topological data of the decomposition, together with the intersection angles between the `empty disks' circumscribing the regions of the decomposition. The main result of this paper is a characterization of when a given topological decom…
Study shows OAT decomposition generates unexplained profit and loss, while SU decompositions depend on risk factor order.
problem Understanding profit and loss attribution in financial markets.
method Used financial market data from 2003 to 2022 to compare OAT, SU, and ASU decompositions.
result SU decompositions are sensitive to risk factor order and cannot identify all relevant risk factors.
A new algorithm speeds up CP decomposition for large tensors.
problem Efficiently processing large-scale tensors in real-time.
method Randomized online CP decomposition (ROCP) algorithm.
result ROCP reduces computing time and memory usage significantly.
Paper characterizes optimization landscape of Tucker decomposition.
problem Finding exact Tucker decomposition is a nonconvex optimization problem.
method Characterized the optimization landscape and provided a local search algorithm.
result All local minima are globally optimal if tensor has an exact Tucker decomposition.
A double pants decomposition of a 2-dimensional surface is a collection of two pants decomposition of this surface introduced in arXiv:1005.0073v2. There are two natural operations acting on double pants decompositions: flips and handle twists. It is shown in arXiv:1005.0073v2 that the groupoid generated by flips and h…
Smooth 4-manifolds have simple horizontal decompositions.
problem Classifying smooth, closed, orientable 4-manifolds.
method Horizontal handlebody decomposition.
result Simplest horizontal decompositions classify closed 4-manifolds.
Let J1 be the real form of a complex simple Jordan algebra such that the automorphism group is F4(−20). By using some orbit types of F4(−20) on J1, for F4(−20), explicitly, we give the Iwasawa decomposition, the Oshima--Sekiguchi's Kε−Iwasawa decomp…
We study the topological types of pants decompositions of a surface by associating to any pants decomposition P, in a natural way its pants decomposition graph, Γ(P). This perspective provides a convenient way to analyze the maximum distance in the pants complex of any pants decomposition to a pants decomposition c…
New method uses random decompositions for high-dimensional Bayesian optimization.
problem Learning accurate decompositions for high-dimensional black-box functions.
method Data-independent random tree-based decomposition sampling.
result Random decomposition upper-confidence bound algorithm (RDUCB) yields significant empirical gains.