Paper identifies key function spaces for ReLU networks based on Fisher information.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Develops a framework for learning nonlinear operators using Mercer kernels.
Uniform bounds for neural networks' generalization error in overparameterized settings.
Bayesian neural networks with Mercer priors for interpretable uncertainty quantification.
Survey of kernels, RKHS, and their applications in machine learning.
New method approximates MMD using pseudo-differential operators and singular values.
New asymmetric kernel methods improve feature learning.
Study bounds on kernel function entropy for finite measures.
The study assesses low-rank approximations in Gaussian Process regression.
The study assesses low-rank approximations in Gaussian Process regression.
Transformers are explained as infinite-dimensional kernel machines.
Lecture notes on kernel functions and Random Fourier Features.
Kernel interpolation improved with continuous volume sampling.
A new PCA method for analyzing point processes.
Overlapping clustering problem is an important learning issue in which clusters are not mutually exclusive and each object may belongs simultaneously to several clusters. This paper presents a kernel based method that produces overlapping clusters on a high feature space using mercer kernel techniques to improve separa…
This paper presents a unified framework to tackle estimation problems in Digital Signal Processing (DSP) using Support Vector Machines (SVMs). The use of SVMs in estimation problems has been traditionally limited to its mere use as a black-box model. Noting such limitations in the literature, we take advantage of sever…
Devoted to multi-task learning and structured output learning, operator-valued kernels provide a flexible tool to build vector-valued functions in the context of Reproducing Kernel Hilbert Spaces. To scale up these methods, we extend the celebrated Random Fourier Feature methodology to get an approximation of operator-…
We consider the problem of cost sensitive multiclass classification, where we would like to increase the sensitivity of an important class at the expense of a less important one. We adopt an {\em apportioned margin} framework to address this problem, which enables an efficient margin shift between classes that share th…
New learning rates derived for Tikhonov-regularized problems without kernel assumptions.
We investigate a generic problem of learning pairwise exponential family graphical models with pairwise sufficient statistics defined by a global mapping function, e.g., Mercer kernels. This subclass of pairwise graphical models allow us to flexibly capture complex interactions among variables beyond pairwise product. …
Interest in multioutput kernel methods is increasing, whether under the guise of multitask learning, multisensor networks or structured output data. From the Gaussian process perspective a multioutput Mercer kernel is a covariance function over correlated output functions. One way of constructing such kernels is based …
As a robust nonlinear similarity measure in kernel space, correntropy has received increasing attention in domains of machine learning and signal processing. In particular, the maximum correntropy criterion (MCC) has recently been successfully applied in robust regression and filtering. The default kernel function in c…
In machine learning or statistics, it is often desirable to reduce the dimensionality of a sample of data points in a high dimensional space . This paper introduces a dimensionality reduction method where the embedding coordinates are the eigenvectors of a positive semi-definite kernel obtained as the sol…
Recently, there has been emerging interest in constructing reproducing kernel Banach spaces (RKBS) for applied and theoretical purposes such as machine learning, sampling reconstruction, sparse approximation and functional analysis. Existing constructions include the reflexive RKBS via a bilinear form, the semi-inner-p…
This work analyzes how different layers in deep neural networks contribute to generalization error.
A simple framework Probabilistic Multi-view Graph Embedding (PMvGE) is proposed for multi-view feature learning with many-to-many associations so that it generalizes various existing multi-view methods. PMvGE is a probabilistic model for predicting new associations via graph embedding of the nodes of data vectors with …
Sequential modelling with self-attention has achieved cutting edge performances in natural language processing. With advantages in model flexibility, computation complexity and interpretability, self-attention is gradually becoming a key component in event sequence models. However, like most other sequence models, self…
In data science, determining proximity between observations is critical to many downstream analyses such as clustering, information retrieval and classification. However, when the underlying structure of the data probability space is unclear, the function used to compute similarity between data points is often arbitrar…
We reformulate unsupervised dimension reduction problem (UDR) in the language of tempered distributions, i.e. as a problem of approximating an empirical probability density function by another tempered distribution, supported in a -dimensional subspace. We show that this task is connected with another classical prob…
A new deep neural network tackles nonlinear functional regression with improved dimensionality reduction.
Researchers describe and compare decompositions of Poincaré duality pairs.
The paper proposes and discusses semiorthogonal decompositions for moduli spaces of vector bundles.
The paper classifies decompositions of 3-sphere and lens spaces with handlebodies.
We combine aspects of the notions of finite decomposition complexity and asymptotic property C into a notion that we call finite APC-decomposition complexity. Any space with finite decomposition complexity has finite APC-decomposition complexity and any space with asymptotic property C has finite APC-decomposition comp…
Paper develops a new algorithm for distribution regression with optimal learning rates.
This paper generalizes octahedral decomposition to links in thickened surfaces.
Paper learns optimal kernels for Gaussian process regression in aerodynamics.
Researchers compute Goeritz groups for all (1,1)-link decompositions.
Study concordance of decompositions from defining sequences in 3-sphere.
Given a Delaunay decomposition of a compact hyperbolic surface, one may record the topological data of the decomposition, together with the intersection angles between the `empty disks' circumscribing the regions of the decomposition. The main result of this paper is a characterization of when a given topological decom…
Study shows OAT decomposition generates unexplained profit and loss, while SU decompositions depend on risk factor order.
A new algorithm speeds up CP decomposition for large tensors.
Paper characterizes optimization landscape of Tucker decomposition.
A double pants decomposition of a 2-dimensional surface is a collection of two pants decomposition of this surface introduced in arXiv:1005.0073v2. There are two natural operations acting on double pants decompositions: flips and handle twists. It is shown in arXiv:1005.0073v2 that the groupoid generated by flips and h…
Smooth 4-manifolds have simple horizontal decompositions.
Let be the real form of a complex simple Jordan algebra such that the automorphism group is . By using some orbit types of on , for , explicitly, we give the Iwasawa decomposition, the Oshima--Sekiguchi's Iwasawa decomp…
We study the topological types of pants decompositions of a surface by associating to any pants decomposition in a natural way its pants decomposition graph, This perspective provides a convenient way to analyze the maximum distance in the pants complex of any pants decomposition to a pants decomposition c…
New method uses random decompositions for high-dimensional Bayesian optimization.