The study shows inner-product kernels behave similarly to binary kernels in high dimensions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New test for binary treatment effects using kernel methods.
The term "CoRE kernel" stands for correlation-resemblance kernel. In many applications (e.g., vision), the data are often high-dimensional, sparse, and non-binary. We propose two types of (nonlinear) CoRE kernels for non-binary sparse data and demonstrate the effectiveness of the new kernels through a classification ex…
With the advent of kernel methods, automating the task of specifying a suitable kernel has become increasingly important. In this context, the Multiple Kernel Learning (MKL) problem of finding a combination of pre-specified base kernels that is suitable for the task at hand has received significant attention from resea…
Transformers are explained as infinite-dimensional kernel machines.
New tests for binary classification regression functions without distribution assumptions.
Kauri is a novel unsupervised binary tree for clustering that outperforms existing methods.
Paper introduces Laplace-HDC for better binary hyperdimensional computing.
We present generalization bounds for the TS-MKL framework for two stage multiple kernel learning. We also present bounds for sparse kernel learning formulations within the TS-MKL framework.
Support Vector Machines (SVMs) are powerful learners that have led to state-of-the-art results in various computer vision problems. SVMs suffer from various drawbacks in terms of selecting the right kernel, which depends on the image descriptors, as well as computational and memory efficiency. This paper introduces a n…
The min-max kernel is a generalization of the popular resemblance kernel (which is designed for binary data). In this paper, we demonstrate, through an extensive classification study using kernel machines, that the min-max kernel often provides an effective measure of similarity for nonnegative data. As the min-max ker…
Improved text classification performance through conformal transformations of kernels.
Geometric theory connects machine learning classifiers to differential geometry.
Kernel-based learning predicts ICU escalation from COVID-19 chest X-rays.
Random Forest kernels improve performance in various regression and survival tasks.
This paper uses MIO to select features for kernel SVM classification.
Proposes a gradient-based variable selection method for binary classification in RKHS.
Paper proposes a simple estimator for DPP correlation kernels.
Tree ensembles like RF and GBT can be seen as kernels, improving regression and classification performance.
New method uses path signatures for efficient likelihood estimation in time-series data.
A scalable ROC-SVM variant reduces training time for imbalanced binary classification.
Extends Tanimoto kernel to real-valued functions.
The paper develops efficient estimators for semi-parametric binary models in distributed computing.
This paper aims at refined error analysis for binary classification using support vector machine (SVM) with Gaussian kernel and convex loss. Our first result shows that for some loss functions such as the truncated quadratic loss and quadratic loss, SVM with Gaussian kernel can reach the almost optimal learning rate, p…
We consider a Gaussian process formulation of the multiple kernel learning problem. The goal is to select the convex combination of kernel matrices that best explains the data and by doing so improve the generalisation on unseen data. Sparsity in the kernel weights is obtained by adopting a hierarchical Bayesian approa…
The paper studies binary classification and aims at estimating the underlying regression function which is the conditional expectation of the class labels given the inputs. The regression function is the key component of the Bayes optimal classifier, moreover, besides providing optimal predictions, it can also assess t…
We establish upper bounds for the minimal number of hidden units for which a binary stochastic feedforward network with sigmoid activation probabilities and a single hidden layer is a universal approximator of Markov kernels. We show that each possible probabilistic assignment of the states of output units, given t…
Sequential hypothesis testing is a desirable decision making strategy in any time sensitive scenario. Compared with fixed sample-size testing, sequential testing is capable of achieving identical probability of error requirements using less samples in average. For a binary detection problem, it is well known that for k…
Efficient methods for sparse random projections improve classification accuracy in very high-dimensional data.
We pose causal inference as the problem of learning to classify probability distributions. In particular, we assume access to a collection , where each is a sample drawn from the probability distribution of , and is a binary label indicating whether "" or …
This paper introduces a probability density estimator based on Green's function identities. A density model is constructed under the sole assumption that the probability density is differentiable. The method is implemented as a binary likelihood estimator for classification purposes, so issues such as mis-modeling and …
For binary classification we establish learning rates up to the order of for support vector machines (SVMs) with hinge loss and Gaussian RBF kernels. These rates are in terms of two assumptions on the considered distributions: Tsybakov's noise assumption to establish a small estimation error, and a new geometr…
In this chapter we review the main literature related to kernel spectral clustering (KSC), an approach to clustering cast within a kernel-based optimization setting. KSC represents a least-squares support vector machine based formulation of spectral clustering described by a weighted kernel PCA objective. Just as in th…
BKP R package models spatially varying binomial probabilities efficiently.
Method reveals dissimilarity in alloys' Curie temperatures.
Paper shows similarity learning can lead to strong binary classification performance.
Deep learning predicts drug prescriptions across global health records.
Generalizes NTK for surrogate gradient learning in neural networks.
New method finds 198,846 toric-colorable seeds of Picard number 5.
Model detects electricity theft with high accuracy.
In many applications (in particular information systems, such as pattern recognition, machine learning, cheminformatics, bioinformatics to name but a few) the assessment of uncertainty is essential - i.e., the estimation of the underlying probability distribution function. More often than not, the form of this function…
Anomaly detection based on one-class classification algorithms is broadly used in many applied domains like image processing (e.g. detection of whether a patient is "cancerous" or "healthy" from mammography image), network intrusion detection, etc. Performance of an anomaly detection algorithm crucially depends on a ke…
In this paper we measured the stability of stochastic gradient method (SGM) for learning an approximated Fourier primal support vector machine. The stability of an algorithm is considered by measuring the generalization error in terms of the absolute difference between the test and the training error. Our problem is to…
Model detects patterns in noisy binary data, explaining neuron activity in terms of cell assemblies.
Study preference-based reinforcement learning in episodic kernel MDPs.
We consider binary classification problems with positive definite kernels and square loss, and study the convergence rates of stochastic gradient methods. We show that while the excess testing loss (squared loss) converges slowly to zero as the number of observations (and thus iterations) goes to infinity, the testing …
Multiple kernel learning (MKL) algorithms combine different base kernels to obtain a more efficient representation in the feature space. Focusing on discriminative tasks, MKL has been used successfully for feature selection and finding the significant modalities of the data. In such applications, each base kernel repre…
We extend kernelized matrix factorization with a fully Bayesian treatment and with an ability to work with multiple side information sources expressed as different kernels. Kernel functions have been introduced to matrix factorization to integrate side information about the rows and columns (e.g., objects and users in …