Flexible classifier using Mahalanobis distances for non-elliptical distributions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In high dimension, low sample size (HDLSS)settings, the simple average distance classifier based on the Euclidean distance performs poorly if differences between the locations get masked by the scale differences. To rectify this issue, modifications to the average distance classifier was proposed by Chan and Hall (2009…
We consider classifiers for high-dimensional data under the strongly spiked eigenvalue (SSE) model. We first show that high-dimensional data often have the SSE model. We consider a distance-based classifier using eigenstructures for the SSE model. We apply the noise reduction methodology to estimation of the eigenvalue…
We consider general non-Euclidean distance measures between real world objects that need to be classified. It is assumed that objects are represented by distances to other objects only. Conditions for zero-error dissimilarity based classifiers are derived. Additional conditions are given under which the zero-error deci…
The paper classifies singularities of plane congruences and affine distance functions.
Paper defines a new distance metric for comparing learning tasks.
The -nearest neighbour (-NN) classifier is one of the oldest and most important supervised learning algorithms for classifying datasets. Traditionally the Euclidean norm is used as the distance for the -NN classifier. In this thesis we investigate the use of alternative distances for the -NN classifier. We …
Paper explains distance-based classifiers using neural network structures.
A new classifier uses Fermat distance for semi-supervised learning in high dimensions.
RS-Del provides robustness for sequence classifiers against edit distance attacks.
This thesis uses Kantorovich-Rubinstein distance for classifying points based on their measures.
Novel neural network approximates exact distance for robust classification.
We introduce a new metric to evaluate corruption robustness of ML classifiers.
Time series classification is an increasing research topic due to the vast amount of time series data that are being created over a wide variety of fields. The particularity of the data makes it a challenging task and different approaches have been taken, including the distance based approach. 1-NN has been a widely us…
We define generalized distance-squared mappings, and we concentrate on the plane to plane case. We classify generalized distance-squared mappings of the plane into the plane in a recognizable way.
PAC-Bayesian set up involves a stochastic classifier characterized by a posterior distribution on a classifier set, offers a high probability bound on its averaged true risk and is robust to the training sample used. For a given posterior, this bound captures the trade off between averaged empirical risk and KL-diverge…
The dynamic time warping (dtw) distance fails to satisfy the triangle inequality and the identity of indiscernibles. As a consequence, the dtw-distance is not warping-invariant, which in turn results in peculiarities in data mining applications. This article converts the dtw-distance to a semi-metric and shows that its…
Proposes a new scoring function for linear classifiers to improve object positioning in feature space.
This guide explains statistical distances for evaluating generative models.
If a simple 3-manifold M admits a reducible and a toroidal Dehn filling, the distance between the filling slopes is known to be bounded by three. In this paper, we classify all manifolds which admit a reducible Dehn filling and a toroidal Dehn filling with distance 3.
The nearest neighbor classifier fails in high dimensions, leading to this study.
Detecting adversarial examples is as hard as classifying them.
Study shows -NN classifier is not universally consistent on but consistent on discrete and specific measure spaces.
A statistical model predicts generalization in few-shot learning.
A scalable version of MADD improves big-data classification speed.
PolyGraph Discrepancy improves graph generative model evaluation.
Though deep learning has been applied successfully in many scenarios, malicious inputs with human-imperceptible perturbations can make it vulnerable in real applications. This paper proposes an error-correcting neural network (ECNN) that combines a set of binary classifiers to combat adversarial examples in the multi-c…
The Lorentzian length, which is one of the most significant functions in Lorentzian geometry, is a complex-valued function. Its square gives a real-valued non-degenerate quadratic function. In this paper, we define naturally extended mappings of Lorentzian distance-squared functions, wherein each component is a Lorentz…
The reliable measurement of confidence in classifiers' predictions is very important for many applications and is, therefore, an important part of classifier design. Yet, although deep learning has received tremendous attention in recent years, not much progress has been made in quantifying the prediction confidence of…
New classifiers for HDLSS data classify without tuning, robustly.
Distance metric learning is a successful way to enhance the performance of the nearest neighbor classifier. In most cases, however, the distribution of data does not obey a regular form and may change in different parts of the feature space. Regarding that, this paper proposes a novel local distance metric learning met…
New bounds for average graph distance using curvature and centrality.
Algorithm classifies point clouds using deep set linearized optimal transport.
The nearest-centroid classifier is a simple linear-time classifier based on computing the centroids of the data classes in the training phase, and then assigning a new datum to the class corresponding to its nearest centroid. Thanks to its very low computational cost, the nearest-centroid classifier is still widely use…
Study compares chi-squared divergence and KL-divergence posteriors for PAC-Bayesian bounds.
Here, a non-linear analysis method is applied rather than classical one to study projective Finsler geometry. More intuitively, by means of an inequality on Ricci-Finsler curvature, a projectively invariant pseudo-distance is introduced and an analogous of Schwarz' lemma in Finsler geometry is proved. Next, the Schwarz…
We propose and evaluate alternative ensemble schemes for a new instance based learning classifier, the Randomised Sphere Cover (RSC) classifier. RSC fuses instances into spheres, then bases classification on distance to spheres rather than distance to instances. The randomised nature of RSC makes it ideal for use in en…
In the research area of time series classification, the ensemble shapelet transform algorithm is one of state-of-the-art algorithms for classification. However, its high time complexity is an issue to hinder its application since its base classifier shapelet transform includes a high time complexity of a distance calcu…
We discuss theoretical aspects of the product rule for classification problems in supervised machine learning for the case of combining classifiers. We show that (1) the product rule arises from the MAP classifier supposing equivalent priors and conditional independence given a class; (2) under some conditions, the pro…
Social messages classification is a research domain that has attracted the attention of many researchers in these last years. Indeed, the social message is different from ordinary text because it has some special characteristics like its shortness. Then the development of new approaches for the processing of the social…
Defines contact surgery distance and shows it's bounded by topological surgery distance by 5.
Paper develops multivariate time series similarity and distance measures.
The literature postulates that the dynamic time warping (dtw) distance can cope with temporal variations but stores and processes time series in a form as if the dtw-distance cannot cope with such variations. To address this inconsistency, we first show that the dtw-distance is not warping-invariant. The lack of warpin…
Mathematical conditions and practical computations for adversarial robustness measures are established.
Uniform measures have played a fundamental role in geometric measure theory since they naturally appear as tangent objects. For instance, they were essential in the groundbreaking work of Preiss on the rectifiability of Radon measures. However, relatively little is understood about the structure of general uniform meas…
We classify generalized distance-squared mappings of into () having generic central points. Moreover, we show that there does not exist a universal bad set in the case of this dimension-pair.
Classifiers trained on data sets possessing an imbalanced class distribution are known to exhibit poor generalisation performance. This is known as the imbalanced learning problem. The problem becomes particularly acute when we consider incremental classifiers operating on imbalanced data streams, especially when the l…
Real-world data such as digital images, MRI scans and electroencephalography signals are naturally represented as matrices with structural information. Most existing classifiers aim to capture these structures by regularizing the regression matrix to be low-rank or sparse. Some other methodologies introduce factorizati…