Improved average distance classifier for HDLSS settings with multiple population differences.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New bounds for average graph distance using curvature and centrality.
PAC-Bayesian set up involves a stochastic classifier characterized by a posterior distribution on a classifier set, offers a high probability bound on its averaged true risk and is robust to the training sample used. For a given posterior, this bound captures the trade off between averaged empirical risk and KL-diverge…
In this work, we addressed the issue of combining linear classifiers using their score functions. The value of the scoring function depends on the distance from the decision boundary. Two score functions have been tested and four different combination strategies were investigated. During the experimental study, the pro…
A simple method flags images as out-of-distribution based on their distance to nearest neighbors.
DW-KNN improves KNN by integrating distance and neighbor reliability for better prediction accuracy.
The paper introduces tests for high-dimensional independence using maximum and average distance correlations.
tl;dr: no, it cannot, at least not on average on the standard archive problems. We assess whether using six smoothing algorithms (moving average, exponential smoothing, Gaussian filter, Savitzky-Golay filter, Fourier approximation and a recursive median sieve) could be automatically applied to time series classificatio…
Flexible classifier using Mahalanobis distances for non-elliptical distributions.
Paper proposes a method to classify EEG signals with missing data.
We consider classifiers for high-dimensional data under the strongly spiked eigenvalue (SSE) model. We first show that high-dimensional data often have the SSE model. We consider a distance-based classifier using eigenstructures for the SSE model. We apply the noise reduction methodology to estimation of the eigenvalue…
We consider general non-Euclidean distance measures between real world objects that need to be classified. It is assumed that objects are represented by distances to other objects only. Conditions for zero-error dissimilarity based classifiers are derived. Additional conditions are given under which the zero-error deci…
We propose the Wasserstein-Fourier (WF) distance to measure the (dis)similarity between time series by quantifying the displacement of their energy across frequencies. The WF distance operates by calculating the Wasserstein distance between the (normalised) power spectral densities (NPSD) of time series. Yet this ratio…
The paper classifies singularities of plane congruences and affine distance functions.
Study on predicting graph labels at nodes using local averaging and distance estimation.
Paper defines a new distance metric for comparing learning tasks.
The -nearest neighbour (-NN) classifier is one of the oldest and most important supervised learning algorithms for classifying datasets. Traditionally the Euclidean norm is used as the distance for the -NN classifier. In this thesis we investigate the use of alternative distances for the -NN classifier. We …
Paper explains distance-based classifiers using neural network structures.
A new classifier uses Fermat distance for semi-supervised learning in high dimensions.
Paper proves CLTs for Q-learning with asynchronous updates.
RS-Del provides robustness for sequence classifiers against edit distance attacks.
We present an information-theoretic framework for bounding the number of labeled samples needed to train a classifier in a parametric Bayesian setting. We derive bounds on the average distance between the learned classifier and the true maximum a posteriori classifier, which are well-established surrogates for th…
This thesis uses Kantorovich-Rubinstein distance for classifying points based on their measures.
Novel neural network approximates exact distance for robust classification.
In his seminal work, Schapire (1990) proved that weak classifiers could be improved to achieve arbitrarily high accuracy, but he never implied that a simple majority-vote mechanism could always do the trick. By comparing the asymptotic misclassification error of the majority-vote classifier with the average individual …
Clusters of financial market states identified over 2006-2019.
We introduce a new metric to evaluate corruption robustness of ML classifiers.
Single particle reconstruction (SPR) from cryo-electron microscopy (EM) is a technique in which the 3D structure of a molecule needs to be determined from its contrast transfer function (CTF) affected, noisy 2D projection images taken at unknown viewing directions. One of the main challenges in cryo-EM is the typically…
Adversarial robustness has become an important research topic given empirical demonstrations on the lack of robustness of deep neural networks. Unfortunately, recent theoretical results suggest that adversarial training induces a strict tradeoff between classification accuracy and adversarial robustness. In this paper,…
Time series classification is an increasing research topic due to the vast amount of time series data that are being created over a wide variety of fields. The particularity of the data makes it a challenging task and different approaches have been taken, including the distance based approach. 1-NN has been a widely us…
In this paper we study the geometry of metric spheres in the curve complex of a surface, with the goal of determining the "average" distance between points on a given sphere. Averaging is not technically possible because metric spheres in the curve complex are countably infinite and do not support any invariant probabi…
We define generalized distance-squared mappings, and we concentrate on the plane to plane case. We classify generalized distance-squared mappings of the plane into the plane in a recognizable way.
This paper provides new insight into maximizing F1 scores in the context of binary classification and also in the context of multilabel classification. The harmonic mean of precision and recall, F1 score is widely used to measure the success of a binary classifier when one class is rare. Micro average, macro average, a…
Proposes a non-convex optimization method for a parsimonious weighted naive Bayes classifier.
The dynamic time warping (dtw) distance fails to satisfy the triangle inequality and the identity of indiscernibles. As a consequence, the dtw-distance is not warping-invariant, which in turn results in peculiarities in data mining applications. This article converts the dtw-distance to a semi-metric and shows that its…
Proposes a new scoring function for linear classifiers to improve object positioning in feature space.
New fairness criteria for algorithmic recourse actions that consider causal relationships.
Revises Bayesian model averaging for foundation models.
Mahalanobis distance detects anomalies well, but not for classification.
Proposes MFSWB for marginal fairness in SWB, improving efficiency and performance.
A clustering procedure, based on the Hausdorff distance, is introduced and tested on the financial time series of the Dow Jones Industrial Average (DJIA) index.
This guide explains statistical distances for evaluating generative models.
If a simple 3-manifold M admits a reducible and a toroidal Dehn filling, the distance between the filling slopes is known to be bounded by three. In this paper, we classify all manifolds which admit a reducible Dehn filling and a toroidal Dehn filling with distance 3.
CADM proposes a cluster-specific distance metric for categorical data clustering.
The nearest neighbor classifier fails in high dimensions, leading to this study.
Model tracks structural changes in Brownian particle configurations on a sphere.
We prove that S^2 x S^2 satisfies an intermediate condition between having metrics with positive Ricci and positive sectional curvature. Namely, there exist metrics for which the average of the sectional curvatures of any two planes tangent at the same point, but separated by a minimum distance in the 2-Grassmannian, i…
Detecting adversarial examples is as hard as classifying them.