The article examines different thresholding methods for improving PAM algorithm in cancer classification.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new classification method using disjoint centroids and normalized distance.
The nearest-centroid classifier is a simple linear-time classifier based on computing the centroids of the data classes in the training phase, and then assigning a new datum to the class corresponding to its nearest centroid. Thanks to its very low computational cost, the nearest-centroid classifier is still widely use…
A conceptually simple way to classify images is to directly compare test-set data and training-set data. The accuracy of this approach is limited by the method of comparison used, and by the extent to which the training-set data cover configuration space. Here we show that this coverage can be substantially increased u…
Multilayer bootstrap network builds a gradually narrowed multilayer nonlinear network from bottom up for unsupervised nonlinear dimensionality reduction. Each layer of the network is a nonparametric density estimator. It consists of a group of k-centroids clusterings. Each clustering randomly selects data points with r…
A fast algorithm for -means clustering using subsampled SDP.
CDF uses centroids to split features for high-dimensional classification.
The study characterizes quadrics among affine hyperspheres based on centroid collinearity of sections.
Ball k-means reduces point-centroid distance computations for faster k-means clustering.
We define a new method to estimate centroid for text classification based on the symmetric KL-divergence between the distribution of words in training documents and their class centroids. Experiments on several standard data sets indicate that the new method achieves substantial improvements over the traditional classi…
Centroid Transformers reduce memory and computation by summarizing inputs into centroids.
EKM addresses imbalanced data clustering by repelling centroids in large clusters.
Centroid-Encoder reduces high-dimensional data for better visualization.
Empty core found in max-loss non-centroid clustering.
Due to the success of the bag-of-word modeling paradigm, clustering histograms has become an important ingredient of modern information processing. Clustering histograms can be performed using the celebrated -means centroid-based algorithm. From the viewpoint of applications, it is usually required to deal with symm…
The paper proves a theorem linking convex body centroids and category theory.
Centroids Matching tackles catastrophic forgetting by matching feature vectors to class centroids.
Optimizes a small set of centroid points to approximate bootstrap distribution.
This paper proposes the use of an optimization algorithm, namely PSO to decide the initial centroids in K-means, to eventually get better accuracy. The vectorized notation of the optimal centroids can be thought of as entities in an optimization space, where the accuracy of K-means over a random subset of the data coul…
Text clustering method replaces centroids with summaries for interpretability and scalability.
We formally prove the connection between k-means clustering and the predictions of neural networks based on the softmax activation layer. In existing work, this connection has been analyzed empirically, but it has never before been mathematically derived. The softmax function partitions the transformed input space into…
The paper generalizes the second Pappus-Guldin theorem for calculating volumes of bodies.
Sharp Lp affine isoperimetric inequalities are established for the entire class of Lp projection bodies and the entire class of Lp centroid bodies. These new inequalities strengthen the Lp Petty projection and the Lp Busemann--Petty centroid inequality.
In addition to finding meaningful clusters, centroid-based clustering algorithms such as K-means or mean-shift should ideally find centroids that are valid patterns in the input space, representative of data in their cluster. This is challenging with data having a nonconvex or manifold structure, as with images or text…
K-means -- and the celebrated Lloyd algorithm -- is more than the clustering method it was originally designed to be. It has indeed proven pivotal to help increase the speed of many machine learning and data analysis techniques such as indexing, nearest-neighbor search and prediction, data compression; its beneficial u…
CCC clusters with controlled spread, outperforming standard methods.
We study a notion of a Lipschitz, permutation-invariant "centroid" for triples of points in mapping class groups MCG(S), which satisfies a certain polynomial growth bound. A consequence (via work of Drutu-Sapir or Chatterji-Ruane) is the Rapid Decay Property for MCG(S).
Method counters noisy labels by discounting distant samples.
New meta-learning method improves domain generalization by balancing parameters closer to domain centroids.
New clustering method reduces data redundancy for better summaries.
SIVF k-means algorithm speeds up sparse data clustering.
We introduce a new volume definition on normed vector spaces. We show that the induced -area functionals are convex for all . In the particular case , our theorem implies that Busemann's 2-volume density is convex, which was recently shown by Burago-Ivanov. We also show how the new volume definition is relat…
HIV RNA viral load (VL) is an important outcome variable in studies of HIV infected persons. There exists only a handful of methods which classify patients by viral load patterns. Most methods place limits on the use of viral load measurements, are often specific to a particular study design, and do not account for com…
This study evaluates cluster search algorithms using Gaussian mixture models.
Paper presents robust clustering methods for general mixture models.
The paper explores centroids and static equilibrium points in non-Euclidean geometries.
We investigate the use of Deep Neural Networks for the classification of image datasets where texture features are important for generating class-conditional discriminative representations. To this end, we first derive the size of the feature space for some standard textural features extracted from the input dataset an…
A new method clusters complex networks using topological and geometric structure.
Proposes a robust clustering method using the Median-of-Means estimator.
Pedal curves derived from ellipses are invariant in area.
Archimedes showed that the area between a parabola and any chord on the parabola is four thirds of the area of triangle , where P is the point on the parabola at which the tangent is parallel to the chord . Recently, this property of parabolas was proved to be a characteristic property of parabolas. With…
Optimal inequality on sphere for convex bodies.
HD-BWDM improves clustering validation in high-dimensional data.
Individual's semantics have been used for guiding the learning process of Genetic Programming solving supervised learning problems. The semantics has been used to proposed novel genetic operators as well as different ways of performing parent selection. The latter is the focus of this contribution by proposing three he…
Method summarizes and predicts time series data for COVID-19 cases and deaths.
This paper introduces Laplace techniques for designing a neural network, with the goal of estimating simplex-constraint sparse vectors from compressed measurements. To this end, we recast the problem of MMSE estimation (w.r.t. a pre-defined uniform input distribution) as the problem of computing the centroid of some po…
New method improves fairness of facial recognition systems.
Deep neural networks have been shown to suffer from a surprising weakness: their classification outputs can be changed by small, non-random perturbations of their inputs. This adversarial example phenomenon has been explained as originating from deep networks being "too linear" (Goodfellow et al., 2014). We show here t…