The adjusted Rand index (ARI) is commonly used in cluster analysis to measure the degree of agreement between two data partitions. Since its introduction, exploring the situations of extreme agreement and disagreement under different circumstances has been a subject of interest, in order to achieve a better understandi…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Proposes new random models for fuzzy clustering similarity measures.
The misclassification error distance and the adjusted Rand index are two of the most commonly used criteria to evaluate the performance of clustering algorithms. This paper provides an in-depth comparison of the two criteria, aimed to better understand exactly what they measure, their properties and their differences. …
Meta-learning neural networks for better clustering representations.
In unsupervised machine learning, agreement between partitions is commonly assessed with so-called external validity indices. Researchers tend to use and report indices that quantify agreement between two partitions for all clusters simultaneously. Commonly used examples are the Rand index and the adjusted Rand index. …
Unified framework for comparing clusterings from information-theoretic and pair-counting perspectives.
The goal of lifetime clustering is to develop an inductive model that maps subjects into clusters according to their underlying (unobserved) lifetime distribution. We introduce a neural-network based lifetime clustering model that can find cluster assignments by directly maximizing the divergence between the empiri…
The main goal of this study is to extract a set of brain networks in multiple time-resolutions to analyze the connectivity patterns among the anatomic regions for a given cognitive task. We suggest a deep architecture which learns the natural groupings of the connectivity patterns of human brain in multiple time-resolu…
New metric improves clustering in persistent homology.
Paper addresses xVA models for market-implied skew and smile.
A new measure DCSI quantifies separability for density-based clustering.
Randomizes AD models for better option pricing.
It has been noticed that some external CVIs exhibit a preferential bias towards a larger or smaller number of clusters which is monotonic (directly or inversely) in the number of clusters in candidate partitions. This type of bias is caused by the functional form of the CVI model. For example, the popular Rand index (R…
Unsupervised image segmentation aims at clustering the set of pixels of an image into spatially homogeneous regions. We introduce here a class of Bayesian nonparametric models to address this problem. These models are based on a combination of a Potts-like spatial smoothness component and a prior on partitions which is…
In mixture model-based clustering applications, it is common to fit several models from a family and report clustering results from only the `best' one. In such circumstances, selection of this best model is achieved using a model selection criterion, most often the Bayesian information criterion. Rather than throw awa…
CDL index improves clustering validation for non-convex data.
In the travel industry, online customers book their travel itinerary according to several features, like cost and duration of the travel or the quality of amenities. To provide personalized recommendations for travel searches, an appropriate segmentation of customers is required. Clustering ensemble approaches were dev…
A system is presented that segments, clusters and predicts musical audio in an unsupervised manner, adjusting the number of (timbre) clusters instantaneously to the audio input. A sequence learning algorithm adapts its structure to a dynamically changing clustering tree. The flow of the system is as follows: 1) segment…
Study EM and GD for clustering with penalties for misspecification and high dimensions.
Adjusted for chance measures are widely used to compare partitions/clusterings of the same data set. In particular, the Adjusted Rand Index (ARI) based on pair-counting, and the Adjusted Mutual Information (AMI) based on Shannon information theory are very popular in the clustering community. Nonetheless it is an open …
Two types of zeroth-order stochastic algorithms have recently been designed for nonconvex optimization respectively based on the first-order techniques SVRG and SARAH/SPIDER. This paper addresses several important issues that are still open in these methods. First, all existing SVRG-type zeroth-order algorithms suffer …
STICC clusters geographic objects considering both spatial contiguity and attributes.
Mixtures of Unigrams are one of the simplest and most efficient tools for clustering textual data, as they assume that documents related to the same topic have similar distributions of terms, naturally described by Multinomials. When the classification task is particularly challenging, such as when the document-term ma…
Study compares clustering methods for mixed-type data.
A new measure normalizes clustering accuracy to evaluate algorithms better.
FCM clustering adapts to persistence diagrams for topological data analysis.
Clustering is a central approach for unsupervised learning. After clustering is applied, the most fundamental analysis is to quantitatively compare clusterings. Such comparisons are crucial for the evaluation of clustering methods as well as other tasks such as consensus clustering. It is often argued that, in order to…
New k-means method handles random data better than traditional techniques.
The study of genetic variants can help find correlating population groups to identify cohorts that are predisposed to common diseases and explain differences in disease susceptibility and how patients react to drugs. Machine learning algorithms are increasingly being applied to identify interacting GVs to understand th…
Visual summarization of clinical data collected on patients contained within the electronic health record (EHR) may enable precise and rapid triage at the time of patient presentation to an emergency department (ED). The triage process is critical in the appropriate allocation of resources and in anticipating eventual …
The paper investigates how irrelevant features affect clustering performance.
New method clusters matrix-valued data by latent variables.
The paper introduces a new model to correct bias in treatment effect estimates due to sample selection.
The paper critiques and expands on common evaluation metrics in machine learning.
We enhance short-rate models to control implied volatility analytically.
A new clustering evaluation index based on density estimation.
Study on symmetric operators on non-compact manifolds, focusing on their index modulo 2.
New methods reduce communication in distributed training for variational inequalities.
New index formula connects numerical and -theoretic indices.
The paper explores global index formulas for one-dimensional holomorphic foliations.
Explain Arnold's proof of the Morse index theorem using Maslov index.
Paper introduces danceability index as a new bridge index definition.
The p-index improves investment performance for NYSE stocks but not for SSE stocks.
Study Whittle index learning algorithms for restless bandits with constant stepsizes.
A new index rebalancing strategy reduces large constituent weights without undesirable effects.
We study bounded pseudoconvex domains in complex Euclidean space. We define an index associated to the boundary and show this new index is equivalent to the Diederich-Fornæss index defined in 1977. This connects the Diederich-Fornæss index to boundary conditions and refines the Levi pseudoconvexity. We also prove the $…
Study on symmetric braid index of ribbon knots, deriving bounds and characterizations.
Study proves bridge and braid indices match for twist positive knots.