New criterion selects optimal number of clusters based on stability.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Cluster stability selection improves feature selection in correlated data.
reval package selects best clustering solutions via stability-based validation.
Study examines how cluster number affects short-text clustering, introducing a stability metric.
Randomized hierarchical clustering tests for stability and detects clusters.
Discussing issues in robust clustering, especially with Gaussian models.
A popular method for selecting the number of clusters is based on stability arguments: one chooses the number of clusters such that the corresponding clustering results are "most stable". In recent years, a series of papers has analyzed the behavior of this method from a theoretical point of view. However, the results …
CARVE validates clustering results using resampling and stability analysis.
We consider two interesting spaces associated to a quiver with potential: a space of stability conditions and a cluster variety. In the case where the quiver with potential arises from an ideal triangulation of a marked bordered surface, we construct a natural map from a dense subset of the space of stability condition…
This paper introduces cluster exchange groupoids for Coxeter-Dynkin diagrams and finds their fundamental groups are braid groups.
We assess cluster stability by trimming extreme points and tracking data range reduction.
The paper provides presentations for mapping class groups and cluster automorphism groups of surfaces.
Characterizes pseudo-Anosov mapping classes on general marked surfaces.
Stable density-based clustering via multiparameter persistence.
We investigate the role of the initialization for the stability of the k-means clustering algorithm. As opposed to other papers, we consider the actual k-means algorithm and do not ignore its property of getting stuck in local optima. We are interested in the actual clustering, not only in the costs of the solution. We…
Study clusters Kenyan medical insurance companies based on financial performance and reporting consistency.
Determining the number of clusters present in a dataset is an important problem in cluster analysis. Conventional clustering techniques generally assume this parameter to be provided up front. %user supplied. %Recently, robustness of any given clustering algorithm is analyzed to measure cluster stability/instability wh…
Characterizes pseudo-Anosov mapping classes using cluster algebra techniques.
We present in this paper an empirical framework motivated by the practitioner point of view on stability. The goal is to both assess clustering validity and yield market insights by providing through the data perturbations we propose a multi-view of the assets' clustering behaviour. The perturbation framework is illust…
Paper translates train track concepts to cluster algebras for pseudo-Anosov mapping classes.
S3VDC improves DC methods for scalability, stability, and simplicity.
In this paper, we study different discrete data clustering methods, which use the Model-Based Clustering (MBC) framework with the Multinomial distribution. Our study comprises several relevant issues, such as initialization, model estimation and model selection. Additionally, we propose a novel MBC method by efficientl…
Crowded trades by similarly trading peers influence the dynamics of asset prices, possibly creating systemic risk. We propose a market clustering measure using granular trading data. For each stock the clustering measure captures the degree of trading overlap among any two investors in that stock. We investigate the ef…
Empty core found in max-loss non-centroid clustering.
MAS scores cluster size consistency from points, robust to label changes.
New models learn stable latent clusters without side info.
The paper provides guarantees for clustering validity without distributional assumptions.
Clust-PSI-PFL uses PSI to improve accuracy and fairness in federated learning.
We study the spaces of polynomials stratified into the sets of polynomial with fixed number of roots inside certain semialgebraic region , on its border, and at the complement to its closure. Presented approach is a generalisation, unification and development of several classical approaches to stability problems in …
We present a graph-theoretical approach to data clustering, which combines the creation of a graph from the data with Markov Stability, a multiscale community detection framework. We show how the multiscale capabilities of the method allow the estimation of the number of clusters, as well as alleviating the sensitivity…
Clustering ensemble, or consensus clustering, has emerged as a powerful tool for improving both the robustness and the stability of results from individual clustering methods. Weighted clustering ensemble arises naturally from clustering ensemble. One of the arguments for weighted clustering ensemble is that elements (…
Future grid scenario analysis requires a major departure from conventional power system planning, where only a handful of most critical conditions is typically analyzed. To capture the inter-seasonal variations in renewable generation of a future grid scenario necessitates the use of computationally intensive time-seri…
Recurrent neural networks trained on regular languages exhibit stable states that can recover from noise.
This paper introduces hierarchical quasi-clustering methods, a generalization of hierarchical clustering for asymmetric networks where the output structure preserves the asymmetry of the input data. We show that this output structure is equivalent to a finite quasi-ultrametric space and study admissibility with respect…
We introduce the cluster exchange groupoid associated to a non-degenerate quiver with potential, as an enhancement of the cluster exchange graph. In the case that arises from an (unpunctured) marked surface, where the exchange graph is modelled on the graph of triangulations of the marked surface, we show that the univ…
t-SNE and hierarchical clustering are popular methods of exploratory data analysis, particularly in biology. Building on recent advances in speeding up t-SNE and obtaining finer-grained structure, we combine the two to create tree-SNE, a hierarchical clustering and visualization algorithm based on stacked one-dimension…
Bi-filtration stabilizes TDA mapper results under noise.
Here, we propose a clustering technique for general clustering problems including those that have non-convex clusters. For a given desired number of clusters , we use three stages to find a clustering. The first stage uses a hybrid clustering technique to produce a series of clusterings of various sizes (randomly se…
To devise efficient solutions for approximating a mean partition in consensus clustering, Dimitriadou et al. [3] presented a necessary condition of optimality for a consensus function based on least square distances. We show that their result is pivotal for deriving interesting properties of consensus clustering beyond…
Convex clustering, a convex relaxation of k-means clustering and hierarchical clustering, has drawn recent attentions since it nicely addresses the instability issue of traditional nonconvex clustering methods. Although its computational and statistical properties have been recently studied, the performance of convex c…
An adaptive clustering algorithm learns from evolving data without manual tuning.
High density clusters can be characterized by the connected components of a level set of the underlying probability density function generating the data, at some appropriate level . The complete hierarchical clustering can be characterized by a cluster tree ${\cal T}= \bigcup_λ L(λ)…
Efficient clustering for large datasets using a sampling-based approach.
The clustering algorithms that view each object data as a single sample drawn from a certain distribution, Gaussian distribution, for example, has been a hot topic for decades. Many clustering algorithms: such as k-means and spectral clustering are proposed based on the single sample assumption. However, in real life, …
We introduce a property of mutation loops, called the sign stability, with a focus on an asymptotic behavior of the iteration of the tropical -transformation. A sign-stable mutation loop has a numerical invariant which we call the cluster stretch factor, in analogy with that of a pseudo-Anosov mapping clas…
We construct a framework for studying clustering algorithms, which includes two key ideas: persistence and functoriality. The first encodes the idea that the output of a clustering scheme should carry a multiresolution structure, the second the idea that one should be able to compare the results of clustering algorithm…
Cluster analysis is a fundamental tool for pattern discovery of complex heterogeneous data. Prevalent clustering methods mainly focus on vector or matrix-variate data and are not applicable to general-order tensors, which arise frequently in modern scientific and business applications. Moreover, there is a gap between …
Framework for multi-scale clustering using phase transitions.