New criterion selects optimal number of clusters based on stability.
problem Challenges in selecting optimal number of clusters in non-parametric clustering.
method Proposes a stability-based validation criterion combining between-cluster and within-cluster stability.
result Empirically demonstrates effectiveness in selecting optimal number of clusters.
Cluster stability selection improves feature selection in correlated data.
problem Feature selection stability in correlated data.
method Cluster stability selection exploiting known cluster structure.
result Better predictive performance than lasso alone and stability selection.
reval package selects best clustering solutions via stability-based validation.
problem Challenges in determining best clustering solutions due to lack of validation methods.
method Stability-based relative clustering validation methods.
result Determines best clustering solutions that generalize to unseen data.
Proposes a method to predict cluster number and cluster representatives using cluster stability analysis.
problem Determining the number of clusters in a dataset.
method Analyzes cluster stability using Monte-Carlo simulation to predict cluster number and find cluster representatives.
result Significant improvement in predicting cluster numbers and cluster composition in large datasets.
Study examines how cluster number affects short-text clustering, introducing a stability metric.
problem Challenges in finding meaningful clusters in short-text data.
method Introduces a stability metric to determine cluster robustness and visualizes cluster subdivisions.
result Choosing a cluster number involves balancing informativeness and complexity, not seeking a single 'optimal' solution.
Proposes a method to assess clustering stability in financial time series.
problem Assessing the validity of clustering in financial data.
method Empirical framework with data perturbations.
result Provides insights into assets' clustering behavior through multi-view analysis.
Randomized hierarchical clustering tests for stability and detects clusters.
problem Greedy hierarchical clustering's sensitivity to data perturbations.
method Randomization scheme and p-values at each node.
result Valid hypothesis testing procedures for clustering results.
Discussing issues in robust clustering, especially with Gaussian models.
problem Handling outliers and ambiguity in clustering groups.
method Focus on Gaussian mixture model, examining formal definitions, interactions, and tuning decisions.
result Outliers can confuse clustering groups and existing stability measures fail with them.
Solves Riemann-Hilbert problems on surface triangulations.
problem Riemann-Hilbert problems in Donaldson-Thomas theory.
method Map from stability conditions to cluster variety.
result Constructs solutions to Riemann-Hilbert problems.
Study compares discrete data clustering methods using Multinomial distribution.
problem Discrete data clustering with Multinomial distribution.
method Model-Based Clustering (MBC) framework with Multinomial distribution, combining partitional and hierarchical clustering techniques.
result Proposed method is competitive in clustering accuracy and better in stability and computation time.
A popular method for selecting the number of clusters is based on stability arguments: one chooses the number of clusters such that the corresponding clustering results are "most stable". In recent years, a series of papers has analyzed the behavior of this method from a theoretical point of view. However, the results …
Crowded trades cluster investors, affecting stock price stability.
problem Crowded trades lead to price instability and systemic risk.
method Market clustering measure using granular trading data.
result Market clustering has a causal effect on stock return distribution tails, especially positive tail.
CARVE validates clustering results using resampling and stability analysis.
problem Inconsistent and unreliable clustering results due to algorithm, preprocessing, and k sensitivity. method CARVE uses resampling-based validation and stability analysis to evaluate multiple clustering algorithms and hyperparameters.
result CARVE consistently recovers near-optimal clusterings and finer biological structure.
Study polynomials' root clustering via symmetric products, unifying stability approaches.
problem Stability of polynomials with roots in specific regions.
method Interpreting roots and coefficients as symmetric product morphism, analyzing topology up to homeomorphism.
result Description of strata topology and adjacency, explaining classical stability problems.
Graph clustering uses multiscale community detection for improved performance.
problem Improving data clustering accuracy and robustness.
method Graph-theoretical approach combining multiscale community detection.
result Multiscale graph-based clustering achieves better performance than traditional methods.
This paper introduces cluster exchange groupoids for Coxeter-Dynkin diagrams and finds their fundamental groups are braid groups.
problem Understanding the fundamental groups of cluster exchange groupoids for Coxeter-Dynkin diagrams.
method Introduced cluster exchange groupoids for Coxeter-Dynkin diagrams and showed the fundamental group isomorphic to braid groups.
result The fundamental group of the exchange groupoid for a Coxeter-Dynkin diagram is the braid group associated with the diagram.
We assess cluster stability by trimming extreme points and tracking data range reduction.
problem Assessing stability of one-dimensional clusters.
method Probabilistic method using diameter-shrinkage ratio to track data range reduction.
result Our method achieves higher accuracy than classical tests in small or noisy samples.
The paper provides presentations for mapping class groups and cluster automorphism groups of surfaces.
problem Presentations of mapping class groups of surfaces stabilizing boundaries.
method Gave presentations of mapping class groups of marked surfaces stabilizing boundaries.
result Presented cluster automorphism groups of cluster algebras from surfaces.
Characterizes pseudo-Anosov mapping classes on general marked surfaces.
problem Stability of mapping classes on marked surfaces.
method Cluster algebraic description and reduction procedure of mapping classes.
result Characterizes pseudo-Anosov mapping classes in terms of uniform sign stability.
Stable density-based clustering via multiparameter persistence.
problem Density-based clustering stability to data perturbations.
method Degree-Rips construction, correspondence-interleaving distance, multiparameter stability analysis.
result Persistable pipeline yields stable, consistent density-based clustering.
We investigate the role of the initialization for the stability of the k-means clustering algorithm. As opposed to other papers, we consider the actual k-means algorithm and do not ignore its property of getting stuck in local optima. We are interested in the actual clustering, not only in the costs of the solution. We…
Study clusters Kenyan medical insurance companies based on financial performance and reporting consistency.
problem Identifying financial health and reporting consistency in Kenyan medical insurance companies.
method Advanced clustering techniques (KMeans, DTW) on financial ratios and time series data.
result Four distinct clusters identified, each representing different financial performance and reporting consistency combinations.
Enhances consensus clustering with a stronger Mean Partition Theorem.
problem Improving consensus clustering solutions for mean partition.
method Presented a stronger Mean Partition Theorem and Expected Partition Theorem.
result Shows versatility of the Mean Partition Theorem in multiple applications.
Enhances understanding of stability conditions on surfaces.
problem Understanding stability conditions on surfaces.
method Introduces cluster exchange groupoid and uses triangulation covering graphs.
result Space of stability conditions is simply connected.
Characterizes pseudo-Anosov mapping classes using cluster algebra techniques.
problem Characterize pseudo-Anosov mapping classes purely in terms of shear coordinates.
method Uses cluster algebraic generalization and tropical cluster transformations.
result Algebraic entropies of cluster transformations match topological entropy.
S3VDC improves DC methods for scalability, stability, and simplicity.
problem Poor scalability, instability, and lack of simplicity in DC methods.
method Four algorithmic improvements: initial γ-training, periodic β-annealing, mini-batch GMM initialization, and inverse min-max transform. S3VDC incorporates all improvements. result S3VDC outperforms state-of-the-art methods on benchmark and industrial datasets.
Paper translates train track concepts to cluster algebras for pseudo-Anosov mapping classes.
problem Understanding pseudo-Anosov mapping classes on surfaces.
method Using Goncharov--Shen's potential function, the paper translates train track concepts into cluster algebra language.
result Proves sign stability of general pseudo-Anosov mapping classes.
Sparse Convex Clustering improves clustering performance in high-dimensional data.
problem Distortion in convex clustering performance with uninformative features.
method Introduces Sparse Convex Clustering with an adaptive group-lasso penalty and a tuning criterion based on clustering stability.
result Demonstrates improved clustering performance through feature selection.
A new t-k-means algorithm improves clustering stability and robustness.
problem Poor performance and instability of standard k-means on heavy-tailed data. method Proposes t-k-means, a robust and stable variant of k-means. result Demonstrates improved stability and robustness on datasets with outliers.
Empty core found in max-loss non-centroid clustering.
problem Core stability in non-centroid clustering under max-loss objective.
method Proof for all k≥3 and n≥9 agents, computer-aided proof for 2D Euclidean points.
result Core can be empty in non-centroid clustering under max-loss objective.
MAS scores cluster size consistency from points, robust to label changes.
problem Desired uniformity in cluster sizes, stability under label perturbations.
method Mass Agreement Score (MAS) measures point-centric cluster size consistency, robust to label changes.
result MAS yields similar scores for partitions with similar bulk structure, sensitive to genuine redistribution of cluster mass.
New models learn stable latent clusters without side info.
problem Stability of non-linear ICA representations without side information.
method Deep generative models with latent clusterings, compared to standard VAEs and auxiliary labeled models.
result Deep generative models with latent clusterings are as stable as models with side information.
The paper provides guarantees for clustering validity without distributional assumptions.
problem Validating clustering results without distributional assumptions.
method Generic method to obtain post-inference guarantees of near-optimality and stability for clustering.
result The guarantees do not depend on distributional assumptions but depend on the data set admitting a stable clustering.
DAOC provides stable clustering for large networks.
problem Stable clustering of large networks with accuracy and robustness.
method DAOC uses Overlap Decomposition for deterministic fine-grained clusters and Mutual Maximal Gain for robustness.
result DAOC yields stable clusters that are 25% more accurate than state-of-the-art deterministic algorithms.
Quickshift++ improves clustering stability and performance.
problem Improving initial seedings for clustering algorithms.
method Provably good initial seedings for Quick Shift clustering.
result Statistical consistency and strong clustering performance.
Clust-PSI-PFL uses PSI to improve accuracy and fairness in federated learning.
problem Non-IID data biases federated learning performance.
method Clust-PSI-PFL uses clustering and PSI to form homogeneous groups of clients.
result Clust-PSI-PFL delivers up to 18% higher global accuracy and improves client fairness.
Paper proposes a fast stability scanning method for future grid scenarios.
problem Capturing inter-seasonal variations in renewable generation.
method Novel feature selection algorithm and self-adaptive PSO-k-means clustering.
result Reduced computational burden up to ten times with acceptable accuracy.
New method improves reliability of LDA topic modeling by assessing stability across replicated runs.
problem LDA's reproducibility issues due to initial values and Gibbs sampling.
method Cluster replicated LDA runs using modified Jaccard coefficient and pruning algorithm.
result New measure S-CLOP quantifies LDA topic stability, improving reproducibility.
Recurrent neural networks trained on regular languages exhibit stable states that can recover from noise.
problem Stability of internal states in recurrent neural networks trained on regular languages.
method Empirical study with analysis of network activation and transitions between states.
result Recurrent neural networks trained on regular languages can recover from random perturbations and maintain stable states.
This paper reviews weighted clustering ensemble methods.
problem Improving clustering results from individual methods.
method Different types of weights and approaches to determining weight values.
result Unified framework for selecting appropriate weighting mechanisms.
This paper introduces hierarchical quasi-clustering methods, a generalization of hierarchical clustering for asymmetric networks where the output structure preserves the asymmetry of the input data. We show that this output structure is equivalent to a finite quasi-ultrametric space and study admissibility with respect…
Proposes a deterministic LIME for CAD systems.
problem Instability in LIME explanations.
method Uses agglomerative HC and KNN to select relevant clusters and trains a linear model.
result DLIME is more stable than LIME.
Bi-filtration stabilizes TDA mapper results under noise.
problem Stability issues in TDA mapper results under data perturbation.
method Introduced bi-filtration approach to stabilize mapper graphs.
result Persistent homology of perturbed data set is 2δ-interleaved with original.
New method reduces spectral clustering complexity by sparsifying graphs.
problem Computational bottleneck in spectral clustering due to eigendeomposition of NN graph Laplacian matrices.
method Spectrum-preserving graph sparsification via low-stretch spanning trees and spectral off-tree embedding.
result Ultra-sparse NN graphs with preserved first few eigenvectors for scalable spectral clustering.
Here, we propose a clustering technique for general clustering problems including those that have non-convex clusters. For a given desired number of clusters K, we use three stages to find a clustering. The first stage uses a hybrid clustering technique to produce a series of clusterings of various sizes (randomly se…
Tree-SNE combines t-SNE and hierarchical clustering for data visualization.
problem Data visualization and clustering in complex datasets.
method Stacked one-dimensional t-SNE embeddings and alpha-clustering.
result Effective hierarchical clustering and visualization of various datasets.
An adaptive clustering algorithm learns from evolving data without manual tuning.
problem Clustering in dynamic data environments where distributions change over time.
method ART-based topological clustering with self-adjusting vigilance parameter.
result The algorithm outperforms state-of-the-art methods in clustering performance and continual learning.
High density clusters can be characterized by the connected components of a level set L(λ)={x: p(x)>λ} of the underlying probability density function p generating the data, at some appropriate level λ≥0. The complete hierarchical clustering can be characterized by a cluster tree ${\cal T}= \bigcup_λ L(λ)…