Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

57113170226 · Jun 202019922001200920182026
48 results for clustering stability

New criterion selects optimal number of clusters based on stability.

problem Challenges in selecting optimal number of clusters in non-parametric clustering.
method Proposes a stability-based validation criterion combining between-cluster and within-cluster stability.
result Empirically demonstrates effectiveness in selecting optimal number of clusters.

reval package selects best clustering solutions via stability-based validation.

problem Challenges in determining best clustering solutions due to lack of validation methods.
method Stability-based relative clustering validation methods.
result Determines best clustering solutions that generalize to unseen data.

Proposes a method to predict cluster number and cluster representatives using cluster stability analysis.

problem Determining the number of clusters in a dataset.
method Analyzes cluster stability using Monte-Carlo simulation to predict cluster number and find cluster representatives.
result Significant improvement in predicting cluster numbers and cluster composition in large datasets.

Study examines how cluster number affects short-text clustering, introducing a stability metric.

problem Challenges in finding meaningful clusters in short-text data.
method Introduces a stability metric to determine cluster robustness and visualizes cluster subdivisions.
result Choosing a cluster number involves balancing informativeness and complexity, not seeking a single 'optimal' solution.

Proposes a method to assess clustering stability in financial time series.

problem Assessing the validity of clustering in financial data.
method Empirical framework with data perturbations.
result Provides insights into assets' clustering behavior through multi-view analysis.

Study compares discrete data clustering methods using Multinomial distribution.

problem Discrete data clustering with Multinomial distribution.
method Model-Based Clustering (MBC) framework with Multinomial distribution, combining partitional and hierarchical clustering techniques.
result Proposed method is competitive in clustering accuracy and better in stability and computation time.

A popular method for selecting the number of clusters is based on stability arguments: one chooses the number of clusters such that the corresponding clustering results are "most stable". In recent years, a series of papers has analyzed the behavior of this method from a theoretical point of view. However, the results …

2010-07-07abs ↗pdf ↗

CARVE validates clustering results using resampling and stability analysis.

problem Inconsistent and unreliable clustering results due to algorithm, preprocessing, and kk sensitivity.
method CARVE uses resampling-based validation and stability analysis to evaluate multiple clustering algorithms and hyperparameters.
result CARVE consistently recovers near-optimal clusterings and finer biological structure.

Study polynomials' root clustering via symmetric products, unifying stability approaches.

problem Stability of polynomials with roots in specific regions.
method Interpreting roots and coefficients as symmetric product morphism, analyzing topology up to homeomorphism.
result Description of strata topology and adjacency, explaining classical stability problems.

This paper introduces cluster exchange groupoids for Coxeter-Dynkin diagrams and finds their fundamental groups are braid groups.

problem Understanding the fundamental groups of cluster exchange groupoids for Coxeter-Dynkin diagrams.
method Introduced cluster exchange groupoids for Coxeter-Dynkin diagrams and showed the fundamental group isomorphic to braid groups.
result The fundamental group of the exchange groupoid for a Coxeter-Dynkin diagram is the braid group associated with the diagram.

We assess cluster stability by trimming extreme points and tracking data range reduction.

problem Assessing stability of one-dimensional clusters.
method Probabilistic method using diameter-shrinkage ratio to track data range reduction.
result Our method achieves higher accuracy than classical tests in small or noisy samples.

The paper provides presentations for mapping class groups and cluster automorphism groups of surfaces.

problem Presentations of mapping class groups of surfaces stabilizing boundaries.
method Gave presentations of mapping class groups of marked surfaces stabilizing boundaries.
result Presented cluster automorphism groups of cluster algebras from surfaces.

Stable density-based clustering via multiparameter persistence.

problem Density-based clustering stability to data perturbations.
method Degree-Rips construction, correspondence-interleaving distance, multiparameter stability analysis.
result Persistable pipeline yields stable, consistent density-based clustering.

We investigate the role of the initialization for the stability of the k-means clustering algorithm. As opposed to other papers, we consider the actual k-means algorithm and do not ignore its property of getting stuck in local optima. We are interested in the actual clustering, not only in the costs of the solution. We…

2009-07-31abs ↗pdf ↗

Study clusters Kenyan medical insurance companies based on financial performance and reporting consistency.

problem Identifying financial health and reporting consistency in Kenyan medical insurance companies.
method Advanced clustering techniques (KMeans, DTW) on financial ratios and time series data.
result Four distinct clusters identified, each representing different financial performance and reporting consistency combinations.

Characterizes pseudo-Anosov mapping classes using cluster algebra techniques.

problem Characterize pseudo-Anosov mapping classes purely in terms of shear coordinates.
method Uses cluster algebraic generalization and tropical cluster transformations.
result Algebraic entropies of cluster transformations match topological entropy.

S3VDC improves DC methods for scalability, stability, and simplicity.

problem Poor scalability, instability, and lack of simplicity in DC methods.
method Four algorithmic improvements: initial γγ-training, periodic ββ-annealing, mini-batch GMM initialization, and inverse min-max transform. S3VDC incorporates all improvements.
result S3VDC outperforms state-of-the-art methods on benchmark and industrial datasets.

Paper translates train track concepts to cluster algebras for pseudo-Anosov mapping classes.

problem Understanding pseudo-Anosov mapping classes on surfaces.
method Using Goncharov--Shen's potential function, the paper translates train track concepts into cluster algebra language.
result Proves sign stability of general pseudo-Anosov mapping classes.

Sparse Convex Clustering improves clustering performance in high-dimensional data.

problem Distortion in convex clustering performance with uninformative features.
method Introduces Sparse Convex Clustering with an adaptive group-lasso penalty and a tuning criterion based on clustering stability.
result Demonstrates improved clustering performance through feature selection.

MAS scores cluster size consistency from points, robust to label changes.

problem Desired uniformity in cluster sizes, stability under label perturbations.
method Mass Agreement Score (MAS) measures point-centric cluster size consistency, robust to label changes.
result MAS yields similar scores for partitions with similar bulk structure, sensitive to genuine redistribution of cluster mass.

New models learn stable latent clusters without side info.

problem Stability of non-linear ICA representations without side information.
method Deep generative models with latent clusterings, compared to standard VAEs and auxiliary labeled models.
result Deep generative models with latent clusterings are as stable as models with side information.

The paper provides guarantees for clustering validity without distributional assumptions.

problem Validating clustering results without distributional assumptions.
method Generic method to obtain post-inference guarantees of near-optimality and stability for clustering.
result The guarantees do not depend on distributional assumptions but depend on the data set admitting a stable clustering.

Clust-PSI-PFL uses PSI to improve accuracy and fairness in federated learning.

problem Non-IID data biases federated learning performance.
method Clust-PSI-PFL uses clustering and PSI to form homogeneous groups of clients.
result Clust-PSI-PFL delivers up to 18% higher global accuracy and improves client fairness.

New method improves reliability of LDA topic modeling by assessing stability across replicated runs.

problem LDA's reproducibility issues due to initial values and Gibbs sampling.
method Cluster replicated LDA runs using modified Jaccard coefficient and pruning algorithm.
result New measure S-CLOP quantifies LDA topic stability, improving reproducibility.

Recurrent neural networks trained on regular languages exhibit stable states that can recover from noise.

problem Stability of internal states in recurrent neural networks trained on regular languages.
method Empirical study with analysis of network activation and transitions between states.
result Recurrent neural networks trained on regular languages can recover from random perturbations and maintain stable states.

This paper introduces hierarchical quasi-clustering methods, a generalization of hierarchical clustering for asymmetric networks where the output structure preserves the asymmetry of the input data. We show that this output structure is equivalent to a finite quasi-ultrametric space and study admissibility with respect…

2014-04-17abs ↗pdf ↗

New method reduces spectral clustering complexity by sparsifying graphs.

problem Computational bottleneck in spectral clustering due to eigendeomposition of NN graph Laplacian matrices.
method Spectrum-preserving graph sparsification via low-stretch spanning trees and spectral off-tree embedding.
result Ultra-sparse NN graphs with preserved first few eigenvectors for scalable spectral clustering.

Here, we propose a clustering technique for general clustering problems including those that have non-convex clusters. For a given desired number of clusters KK, we use three stages to find a clustering. The first stage uses a hybrid clustering technique to produce a series of clusterings of various sizes (randomly se…

2015-03-04abs ↗pdf ↗

An adaptive clustering algorithm learns from evolving data without manual tuning.

problem Clustering in dynamic data environments where distributions change over time.
method ART-based topological clustering with self-adjusting vigilance parameter.
result The algorithm outperforms state-of-the-art methods in clustering performance and continual learning.

High density clusters can be characterized by the connected components of a level set L(λ)={x: p(x)>λ}L(λ) = \{x:\ p(x)>λ\} of the underlying probability density function pp generating the data, at some appropriate level λ0λ\geq 0. The complete hierarchical clustering can be characterized by a cluster tree ${\cal T}= \bigcup_λ L(λ)…

2010-11-11abs ↗pdf ↗