The paper improves spectral clustering by analyzing the asymptotic normalised cut value.
problem No agreed method for tuning scaling parameter or automatically determining cluster number.
method Investigates asymptotic value of normalised cut for increasing samples.
result Provides recommendations for improving spectral clustering methodology.
Improved clustering algorithm for large datasets.
problem Finding alternative partitions in large datasets.
method Iterative Spectral Method (ISM) for alternative clustering.
result Significantly improved scalability and computation time.
Proposes ConiVAT for better cluster assessment and clustering with background knowledge.
problem Challenges in cluster assessment and clustering with noise and bridge points.
method Uses background constraints to improve VAT/iVAT for complex datasets.
result Improves clustering accuracy and resolves issues with noise and bridge points.
A scalable framework for clustering large graphs using randomized sketching.
problem Clustering large partially observed graphs efficiently.
method Randomized graph sketching, correlation-based retrieval, uniform and degree-based node sampling.
result Improved phase transitions for clustering with reduced computational complexity and minimum cluster size.
SESSC clusters fuzzy rules for TSK classifiers, improving performance with label info.
problem Lack of supervised clustering for TSK fuzzy classifiers.
method SESSC integrates within-cluster compactness, between-cluster separation, and label information.
result SESSC initialization outperforms other clustering methods, especially for small rule numbers.
A novel approach ODAR detects outliers for clustering.
problem Outliers interfere with clustering algorithms, leading to unreliable results.
method Feature transformation to separate outliers and normal objects into distinct clusters.
result ODAR improves clustering accuracy on 7 out of 10 datasets.
IMPACC improves consensus clustering for bioinformatics data.
problem Consensus clustering's inefficiency and lack of interpretability for large-scale data.
method Ensemble minipatch co-occurrences, adaptive sampling of observations and features.
result Significantly improved accuracy and interpretability with substantial computational savings.
Deep image clustering improved with STN and DAC.
problem Challenges in clustering images, especially with spatial transformations.
method Combining DAC with STN to reduce spatial transformation issues.
result The combined model outperformed baseline models on MNIST and FashionMNIST.
Dual regularized graph Laplacian improves spectral clustering for community detection.
problem Detecting clusters in networks with improved spectral clustering methods.
method Proposes dual regularized graph Laplacian for three spectral clustering approaches.
result Theoretical analysis shows DRSC and DRSLIM yield stable consistent community detection.
Posterior regularization enhances Bayesian hierarchical mixture clustering by improving node separation.
problem High nodal variance in BHMC trees, leading to weak separation between nodes at higher levels.
method Employing Posterior Regularization to impose max-margin constraints on nodes at every level.
result Improves cluster separation in BHMC models, enhancing overall model performance.
Graph clustering uses multiscale community detection for improved performance.
problem Improving data clustering accuracy and robustness.
method Graph-theoretical approach combining multiscale community detection.
result Multiscale graph-based clustering achieves better performance than traditional methods.
Improved spectral clustering for community detection in networks.
problem Community detection in networks.
method Improved spectral clustering (ISC) based on k-means clustering on weighted eigenvectors of a regularized Laplacian matrix.
result ISC yields stable consistent community detection under mild conditions and outperforms classical methods.
DEMVC improves multi-view clustering with collaborative training and deep autoencoders.
problem Existing multi-view clustering methods have high computation and space complexities or lack representation capability.
method DEMVC learns embedded representations of multiple views individually using deep autoencoders and collaboratively trains all views.
result DEMVC achieves significant improvements over state-of-the-art methods on multi-view datasets.
New method improves cluster instability estimation for better k selection.
problem Selecting the optimal number of clusters in cluster analysis.
method Developed a normalized cluster instability measure to correct for cluster size distribution.
result Normalized instability measure outperforms current methods across all possible k. Cluster jackknife improves inference for staggered DID methods.
problem Over-rejection of CSDID in small clusters or treated clusters.
method Cluster jackknife for CSDID inference.
result Cluster jackknife greatly improves inference for CSDID.
Survey of flow-based algorithms for improving clusters.
problem Improving clusters obtained by other methods.
method Flow-based algorithms solving maximum flow problems.
result Efficient implementations and extensive numerical experiments.
Improved clustering accuracy with disentangled latent code representation.
problem Improving k-Means clustering performance.
method Optimizing the entanglement of autoencoder latent code representation using soft nearest neighbor loss with annealing temperature.
result 96.2% test clustering accuracy on MNIST, 85.6% on Fashion-MNIST, and 79.2% on EMNIST Balanced datasets.
Transform learning improves K-means clustering for document analysis.
problem Improving K-means clustering for document analysis.
method Embedding K-means clustering loss into transform learning framework and solving jointly using ADMM.
result Improves over state-of-the-art in document clustering.
VC-PCR improves prediction by clustering correlated variables.
problem Decreased prediction accuracy due to cluster structure in predictor variables.
method Supervised variable selection and clustering to integrate cluster information into a sparse modeling process.
result VC-PCR achieves better prediction, variable selection, and clustering performance.
Although many convex relaxations of clustering have been proposed in the past decade, current formulations remain restricted to spherical Gaussian or discriminative models and are susceptible to imbalanced clusters. To address these shortcomings, we propose a new class of convex relaxations that can be flexibly applied…
New clustering algorithm for mixed data improves applicability and efficiency.
problem Clustering large, mixed data with improved accuracy and efficiency.
method Developed a new clustering algorithm using peak-finding technique, reducing computational complexity.
result Algorithm detects outliers, clusters of lower density, and determines correct number of clusters.
Accelerated algorithm for density-based clustering outperforms DBSCAN.
problem Density-based clustering with variable density clusters and tuning parameters.
method Accelerated HDBSCAN* algorithm improving upon HDBSCAN*.
result Comparable performance to DBSCAN, eliminates tuning parameter.
A new clustering method using autoencoders for improved data representation.
problem Improving clustering of complex data like images and text.
method DAMIC algorithm based on a mixture of deep autoencoders.
result Significant improvement over state-of-the-art methods on image and text corpora.
New A* algorithm improves hierarchical clustering quality.
problem Improving hierarchical clustering quality in large search spaces.
method Combining A* search with a trellis data structure.
result Achieves higher quality results than baselines in particle physics and other benchmarks.
The paper compares clustering methods for improving time series forecasting accuracy.
problem Improving time series forecasting accuracy using neural networks.
method Investigates feature-based and distance-based clustering methods for time series forecasting.
result Feature-based clustering outperforms distance-based clustering in terms of speed and efficiency.
A new clustering method using Bayesian techniques improves robustness and interpretability.
problem Improving clustering techniques for better robustness and interpretability.
method The paper proposes a novel Bayesian clustering method using the proper Bayesian bootstrap, which combines k-means clustering and ensemble clustering.
result The method provides clear indication on the optimal number of clusters and a better representation of the clustered data.
A novel multi-clustering method based on boosting improves hierarchical clustering quality.
problem Improving hierarchical clustering quality in flat clustering problems.
method A boosting iteration with weighted random sampling of elements from the original dataset, followed by hierarchical clustering on each subsample and consensus combination.
result The proposed method provides superior quality solutions compared to standard hierarchical clustering methods.
DPMM-CFL clusters clients for federated learning without fixed K, improving performance.
problem Improving federated learning performance under non-IID client heterogeneity.
method DPMM-CFL uses a Dirichlet Process Mixture Model to infer both cluster number and client assignments.
result DPMM-CFL optimizes per-cluster federated objectives and jointly infers cluster number and assignments.
Cluster-DP improves differential privacy in randomized experiments by clustering data.
problem Reducing variance in causal effect estimation from differentially private data.
method Cluster-DP leverages a given cluster structure to improve the privacy-variance trade-off.
result Selecting higher-quality clusters decreases the variance penalty without compromising privacy guarantees.
Improved algorithm for clustered Federated Learning reduces initialization and hyperparameter requirements.
problem Dichotomy between heterogeneous models and simultaneous training in Federated Learning.
method Proposes a new clustering framework and an improved algorithm ( exttt{SR-FCA}) that removes restrictive assumptions.
result Improves clustering accuracy and removes the need for good initialization and hyperparameters.
Improves clustering performance by mixing latent representations.
problem Finding well-defined clusters in data representations.
method Mixing Consistent Deep Clustering method that encourages realistic interpolations and semantic consistency.
result Improved clustering performance across various models and datasets.
Elastic co-clustering improves clustering of single-cell genomic data.
problem Improving clustering performance of single-cell genomic datasets.
method Elastic coupled co-clustering in an unsupervised transfer learning framework.
result Our algorithm significantly improves clustering performance over traditional methods.
Improved spectral clustering algorithm for better performance.
problem Improving the performance of spectral clustering algorithms.
method Developed a new performance guarantee under a weaker assumption and evaluated using a different spectral embedding map.
result Better performance guarantee under a weaker assumption and evaluation of a new spectral embedding map.
NLSSC improves clustering by enhancing separability in sparse coding.
problem Improving clustering performance in subspace clustering problems.
method Introduces a novel objective term for local separability in non-negative local sparse coding.
result NLSSC outperforms state-of-the-art methods in clustering benchmarks.
This paper proposes a spectral clustering algorithm for hyperbolic spaces, improving efficiency over Euclidean methods.
problem Inefficient clustering in Euclidean spaces for complex data structures.
method Developed a spectral clustering algorithm using hyperbolic similarity matrices.
result The algorithm converges at least as fast as Euclidean spectral clustering and performs better on complex datasets.
Enhances clustering by using external categorical data.
problem Improving clustering outcomes using external categorical evidence.
method Evidence transfer method that manipulates autoencoder latent representations based on external categorical data.
result Our method effectively manipulates latent representations with real evidence and remains robust with low quality evidence.
A new algorithm improves SSC clustering accuracy with low complexity.
problem Sparse Subspace Clustering accuracy loss in time efficiency.
method Active Orthogonal Matching Pursuit (Active OMP-SSC) for improved clustering accuracy.
result Improves clustering accuracy of OMP-SSC with low computational complexity.
New algorithm improves online clustering of bandits with minimal frequency constraints.
problem Online clustering of bandits with non-uniform user frequencies.
method Proposes an efficient algorithm with simple set structures to represent clusters, proving a regret bound free of minimal frequency constraints.
result The new algorithm consistently outperforms existing methods in experiments on synthetic and real datasets.
C3L clusters data with user-controlled leakage, improving semi-supervised models.
problem Finding clusters in partially categorized data sets.
method Semi-supervised Gaussian mixture model with user-defined leakage level.
result C3L finds high-quality clustering models with controlled inconsistency.
COBRAS-TS improves semi-supervised clustering for time series.
problem Semi-supervised clustering of time series data.
method Adapting COBRAS for time series data with semi-supervision.
result COBRAS-TS outperforms existing methods in time series clustering.
ARMED models improve deep learning interpretability and generalize better on clustered data.
problem Clustered data leads to spurious associations and poor model fitting.
method Adversarial regularization and mixed effects subnetworks.
result ARMED models outperform conventional methods in accuracy and generalization.
Proposes an MTL method with clustering to improve regression accuracy.
problem Improving regression accuracy by sharing information among related tasks.
method Centroid parameter for clustering tasks, separating regression and clustering parameters.
result Improves estimation and prediction accuracy for regression coefficient vectors.
PET-TURTLE improves clustering accuracy for imbalanced data.
problem Imbalanced data causes clustering errors.
method Generalizes cost function and introduces sparse logits.
result PET-TURTLE enhances overall clustering accuracy for imbalanced data.
A new method, InfoGuide, improves automatic clustering analysis.
problem Lack of automatic clustering analysis frameworks.
method Capturing traces of information gain between clustering retrievals.
result InfoGuide can enable more automatic clustering analysis.
New methods improve clustering accuracy in noisy data sets.
problem Improving clustering accuracy in data sets with noise features.
method Feature rescaling factors to enhance clustering validity indexes.
result Our methods increase the likelihood of estimating the true number of clusters.
A novel k-means method for MNAR data improves clustering accuracy.
problem Improving k-means clustering for data missing not at random.
method A magnitude-decaying MNAR scenario-based k-means method with size constraints.
result The method reduces bias in estimated cluster centers and improves clustering accuracy.
Proposes Contrastive Clustering for improved clustering performance.
problem Improving clustering performance on various datasets.
method Instance- and cluster-level contrastive learning through data augmentations and feature space projections.
result Contrastive Clustering achieves significant improvements over 17 competitive methods.
Cluster-aware model improves generative performance with unlabeled data.
problem Lack of labelled data in real-world datasets.
method Uses unlabelled data to infer latent clustering, labelled data to refine.
result Significant improvement in log-likelihood (-79.38 nats on MNIST).