Discussing issues in robust clustering, especially with Gaussian models.
problem Handling outliers and ambiguity in clustering groups.
method Focus on Gaussian mixture model, examining formal definitions, interactions, and tuning decisions.
result Outliers can confuse clustering groups and existing stability measures fail with them.
Total variation minimization clusters partially labeled data points.
problem Clustering partially labeled data points in stochastic block models.
method Total variation minimization as a clustering method.
result Total variation minimization allows for accurate clustering under certain model parameters.
Paper proposes a new clustering model that preserves cluster recovery with fewer dimensions.
problem Clustering high-dimensional data with limited embedding dimensions.
method Randomly projected convex clustering model with improved embedding dimension.
result Cluster recovery can be preserved with fewer dimensions, independent of data points.
Proposes a new model for unsupervised clustering with latent variables.
problem The challenge of unsupervised clustering in machine learning.
method Clustered Generator Model with continuous and discrete latent variables.
result Achieves competitive unsupervised clustering accuracy and disentangled latent representations.
New random models improve clustering similarity assessment.
problem Improper random models affect clustering similarity assessments.
method Derived corrected Rand index and Mutual Information measures for varying cluster sizes.
result Random model choice drastically impacts clustering similarity rankings.
AMOS automates model order selection for spectral graph clustering.
problem Automated selection of the correct number of clusters in spectral graph clustering.
method Incrementally increases the number of clusters, estimates cluster quality, and provides reliability tests.
result AMOS outputs clusters of minimal model order with statistical guarantees.
Paper introduces methods to use mixture models for modal clustering.
problem Clarity between clusters and mixture components in nonparametric mixture modeling.
method Two methods to adopt modal clustering after mixture model fitting.
result Mixture modeling can be used for clustering in a nonparametric sense.
Paper introduces NICc for fast cluster-based validation of prediction models.
problem Validation of prediction models on clustered data.
method Derived NICc to approximate leave-one-cluster-out deviance for standard regression models.
result NICc provides more accurate model size and variable selection, especially with strong clustering.
clusterBMA combines clustering results from multiple models using Bayesian model averaging.
problem Uncertainty in model selection for clustering.
method Bayesian model averaging to combine results from multiple clustering algorithms.
result ClusterBMA offers probabilistic cluster allocations and quantifies model-based uncertainty.
This study evaluates cluster search algorithms using Gaussian mixture models.
problem Determining the optimal number of clusters in data sets generated by Gaussian mixture models.
method Examined centroid- and model-based cluster search algorithms in various cases.
result Model-based algorithms are more robust to cluster overlap and covariance type than centroid-based methods.
A new log-volatility factor model reduces dimensionality and identifies cluster contributions to volatility clustering.
problem Understanding the sources of volatility clustering in financial markets.
method Introduced a new factor model using Directed Bubble Hierarchical Tree (DBHT) to identify the number of factors and integrated non-parametric proxy to study volatility clustering.
result Clusters contribute to volatility clustering locally, while the market contributes globally.
Document clustering and topic modeling are two closely related tasks which can mutually benefit each other. Topic modeling can project documents into a topic space which facilitates effective document clustering. Cluster labels discovered by document clustering can be incorporated into topic models to extract local top…
A new vine copula mixture model improves clustering accuracy for non-Gaussian data.
problem Finite mixture models struggle with asymmetric tail dependencies and non-elliptical clusters.
method Proposes a vine copula mixture model for clustering non-Gaussian data, addressing model selection and parameter estimation.
result Significant improvement in clustering accuracy for data with asymmetric tail dependencies or non-Gaussian margins.
This paper examines variable selection for clustering using Gaussian mixture models.
problem Modern databases require efficient variable selection for clustering models.
method Recalls basics of clustering, examines variable selection methods for model-based clustering.
result Opportunities for improving variable selection methods are presented.
Bayesian distance clustering improves robustness to kernel choice.
problem Kernel sensitivity in model-based clustering.
method Modeling pairwise distances instead of original data.
result Dramatic gains in cluster inference robustness.
Model-based clustering defines population level clusters relative to a model that embeds notions of similarity. Algorithms tailored to such models yield estimated clusters with a clear statistical interpretation. We take this view here and introduce the class of G-block covariance models as a background model for varia…
Proposes a general model for plane-based clustering with a new loss function.
problem Clustering of data points in a plane-based approach.
method General model containing existing methods, optimization problem with total loss minimization.
result The proposed method effectively clusters data points with a new loss function.
C3L clusters data with user-controlled leakage, improving semi-supervised models.
problem Finding clusters in partially categorized data sets.
method Semi-supervised Gaussian mixture model with user-defined leakage level.
result C3L finds high-quality clustering models with controlled inconsistency.
DPMM-CFL clusters clients for federated learning without fixed K, improving performance.
problem Improving federated learning performance under non-IID client heterogeneity.
method DPMM-CFL uses a Dirichlet Process Mixture Model to infer both cluster number and client assignments.
result DPMM-CFL optimizes per-cluster federated objectives and jointly infers cluster number and assignments.
A neural-network model clusters subjects based on their lifetime distributions.
problem Clustering subjects into clusters based on their lifetime distributions.
method A neural-network based lifetime clustering model that maximizes divergence between empirical lifetime distributions of clusters.
result Significantly better lifetime clusters compared to competing approaches.
Paper solves spectral clustering's dimensionality and complexity model selection issues.
problem Model selection problems in spectral graph clustering.
method Developed a probabilistic model and simultaneous model selection framework.
result Consistent estimates of model parameters for embedding dimension and number of clusters.
Meta-learning model clusters data better than standard methods.
problem Challenges in clustering due to custom loss functions and small datasets.
method Trains a recurrent model to learn clustering from various datasets.
result Meta clustering model outperforms standard clustering techniques on unseen datasets.
New deep learning framework for tabular data clusters with interpretable features.
problem Need for reliable and interpretable clustering models for tabular data.
method Self-supervised feature selection and gate matrix for cluster-level feature selection.
result Model provides interpretable cluster assignments with driving features.
Study finds the cutoff for exact recovery in Gaussian mixture models.
problem Determining the separation of cluster centers for exact recovery in Gaussian mixture models.
method Used information theory and SDP relaxation of K-means clustering. result Sharp threshold for exact recovery of cluster labels without assuming cluster center symmetry.
Paper characterizes optimal graph clustering limits under a new model.
problem Graph clustering under varying edge density signals.
method Introduced Popularity-Adjusted Block Model (PABM) to address SBM and DCBM limitations.
result Cluster recovery possible even when edge density signals vanish, highlighting local connectivity differences.
Paper detects gradual changes in cluster structure using MC fusion.
problem Detecting gradual changes in cluster structure over time.
method MC fusion for multiple mixture numbers, examining MC transition.
result Accurately captures cluster structure during transitional periods.
Proposes FMC for fair clustering with independent parameters.
problem Finding clusters with balanced sensitive attribute proportions.
method Model-based clustering using finite mixture model with mini-batch learning.
result FMC scales up easily and can handle non-metric data.
FBC clusters data fairly without needing cluster count.
problem Fairness in clustering groups of different sensitive groups.
method Developed a Bayesian model-based clustering method with a fair prior and efficient MCMC algorithm.
result Reasonably infers the number of clusters and achieves a fair utility trade-off.
New model clusters set-valued data without specifying cluster number.
problem Clustering set-valued data and unknown number of clusters.
method Dirichlet Process mixture of Poisson random finite sets, MCMC inference.
result Model discovers extremely unbalanced clusters.
New clustering algorithms capture time-evolving clusters using Markov models.
problem Capturing time-evolving clusters in data.
method Small-variance asymptotic analysis of Markov chain mixture models.
result Two clustering algorithms (D-Means and SD-Means) outperform existing methods in accuracy and computational cost.
Integrates VAEs into EM for deep clustering and generation.
problem Clustering and generating new samples from complex distributions.
method Combines VAEs and EM, updating model parameters and refining cluster assignments.
result Superior clustering performance on MNIST and FashionMNIST.
Paper analyzes conditions for clustering BMMs with unknown clusters.
problem Clustering Bernoulli Mixture Models (BMMs) with unknown number of clusters.
method Theoretical analysis of sample complexity and dimensionality for PAC-clusterability.
result First non-asymptotic bounds on sample complexity for learning or clustering BMMs.
A new method clusters time series based on model prediction accuracy.
problem Clustering time series data effectively.
method Iterative model fitting and assignment based on predictive accuracy.
result The method outperforms other techniques in clustering and predictive accuracy.
New model clusters discrete time series data.
problem Handling discreteness and time series properties in data.
method Finite mixture model with INAR type models.
result Demonstrated clustering on real data.
New method clusters longitudinal data with many time points.
problem Clustering longitudinal data with many time points.
method Extension of the mixture of common factor analyzers model with expectation-maximization algorithm for parameter estimation and Bayesian information criterion for model selection.
result The approach effectively clusters longitudinal data with many time points.
Improved algorithm for clustered Federated Learning reduces initialization and hyperparameter requirements.
problem Dichotomy between heterogeneous models and simultaneous training in Federated Learning.
method Proposes a new clustering framework and an improved algorithm ( exttt{SR-FCA}) that removes restrictive assumptions.
result Improves clustering accuracy and removes the need for good initialization and hyperparameters.
ADEC addresses feature randomness and drift in autoencoder-based clustering.
problem Clustering autoencoders learn unreliable pseudo-labels, distorting latent space and feature randomness.
method Adversarial training to balance reconstruction loss and clustering objective.
result ADEC outperforms state-of-the-art autoencoder-based clustering methods.
ARMED models improve deep learning interpretability and generalize better on clustered data.
problem Clustered data leads to spurious associations and poor model fitting.
method Adversarial regularization and mixed effects subnetworks.
result ARMED models outperform conventional methods in accuracy and generalization.
The paper tackles the problem of identifying same-cluster elements in overlapping clusters with minimum queries.
problem Identifying same-cluster elements in overlapping clusters with minimum queries.
method The paper provides algorithms for identifying same-cluster elements in overlapping clusters with minimum queries, under both arbitrary and statistical modeling assumptions.
result The algorithms are order optimal, parameter free, efficient, and work in the presence of random noise.
New models for microclustering address entity resolution by allowing cluster sizes to grow sublinearly.
problem Need for models where cluster sizes grow sublinearly with data set size.
method Definition of microclustering property and introduction of new models.
result New models can yield clusters whose sizes grow sublinearly, unlike traditional clustering models.
Review of variable selection methods for model-based clustering.
problem Dealing with high-dimensional data in model-based clustering.
method Variable selection techniques to facilitate interpretation.
result Summary and illustration of existing methods.
We present an approach to model-based hierarchical clustering by formulating an objective function based on a Bayesian analysis. This model organizes the data into a cluster hierarchy while specifying a complex feature-set partitioning that is a key component of our model. Features can have either a unique distribution…
K-ARMA models cluster time series data robustly.
problem Clustering time series data effectively.
method Model-based K-ARMA clustering algorithm with robust outlier detection.
result K-ARMA models outperform existing methods for time series clustering.
Proposes a new model for clustering passenger trips considering hierarchical and multi-dimensional data.
problem Clustering passenger trips with hierarchical and multi-dimensional data, especially in large-scale transportation systems.
method Tensor Dirichlet Process Multinomial Mixture (Tensor-DPMM) model, incorporating Dirichlet Process for automatic cluster number determination and tensor representation for multi-mode data.
result Automatic determination of the number of clusters and improved clustering quality.
The paper proposes a parallelizable clustering method for multivariate data.
problem The standard model-based clustering method assumes the same number of clusters per margin, which is often unrealistic.
method Developed a finite mixture model per margin with different numbers of clusters, and used a game-inspired algorithm to cluster multivariate data.
result The proposed method shows good performance in various scenarios and real datasets.
New criterion selects optimal number of clusters based on stability.
problem Challenges in selecting optimal number of clusters in non-parametric clustering.
method Proposes a stability-based validation criterion combining between-cluster and within-cluster stability.
result Empirically demonstrates effectiveness in selecting optimal number of clusters.
Cluster LOCO: A model-agnostic feature importance score for interpreting cluster outputs
problem Interpreting and auditing cluster outputs
method Cluster LOCO (Leave-One-Covariate-Out)
result More reliably recovers informative features than existing methods
A new clustering framework using fixed points for data analysis.
problem Lack of unified understanding and application of clustering algorithms in data analysis.
method Restated model-based clustering using fixed point theory, iteratively constructing contraction maps to find cluster centers.
result Unified clustering framework reveals convergence mechanisms and interconnections among clustering algorithms.