DMClusts discovers multiple clusterings from multi-view data.
problem Finding multiple meaningful and diverse clusterings from multi-view data.
method Deep matrix factorization to gradually factorize multi-view data into representational subspaces and generate one clustering per layer, enforcing diversity through proximity minimization.
result DMClusts outperforms state-of-the-art multiple clustering solutions.
Framework for generating multiple clusterings from multi-view data.
problem Challenges in finding optimal clustering criteria and handling incomplete multi-view data.
method DiMVMC framework that optimizes multiple decoder deep networks to complete data views and generate shared representations.
result DiMVMC outperforms state-of-the-art competitors in generating multiple clusterings with high diversity and quality.
New algorithm clusters multiple samples from hidden distributions.
problem Traditional clustering algorithms cannot handle multiple samples from non-Gaussian distributions.
method Proposes a general framework for multiple sample clustering, generating various algorithms.
result Sufficient statistics improve clustering accuracy and stability.
Proposes MVMC for multi-view clustering, enhancing diversity and quality.
problem Leveraging multi-view data for diverse clustering.
method Adapts multi-view self-representation learning, HSIC for redundancy reduction, matrix factorization.
result Generates multiple high-quality and diverse clusterings from multi-view data.
A fair clustering method for multiple sensitive attributes is proposed.
problem Ensuring fair representation of sensitive attributes in clustering.
method FairKM (Fair K-Means) method inspired by K-Means, using fairness and coherence objectives.
result FairKM clusters significantly better on both quality and fair representation of sensitive attribute groups.
Adapts manifold structure for better clustering performance.
problem Lack of consideration for local manifold structure in existing multiple kernel k-means methods.
method Adopts manifold adaptive kernel to integrate local manifold structure of kernels.
result Proposed method outperforms state-of-the-art methods.
SimpleMKKM improves multi-kernel clustering efficiency.
problem Efficient multi-kernel clustering.
method Re-formulated minimization-maximization problem into a smooth minimization, solved with gradient descent.
result Outperforms state-of-the-art multi-kernel clustering alternatives.
MISC finds multiple independent clusterings in different subspaces.
problem Difficulties in understanding diverse clusterings.
method Two-stage approach using independent subspace analysis and graph regularized semi-nonnegative matrix factorization.
result MISC discovers different clusterings from independent subspaces.
KLIC combines multiple datasets for clustering, down-weighting noisy data.
problem Robustness of COCA in noisy or conflicting datasets.
method Multiple Kernel Learning for Integrative Clustering.
result KLIC down-weights noisy datasets, improving clustering accuracy.
A new method clusters subjects based on brain networks without vectorizing fMRI data.
problem Distortion of clustering results when simplifying fMRI data structure.
method Wishart mixture models for multiple-view clustering of brain networks.
result Identifies multiple underlying pairs of associations between subject clusters and brain sub-networks.
Proposes MFPC for cross-manifold clustering.
problem Cross-manifold clustering challenges traditional methods.
method Multiple Flat Projections Clustering (MFPC).
result MFPC distinguishes cross-manifold clusters from projections.
clusterBMA combines clustering results from multiple models using Bayesian model averaging.
problem Uncertainty in model selection for clustering.
method Bayesian model averaging to combine results from multiple clustering algorithms.
result ClusterBMA offers probabilistic cluster allocations and quantifies model-based uncertainty.
Proposes RNSE for clustering with adaptive similarity matrix learning.
problem Sub-optimal results due to mismatch between stages in Spectral Clustering.
method End-to-end single-stage learning with adaptive similarity matrix and non-negative constraints.
result Superior clustering performance on synthetic and real-world datasets.
DEMVC improves multi-view clustering with collaborative training and deep autoencoders.
problem Existing multi-view clustering methods have high computation and space complexities or lack representation capability.
method DEMVC learns embedded representations of multiple views individually using deep autoencoders and collaboratively trains all views.
result DEMVC achieves significant improvements over state-of-the-art methods on multi-view datasets.
New algorithm clusters hyperspectral images at multiple scales.
problem Clustering hyperspectral images at various scales.
method M-SRDL algorithm using spectral-spatial diffusion distances.
result More accurate clustering labels achieved with spatial regularization.
In the Pioneer 100 (P100) Wellness Project (Price and others, 2017), multiple types of data are collected on a single set of healthy participants at multiple timepoints in order to characterize and optimize wellness. One way to do this is to identify clusters, or subgroups, among the participants, and then to tailor pe…
We consider a problem of grouping multiple graphs into several clusters using singular value thesholding and non-negative factorization. We derive a model selection information criterion to estimate the number of clusters. We demonstrate our approach using "Swimmer data set" as well as simulated data set, and compare i…
Proposes a method for multi-view clustering that integrates consistent and complementary graph regularizers.
problem Multi-view clustering where views have both consistent and complementary information.
method Consistent and complementary graph-regularized multi-view subspace clustering (GRMSC).
result The proposed method outperforms state-of-the-art methods on benchmark datasets.
We propose a novel method for multiple clustering that assumes a co-clustering structure (partitions in both rows and columns of the data matrix) in each view. The new method is applicable to high-dimensional data. It is based on a nonparametric Bayesian approach in which the number of views and the number of feature-/…
To cluster data that are not linearly separable in the original feature space, k-means clustering was extended to the kernel version. However, the performance of kernel k-means clustering largely depends on the choice of kernel function. To mitigate this problem, multiple kernel learning has been introduced into th…
This paper certifies cluster assignments from sum-of-norms clustering algorithms.
problem Certifying the correct cluster assignments from approximate solutions of sum-of-norms clustering.
method Presented a clustering test that identifies and certifies the correct cluster assignment from an approximate solution.
result The correct cluster assignment is guaranteed to be certified by a primal-dual path following algorithm after sufficient iterations.
Driven by the multi-level structure of human intracranial electroencephalogram (iEEG) recordings of epileptic seizures, we introduce a new variant of a hierarchical Dirichlet Process---the multi-level clustering hierarchical Dirichlet Process (MLC-HDP)---that simultaneously clusters datasets on multiple levels. Our sei…
Paper finds methods to accurately determine the number of clusters in data.
problem Finding the correct number of clusters in a dataset.
method Penalized k-means algorithms with ideal clusters and multiplicative penalties.
result K-means with multiplicative penalties provides a clearer indication of the correct number of clusters.
MFCVAE clusters data over multiple facets, improving disentanglement and generation.
problem Clustering high-dimensional data like images over multiple characteristics.
method Variational autoencoder with hierarchical latent variables and Mixture-of-Gaussians priors.
result MFCVAE learns and clusters over multiple aspects of data in a disentangled manner.
Bagging and boosting are proved to be the best methods of building multiple classifiers in classification combination problems. In the area of "flat clustering" problems, it is also recognized that multi-clustering methods based on boosting provide clusterings of an improved quality. In this paper, we introduce a novel…
Incremental clustering approaches have been proposed for handling large data when given data set is too large to be stored. The key idea of these approaches is to find representatives to represent each cluster in each data chunk and final data analysis is carried out based on those identified representatives from all t…
GMBL uses graph embedding to learn binary codes from multiple views for clustering.
problem Lack of complete structure and complementary information from multiple views in single-view hash clustering methods.
method Graph-based Multi-view Binary Learning (GMBL) using Laplacian matrix to preserve data structure and assign weights to views.
result GMBL outperforms previous methods in clustering performance on multiple datasets.
Two new methods improve clustering with missing data.
problem Handling missing data in Gaussian Mixture Models.
method Proposes two methods using Monte Carlo Expectation-Maximization (MCEM) for data augmentation.
result Proposed methods outperform multiple imputation in clustering and density estimation.
ElbowSig assesses clustering structure at multiple scales.
problem Selecting optimal number of clusters in unsupervised learning.
method Formalizes elbow heuristic with a normalized discrete curvature statistic.
result Validates multiscale clustering structure over various resolutions.
This paper tackles fair clustering with multiple types and provides scalable coreset solutions.
problem Fair clustering with multiple, non-disjoint sensitive types.
method Novel constructions of coresets for fair clustering problems.
result First known coreset construction for fair clustering with k-median objective.
Develops a new cluster validity index to find multiple optimal cluster numbers.
problem Finding the optimal number of clusters in real-world data with varying densities, sizes, and shapes.
method A new correlation-based cluster validity index that yields multiple local peaks.
result The new index finds multiple optimal cluster numbers in various scenarios.
A new test for volatility in clustered time series data, robust to distributional assumptions.
problem Volatility issues in clustered multiple time series data, especially in stock market indicators.
method Bootstrap method for multiple time series, accounting for contagion effect.
result The test is correctly sized and powerful, especially for stationary mean and contained volatility in fewer clusters.
Proposes CRG_IMSC for better clustering of multi-view data.
problem Lack of effective connectivity in clustering results.
method Directly obtains clustering result with nonnegative constraint; constructs connectivity matrix based on spectral clustering result; uses multiplicative update algorithm.
result Improves clustering performance on benchmark datasets.
Paper detects gradual changes in cluster structure using MC fusion.
problem Detecting gradual changes in cluster structure over time.
method MC fusion for multiple mixture numbers, examining MC transition.
result Accurately captures cluster structure during transitional periods.
We present a robust multiple manifolds structure learning (RMMSL) scheme to robustly estimate data structures under the multiple low intrinsic dimensional manifolds assumption. In the local learning stage, RMMSL efficiently estimates local tangent space by weighted low-rank matrix factorization. In the global learning …
Proposes a novel multi-view clustering method by aligning partitions.
problem Challenges of integrating multi-view information and information loss.
method Aligns partitions through rotation matrices and assigns weights to views.
result Significant improvement over state-of-the-art methods on real datasets.
The paper quantizes concatenated noisy vectors to a common cluster center, improving performance over naive methods.
problem Clustering concatenated noisy vectors from multiple sources.
method Asymptotic analysis of weighted sum of distances to a common cluster center.
result The clustering approach outperforms naive methods in terms of average distortion.
MLCC clusters data at multiple significance levels, detecting anomalies without distributional assumptions.
problem Clustering and anomaly detection in data with unknown distributions.
method Hierarchical, conformal prediction-based clustering.
result MLCC automatically selects cluster number and detects anomalies robustly.
This paper addresses clustering with missing data using Rubin's rules.
problem How to pool partitions and assess instability when data are incomplete after multiple imputation.
method Consensus clustering and bootstrap theory are used to address the problem.
result New rules for pooling partitions and assessing instability are proposed and validated.
Bayesian models offer great flexibility for clustering applications---Bayesian nonparametrics can be used for modeling infinite mixtures, and hierarchical Bayesian models can be utilized for sharing clusters across multiple data sets. For the most part, such flexibility is lacking in classical clustering methods such a…
Fair HAC algorithms ensure clustering fairness across protected groups.
problem Ensuring clustering fairness in HAC algorithms when datasets contain biases.
method Proposes fair algorithms for HAC that enforce fairness constraints regardless of distance linkage criteria.
result Our fair HAC algorithms find fairer clusterings compared to vanilla HAC and other fair clustering approaches.
We consider grouping as a general characterization for problems such as clustering, community detection in networks, and multiple parametric model estimation. We are interested in merging solutions from different grouping algorithms, distilling all their good qualities into a consensus solution. In this paper, we propo…
New model for clustering graphs with multiple data sources.
problem Graph clustering with multiple data sources.
method Formalized multi-view stochastic block models and developed efficient algorithms.
result Provable improvement over previous approaches in multi-view graph clustering.
Modeling price clustering in financial markets using discrete distributions.
problem Price clustering phenomenon in financial markets.
method Discrete price model based on mixture of double Poisson distributions with dynamic volatility and proportions.
result Higher instantaneous volatility weakens price clustering at ultra-high frequencies.
A method to improve clustering explainability using bagging and feature dropout.
problem Lack of explainability in clustering methods.
method Bagging and feature dropout to generate feature importance scores.
result Improved stability and robustness of cluster definition, especially in small-sample or noisy settings.
ALPCAHUS clusters data from multiple subspaces with varying noise.
problem Heteroscedastic data with varying noise levels.
method Develops a heteroscedastic PCA method for subspace clustering.
result Improves subspace clustering by accounting for sample-wise noise variances.
The clustering ensemble technique aims to combine multiple clusterings into a probably better and more robust clustering and has been receiving an increasing attention in recent years. There are mainly two aspects of limitations in the existing clustering ensemble approaches. Firstly, many approaches lack the ability t…
New locally private algorithm for k-means clustering reduces additive error significantly.
problem Designing a locally private algorithm for k-means clustering with reduced additive error.
method Local differential privacy approach, reducing additive error to nearly n1/2. result Achieves O(1) multiplicative error and n1/2+a additive error, nearly optimal.