KCoreMotif clusters large networks efficiently by exploiting k-core decomposition and motifs.
problem Efficiently clustering large networks for trust evaluation.
method Exploits k-core decomposition and motifs to perform motif-based spectral clustering on k-core subgraphs.
result The proposed algorithm is accurate and efficient for large networks.
Empty core found in max-loss non-centroid clustering.
problem Core stability in non-centroid clustering under max-loss objective.
method Proof for all k≥3 and n≥9 agents, computer-aided proof for 2D Euclidean points.
result Core can be empty in non-centroid clustering under max-loss objective.
Density mode clustering is a nonparametric clustering method. The clusters are the basins of attraction of the modes of a density estimator. We study the risk of mode-based clustering. We show that the clustering risk over the cluster cores --- the regions where the density is high --- is very small even in high dimens…
COREclust detects representative variables in high-dimensional data.
problem Detecting representative variables in high-dimensional data with limited observations.
method CORE-clustering algorithm detects CORE-clusters, variable sets with similar variables, and representative variables are estimated as CORE-cluster centers.
result The CORE-clustering algorithm can handle large datasets efficiently.
New method finds robust clusters with statistical guarantees.
problem Clustering solutions are unstable and lack robustness guarantees.
method Quantifies cluster instability and finds robust clusters (core clusters).
result Core clusters are more stable and robust to changes in data.
This study examines cores within superclusters, highlighting their transitional nature and dynamical state.
problem Understanding the morphology and dynamical properties of cores within superclusters.
method Projected and radial velocity distributions of galaxies, morphological analysis, entropy and mass estimates.
result Cores are transitional structures that evolve towards virialisation but remain gravitationally bound.
New algorithm detects cores in graphs with community structure, improving vertex selection for better clustering.
problem Understanding and detecting core-periphery structures in graphs with community structure.
method Introduces relative centrality to detect cores in graphs with community and core-periphery structures.
result Relative centrality solves bias issues in core detection, leading to better vertex selection and improved clustering performance.
In this paper, we propose a new fuzzy clustering algorithm based on the mode-seeking framework. Given a dataset in Rd, we define regions of high density that we call cluster cores. We then consider a random walk on a neighborhood graph built on top of our data points which is designed to be attracted by hig…
New Ising models improve consensus clustering on specialized hardware.
problem Consensus clustering optimization problems.
method Formulated consensus clustering as Ising models and evaluated on specialized hardware.
result Our Ising models outperform existing techniques on consensus clustering.
A new method classifies high-dimensional images with minimal labels using diffusion geometry.
problem Classifying high-dimensional images efficiently and accurately.
method Spatially-regularized nonlinear diffusion geometry for clustering and active learning.
result High-accuracy labelings achieved with a very small number of training labels.
Quickshift++ improves clustering stability and performance.
problem Improving initial seedings for clustering algorithms.
method Provably good initial seedings for Quick Shift clustering.
result Statistical consistency and strong clustering performance.
A new algorithm reduces graph complexity for better dense subgraph analysis.
problem Mining dense subgraphs in large graphs for better analysis.
method Multi-stage graph peeling algorithm (M-PA) with two-stage data screening.
result M-PA produces similar dense subgraphs to the previous PA but with reduced graph complexity.
Parallelizes word2vec for multi-core and many-core architectures.
problem Efficiently parallelize word2vec for modern multi-core/many-core architectures.
method Proposes HogBatch, improving reuse of data structures through minibatching and negative sample sharing, allowing matrix multiply operations.
result Demonstrates strong scalability up to 32 nodes and near linear scaling across cores and nodes.
Algorithm clusters data without specifying cluster number based on similarity.
problem Clustering data without knowing the number of clusters in advance.
method Centroid-based clustering with a similarity measure to decide cluster assignment.
result Algorithm can handle streaming data and clusters based on predefined similarity level.
Two embedding methods in spectral graph clustering yield different but valid groupings.
problem Clustering vertices of a graph without true groupings.
method Spectral graph clustering using Laplacian or Adjacency spectral embedding.
result Laplacian embedding captures left hemisphere/right hemisphere structure, while adjacency embedding captures gray matter/white matter structure.
New clustering method using point-set kernel measures similarity.
problem Measuring similarity between objects for clustering.
method Point-set kernel for similarity computation; clustering procedure uses this measure.
result Proposed method is more effective and faster than existing algorithms.
This paper generalizes knot polynomials to include the HOMFLY polynomial.
problem Generalizing knot polynomials to include the HOMFLY polynomial.
method Using path posets to directly generalize the construction of the Jones and Alexander polynomials to the HOMFLY polynomial.
result The HOMFLY polynomial is obtained by specializing path posets.
Geometric framework links clustering accuracy to structural recovery.
problem Understanding the trade-off between robustness and sensitivity in clustering.
method Develops a clustering condition number to compare within-cluster scale to the minimum loss increase required to move a point across a cluster boundary.
result Sharp phase transitions for exact recovery under different objectives, providing geometric principle for interpreting low objective values.
A new distributed clustering framework using distributional kernel.
problem Clustering in distributed networks with arbitrary shapes, sizes, and densities.
method Distributed Clustering based on Distributional Kernel (KDC) using similarity of distributions.
result KDC guarantees equivalent clustering outcomes to centralized methods, reduces runtime, and discovers arbitrary clusters.
Given a similarity graph between items, correlation clustering (CC) groups similar items together and dissimilar ones apart. One of the most popular CC algorithms is KwikCluster: an algorithm that serially clusters neighborhoods of vertices, and obtains a 3-approximation ratio. Unfortunately, KwikCluster in practice re…
Simple, scalable sparse k-means for high-dimensional data.
problem Clustering in high-dimensional feature spaces with few relevant features.
method Feature ranking-based sparse k-means algorithm.
result Consistent and convergent sparse k-means clustering.
COBRAS uses super-instances to quickly cluster data with user queries.
problem Clustering data with user-defined pairwise constraints efficiently.
method Top-down construction of super-instances, iterative refinement based on user queries.
result COBRAS produces high-quality clusterings at fast run times.
New method preserves spectral clustering performance under aggressive sparsification and quantization.
problem Maintaining spectral clustering performance with sparse and quantized data.
method Random matrix theory applied to eigenspectrum changes under sparsification and quantization.
result Spectral clustering performance is preserved even with aggressive sparsification and quantization.
The paper introduces group-representative clustering to ensure fair representation of different groups in clusters.
problem Ensuring fair representation of different groups in clusters.
method Developed a new clustering approach called group-representative clustering, which parallels fairness notions in classification.
result Presented approximation algorithms for group representative k-median clustering and evaluated on real-world data. This paper proposes a new AL method that directly uses geometric sampling over clusters.
problem Performance degeneration in uncertainty evaluation for AL with insufficient labeled data.
method Divide-and-conquer approach to AL, transferring it to geometric sampling over clusters.
result The proposed GAL method significantly outperforms state-of-the-art baselines.
This paper introduces GEMINI, a new metric for unsupervised neural network training that avoids the need for regularizations.
problem The mutual information (MI) as a clustering objective does not lead to satisfactory clusters.
method The authors generalised the mutual information by changing its core distance, introducing the Generalised Mutual Information (GEMINI).
result Some GEMINIs do not require regularizations when training and can automatically select the number of clusters.
This paper introduces GEMINI, a new mutual information metric for unsupervised neural network training.
problem The mutual information (MI) as a clustering objective does not lead to satisfactory clusters.
method The authors generalised MI by changing its core distance, introducing GEMINIs that do not require regularizations and can automatically select the number of clusters.
result GEMINIs can automatically select the number of clusters without requiring a priori knowledge of the number of clusters.
FastAMI efficiently approximates AMI and SMI for large datasets.
problem Computational difficulty in comparing clusterings with an adjustment for chance.
method Monte Carlo-based approach to approximate AMI and SMI.
result FastAMI provides accurate results for large datasets.
The paper explains how regularization improves spectral clustering by reducing sensitivity to noise.
problem Spectral clustering's sensitivity to noise in sparse and stochastic graphs.
method Using graph conductance and regularization to improve spectral clustering.
result Regularization reduces sensitivity to small cuts in the graph, improving clustering accuracy and speed.
New clustering method for uncertain data using Wasserstein barycenters.
problem Clustering uncertain and structured data with observational/experimental error.
method Wasserstein barycenters and geodesic criterion for optimal clustering.
result Effective clustering of complex data in astronomy, biology, and remote sensing.
K-means fails in high dimensions with noise and few samples.
problem Clustering in high-dimensional data with noise and limited samples.
method Simple Gaussian Mixture Model (GMM) analysis.
result Almost every partition becomes a fixed point of k-means in high dimensions.
CycleCluster uses clustering to improve deep semi-supervised learning.
problem Difficulty in obtaining labelled data for deep learning.
method Proposes a new framework using clustering regularisation and graph-based pseudo-labels.
result Demonstrates improved predictive capability through numerical results.
A novel unsupervised feature learning architecture using multi-clustering integration and MIRBM.
problem Feature learning without labeled data.
method Multi-clustering integration module with MIRBM, using K-means, affinity propagation, and spectral clustering.
result The proposed architecture outperforms state-of-the-art methods in clustering tasks.
Develops an adversarial clustering algorithm for detecting cyber attacks.
problem Dealing with active adversaries in cyber security data analytics.
method Grid-based adversarial clustering algorithm using game theoretic ideas.
result Identifies normal and attack objects, sub-clusters, overlapping areas, and outliers.
LargeMvC-Net improves scalability of multi-view clustering.
problem Scalability issues in multi-view clustering.
method Deep unfolding of multi-view clustering into a network architecture with three modules.
result LargeMvC-Net consistently outperforms state-of-the-art methods in scalability and effectiveness.
This paper speeds up spectral clustering for large graphs by dilating their eigenspectrum.
problem Slow convergence in spectral clustering due to small eigengaps in graph Laplacians.
method Polynomial approximations to matrix operations that dilate the spectrum without changing eigenvectors.
result Significant acceleration of convergence in spectral clustering.
Survey classifies Clustered Federated Learning into three types of approaches.
problem Non-independent and identically distributed (non-IID) data in Federated Learning.
method Systematic review of CFL literature, principled taxonomy.
result Core CFL and Metadata-based approaches have distinct focuses.
Two clustering algorithms optimize edge controller placement in wireless networks.
problem Optimizing edge controller placement in wireless edge networks.
method Deterministic annealing based clustering algorithms ECP-LL and ECP-LB.
result The algorithms achieve better balance between synchronization and delay costs.
The Dirichlet process (DP) is a fundamental mathematical tool for Bayesian nonparametric modeling, and is widely used in tasks such as density estimation, natural language processing, and time series modeling. Although MCMC inference methods for the DP often provide a gold standard in terms asymptotic accuracy, they ca…
Proposes ICC method for dynamic portfolio optimization.
problem Non-stationarity in market conditions makes traditional portfolio optimization ineffective.
method Inverse Covariance Clustering (ICC) to identify market states and integrate into dynamic optimization.
result ICC-PO generates portfolios with higher Sharpe Ratios and greater robustness.
This paper introduces new methods to improve 1-bit matrix completion by considering cluster effects.
problem Improving 1-bit matrix completion for clustered data.
method Group-Specific 1-bit Matrix Completion (GS1MC) and Cluster Developing Matrix Completion (CDMC).
result GS1MC and CDMC outperform existing methods in synthetic and real-world data.
Trans-GLMC tackles source heterogeneity in transfer learning for structured clusters.
problem Source heterogeneity makes it hard to use multiple related auxiliary sources effectively.
method Trans-GLMC constructs clusters of sources, then combines global fusion, within-cluster refinement, and target debiasing.
result Improves facility-specific prediction and identifies interpretable communities of hospitals with mutual transferability.
New method circumvents curse of dimensionality in Laplacian estimation.
problem High-dimensional data challenges spectral clustering and diffusion maps.
method Kernelized Laplacian estimation via reproducing kernel Hilbert space.
result Non-asymptotic statistical rates show improved performance in high dimensions.
New software package for scalable DPMM inference on large datasets.
problem Scalability and practical adoption of Dirichlet Process Mixture Models.
method Efficient distributed sampling-based inference on CPUs and GPUs.
result Significant speedups and fitting of larger datasets.
New decoder improves robustness of compressive clustering.
problem Designing robust decoders for compressive clustering.
method Inspired by mean shift, proposes a new decoder.
result Significantly improves recovery of clusters from smaller sketches.
Unsupervised segmentation learns features without labels, improving accuracy.
problem Discover and localize semantically meaningful categories in images without annotations.
method Separates feature learning from cluster compactification; distills unsupervised features into discrete semantic labels using a contrastive loss function.
result Significant improvement over prior state of the art on semantic segmentation challenges.
Paper introduces new graph concepts for better modeling of temporal interactions.
problem Graph theory struggles to capture temporal and structural aspects of interactions.
method Generalizes graph concepts to handle both temporal and structural aspects of interactions.
result Formalism allows direct modeling of interactions over time, similar to graph theory.
Proposes methods to find alternative blockmodels in networks.
problem Discover secondary blockmodel representations of networks that are dissimilar to a given blockmodel.
method Incorporates non-negative matrix factorisation (NMF) with inclusion of cannot-link constraints and dissimilarity between image matrices.
result Validated the effectiveness of the proposed methods in discovering alternative blockmodels.