KCoreMotif clusters large networks efficiently by exploiting k-core decomposition and motifs.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Empty core found in max-loss non-centroid clustering.
Density mode clustering is a nonparametric clustering method. The clusters are the basins of attraction of the modes of a density estimator. We study the risk of mode-based clustering. We show that the clustering risk over the cluster cores --- the regions where the density is high --- is very small even in high dimens…
In this paper, we present a new R package COREclust dedicated to the detection of representative variables in high dimensional spaces with a potentially limited number of observations. Variable sets detection is based on an original graph clustering strategy denoted CORE-clustering algorithm that detects CORE-clusters,…
Clustering is a widely used unsupervised learning method for finding structure in the data. However, the resulting clusters are typically presented without any guarantees on their robustness; slightly changing the used data sample or re-running a clustering algorithm involving some stochastic component may lead to comp…
This study examines cores within superclusters, highlighting their transitional nature and dynamical state.
New algorithm detects cores in graphs with community structure, improving vertex selection for better clustering.
In this paper, we propose a new fuzzy clustering algorithm based on the mode-seeking framework. Given a dataset in , we define regions of high density that we call cluster cores. We then consider a random walk on a neighborhood graph built on top of our data points which is designed to be attracted by hig…
New Ising models improve consensus clustering on specialized hardware.
A new method classifies high-dimensional images with minimal labels using diffusion geometry.
A new algorithm reduces graph complexity for better dense subgraph analysis.
We provide initial seedings to the Quick Shift clustering algorithm, which approximate the locally high-density regions of the data. Such seedings act as more stable and expressive cluster-cores than the singleton modes found by Quick Shift. We establish statistical consistency guarantees for this modification. We then…
This paper proposes a centroid-based clustering algorithm which is capable of clustering data-points with n-features, without having to specify the number of clusters to be formed. The core logic behind the algorithm is a similarity measure, which collectively decides whether to assign an incoming data-point to a pre-e…
Clustering is concerned with coherently grouping observations without any explicit concept of true groupings. Spectral graph clustering - clustering the vertices of a graph based on their spectral embedding - is commonly approached via K-means (or, more generally, Gaussian mixture model) clustering composed with either…
Word2vec is a widely used algorithm for extracting low-dimensional vector representations of words. State-of-the-art algorithms including those by Mikolov et al. have been parallelized for multi-core CPU architectures, but are based on vector-vector operations with "Hogwild" updates that are memory-bandwidth intensive …
New clustering method using point-set kernel measures similarity.
Recently, it has been shown that the Jones polynomial, in [LS19], and the Alexander polynomial, in [NT18], of rational knots can be obtained by specializing -polynomials of cluster variables. At the core of both results are continued fractions, which parameterize rational knots and are used to obtain cluster variabl…
Geometric framework links clustering accuracy to structural recovery.
A new distributed clustering framework using distributional kernel.
Given a similarity graph between items, correlation clustering (CC) groups similar items together and dissimilar ones apart. One of the most popular CC algorithms is KwikCluster: an algorithm that serially clusters neighborhoods of vertices, and obtains a 3-approximation ratio. Unfortunately, KwikCluster in practice re…
Constraint-based clustering algorithms exploit background knowledge to construct clusterings that are aligned with the interests of a particular user. This background knowledge is often obtained by allowing the clustering system to pose pairwise queries to the user: should these two elements be in the same cluster or n…
Simple, scalable sparse k-means for high-dimensional data.
New method preserves spectral clustering performance under aggressive sparsification and quantization.
The paper introduces group-representative clustering to ensure fair representation of different groups in clusters.
This paper introduces GEMINI, a new metric for unsupervised neural network training that avoids the need for regularizations.
Active learning (AL) repeatedly trains the classifier with the minimum labeling budget to improve the current classification model. The training process is usually supervised by an uncertainty evaluation strategy. However, the uncertainty evaluation always suffers from performance degeneration when the initial labeled …
This paper introduces GEMINI, a new mutual information metric for unsupervised neural network training.
This paper uses the relationship between graph conductance and spectral clustering to study (i) the failures of spectral clustering and (ii) the benefits of regularization. The explanation is simple. Sparse and stochastic graphs create a lot of small trees that are connected to the core of the graph by only one edge. G…
FastAMI efficiently approximates AMI and SMI for large datasets.
New clustering method for uncertain data using Wasserstein barycenters.
K-means fails in high dimensions with noise and few samples.
CycleCluster uses clustering to improve deep semi-supervised learning.
In this paper, we present a novel unsupervised feature learning architecture, which consists of a multi-clustering integration module and a variant of RBM termed multi-clustering integration RBM (MIRBM). In the multi-clustering integration module, we apply three unsupervised K-means, affinity propagation and spectral c…
LargeMvC-Net improves scalability of multi-view clustering.
This paper speeds up spectral clustering for large graphs by dilating their eigenspectrum.
Survey classifies Clustered Federated Learning into three types of approaches.
Two clustering algorithms optimize edge controller placement in wireless networks.
Nowadays more and more data are gathered for detecting and preventing cyber attacks. In cyber security applications, data analytics techniques have to deal with active adversaries that try to deceive the data analytics models and avoid being detected. The existence of such adversarial behavior motivates the development…
The Dirichlet process (DP) is a fundamental mathematical tool for Bayesian nonparametric modeling, and is widely used in tasks such as density estimation, natural language processing, and time series modeling. Although MCMC inference methods for the DP often provide a gold standard in terms asymptotic accuracy, they ca…
Proposes ICC method for dynamic portfolio optimization.
Matrix completion has a long-time history of usage as the core technique of recommender systems. In particular, 1-bit matrix completion, which considers the prediction as a ``Recommended'' or ``Not Recommended'' question, has proved its significance and validity in the field. However, while customers and products aggre…
Trans-GLMC tackles source heterogeneity in transfer learning for structured clusters.
New method circumvents curse of dimensionality in Laplacian estimation.
We use methods from network science to analyze corruption risk in a large administrative dataset of over 4 million public procurement contracts from European Union member states covering the years 2008-2016. By mapping procurement markets as bipartite networks of issuers and winners of contracts we can visualize and de…
New software package for scalable DPMM inference on large datasets.
Many approaches have been proposed to discover clusters within networks. Community finding field encompasses approaches which try to discover clusters where nodes are tightly related within them but loosely related with nodes of other clusters. However, a community network configuration is not the only possible latent …
New decoder improves robustness of compressive clustering.
Unsupervised segmentation learns features without labels, improving accuracy.