We establish basic properties of cluster algebras associated with oriented bordered surfaces with marked points. In particular, we show that the underlying cluster complex of such a cluster algebra does not depend on the choice of coefficients, describe this complex explicitly in terms of "tagged triangulations" of the…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We try to give a cluster algebraic interpretation of complex volume of knots. We construct the R-operator from the cluster mutations, and we show that it is regarded as a hyperbolic octahedron. The cluster variables are interpreted as edge parameters used by Zickert in computing complex volume.
A clustering algorithm for natural hierarchical clusters with near-linear time complexity.
Proposes a constraint for deep clustering to handle both simple and complex topologies.
A fundamental property of complex networks is the tendency for edges to cluster. The extent of the clustering is typically quantified by the clustering coefficient, which is the probability that a length-2 path is closed, i.e., induces a triangle in the network. However, higher-order cliques beyond triangles are crucia…
We outline a novel clustering scheme for simplicial complexes that produces clusters of simplices in a way that is sensitive to the homology of the complex. The method is inspired by, and can be seen as a higher-dimensional version of, graph spectral clustering. The algorithm involves only sparse eigenproblems, and is …
New concept of mixture complexity helps detect gradual clustering changes.
This paper proposes a spectral clustering algorithm for hyperbolic spaces, improving efficiency over Euclidean methods.
We propose a framework for Semi-Supervised Active Clustering framework (SSAC), where the learner is allowed to interact with a domain expert, asking whether two given instances belong to the same cluster or not. We study the query and computational complexity of clustering in this framework. We consider a setting where…
SPINEX improves clustering with explainable neighbors, outperforming other methods.
DIVA clusters dynamic data without needing cluster count, outperforming baselines.
New method reduces clustering time and improves accuracy.
Clustered attention improves transformer efficiency for large sequences.
Study clusters distributions with known or unknown clusters using distribution testing.
Spectral clustering is a widely studied problem, yet its complexity is prohibitive for dynamic graphs of even modest size. We claim that it is possible to reuse information of past cluster assignments to expedite computation. Our approach builds on a recent idea of sidestepping the main bottleneck of spectral clusterin…
Study examines how cluster number affects short-text clustering, introducing a stability metric.
FCA improves fair clustering by optimizing utility and fairness.
This study evaluates clustering algorithms on high-dimensional data.
CLASSIX is a fast and explainable clustering method that sorts data and merges groups.
New interpretation reconciles country and product complexity.
Kernel methods obtain superb performance in terms of accuracy for various machine learning tasks since they can effectively extract nonlinear relations. However, their time complexity can be rather large especially for clustering tasks. In this paper we define a general class of kernels that can be easily approximated …
A new hierarchical clustering method selects representative points from sub-minimum-spanning-trees.
Study exact partition recovery with same-cluster oracle, bounded error.
A new clustering algorithm reduces density peaks clustering's computational complexity.
A Bernoulli Mixture Model (BMM) is a finite mixture of random binary vectors with independent dimensions. The problem of clustering BMM data arises in a variety of real-world applications, ranging from population genetics to activity analysis in social networks. In this paper, we analyze the clusterability of BMMs from…
SASE improves attributed graph clustering for large graphs with linear time and space complexity.
Paper detects gradual changes in cluster structure using MC fusion.
This article explores and analyzes the unsupervised clustering of large partially observed graphs. We propose a scalable and provable randomized framework for clustering graphs generated from the stochastic block model. The clustering is first applied to a sub-matrix of the graph's adjacency matrix associated with a re…
This paper represents a preliminary (pre-reviewing) version of a sublinear variational algorithm for isotropic Gaussian mixture models (GMMs). Further developments of the algorithm for GMMs with diagonal covariance matrices (instead of isotropic clusters) and their corresponding benchmarking results have been published…
New SDP algorithm recovers large clusters in SBM with small clusters of any size.
Exact cluster recovery with same-cluster queries for arbitrary ellipsoidal clusters.
This paper provides new algorithms for distributed clustering for two popular center-based objectives, k-median and k-means. These algorithms have provable guarantees and improve communication complexity over existing approaches. Following a classic approach in clustering by \cite{har2004coresets}, we reduce the proble…
Discrete random variables are natural components of probabilistic clustering models. A number of VAE variants with discrete latent variables have been developed. Training such methods requires marginalizing over the discrete latent variables, causing training time complexity to be linear in the number clusters. By appl…
We propose a method to compute complex volume of 2-bridge link complements. Our construction sheds light on a relationship between cluster variables with coefficients and canonical decompositions of link complements.
New algorithm improves plant breeding by clustering soybean genotypes more accurately and efficiently.
Proposes new random models for fuzzy clustering similarity measures.
We propose a novel method to quantify the clustering behavior in a complex time series and apply it to a high-frequency data of the financial markets. We find that regardless of used data sets, all data exhibits the volatility clustering properties, whereas those which filtered the volatility clustering effect by using…
Proposes ConiVAT for better cluster assessment and clustering with background knowledge.
Suppose, we are given a set of elements to be clustered into (unknown) clusters, and an oracle/expert labeler that can interactively answer pair-wise queries of the form, "do two elements and belong to the same cluster?". The goal is to recover the optimum clustering by asking the minimum number of quer…
Clustering ensemble, or consensus clustering, has emerged as a powerful tool for improving both the robustness and the stability of results from individual clustering methods. Weighted clustering ensemble arises naturally from clustering ensemble. One of the arguments for weighted clustering ensemble is that elements (…
New matrix reveals cluster info in sparse directed graphs.
PEA improves PCA and k-means for non-linear data and complex clusters.
In this paper we target the class of modal clustering methods where clusters are defined in terms of the local modes of the probability density function which generates the data. The most well-known modal clustering method is the k-means clustering. Mean Shift clustering is a generalization of the k-means clustering wh…
Graph clustering improved using Boltzmann machine heuristics.
One iteration of standard -means (i.e., Lloyd's algorithm) or standard EM for Gaussian mixture models (GMMs) scales linearly with the number of clusters , data points , and data dimensionality . In this study, we explore whether one iteration of -means or EM for GMMs can scale sublinearly with at run…
SNG-DBSCAN clusters data faster with subsampled similarity queries.
We propose infinite mixture prototypes to adaptively represent both simple and complex data distributions for few-shot learning. Our infinite mixture prototypes represent each class by a set of clusters, unlike existing prototypical methods that represent each class by a single cluster. By inferring the number of clust…
A scalable Gaussian process clustering method for large datasets.