Unsupervised clustering of curves according to their shapes is an important problem with broad scientific applications. The existing model-based clustering techniques either rely on simple probability models (e.g., Gaussian) that are not generally valid for shape analysis or assume the number of clusters. We develop an…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New method uses cluster shapes to improve track finding in particle collisions.
Clustering partitions a dataset such that observations placed together in a group are similar but different from those in other groups. Hierarchical and -means clustering are two approaches but have different strengths and weaknesses. For instance, hierarchical clustering identifies groups in a tree-like structure b…
A mixture of Gaussians fit to a single curved or heavy-tailed cluster will report that the data contains many clusters. To produce more appropriate clusterings, we introduce a model which warps a latent mixture of Gaussians to produce nonparametric cluster shapes. The possibly low-dimensional latent mixture model allow…
A mixture of Gaussians fit to a single curved or heavy-tailed cluster will report that the data contains many clusters. To produce more appropriate clusterings, we introduce a model which warps a latent mixture of Gaussians to produce nonparametric cluster shapes. The possibly low-dimensional latent mixture model allow…
Skeleton clustering detects clusters in high-dimensional data without needing prototypes.
A hierarchical clustering algorithm for data clouds without structure assumptions.
New algorithms detect outliers in high-dimensional data with arbitrary shapes.
A novel multi-resolution cluster detection (MCD) method is proposed to identify irregularly shaped clusters in space. Multi-scale test statistic on a single cell is derived based on likelihood ratio statistic for Bernoulli sequence, Poisson sequence and Normal sequence. A neighborhood variability measure is defined to …
funLOCI identifies clusters in functional data.
Proposes a new clustering method based on expectiles for non-spherical clusters.
A Python tool generates synthetic data for cluster analysis from high-level descriptions.
A new distributed clustering framework using distributional kernel.
This paper proposes a new Nystrom-based clustering algorithm for large-scale data.
Paper proposes MMC to avoid high-density bias in clustering.
The paper presents a method for analyzing shape graphs using specific features.
We propose a deep amortized clustering (DAC), a neural architecture which learns to cluster datasets efficiently using a few forward passes. DAC implicitly learns what makes a cluster, how to group data points into clusters, and how to count the number of clusters in datasets. DAC is meta-learned using labelled dataset…
The paper analyzes the emergence of almost-honeycomb structures in low-energy planar clusters.
A new measure -variance captures local distributional shape.
We have measured the dissimilarities among several printed characters of a single page in the Gutenberg 42-line bible and we prove statistically the existence of several different matrices from which the metal types where constructed. This is in contrast with the prevailing theory, which states that only one matrix per…
PEA improves PCA and k-means for non-linear data and complex clusters.
When it comes to clustering nonconvex shapes, two paradigms are used to find the most suitable clustering: minimum cut and maximum density. The most popular algorithms incorporating these paradigms are Spectral Clustering and DBSCAN. Both paradigms have their pros and cons. While minimum cut clusterings are sensitive t…
Consumer Demand Response (DR) is an important research and industry problem, which seeks to categorize, predict and modify consumer's energy consumption. Unfortunately, traditional clustering methods have resulted in many hundreds of clusters, with a given consumer often associated with several clusters, making it diff…
We investigate an efficient context-dependent clustering technique for recommender systems based on exploration-exploitation strategies through multi-armed bandits over multiple users. Our algorithm dynamically groups users based on their observed behavioral similarity during a sequence of logged activities. In doing s…
This paper presents a novel adaptive resonance theory (ART)-based modular architecture for unsupervised learning, namely the distributed dual vigilance fuzzy ART (DDVFA). DDVFA consists of a global ART system whose nodes are local fuzzy ART modules. It is equipped with the distinctive features of distributed higher-ord…
CLASSIX is a fast and explainable clustering method that sorts data and merges groups.
NeuralFLoC unifies registration and clustering of functional data, overcoming phase variation challenges.
Statistical shape analysis can be done in a Riemannian framework by endowing the set of shapes with a Riemannian metric. Sobolev metrics of order two and higher on shape spaces of parametrized or unparametrized curves have several desirable properties not present in lower order metrics, but their discretization is stil…
This paper focuses on density-based clustering, particularly the Density Peak (DP) algorithm and the one based on density-connectivity DBSCAN; and proposes a new method which takes advantage of the individual strengths of these two methods to yield a density-based hierarchical clustering algorithm. Our investigation be…
A new vine copula mixture model improves clustering accuracy for non-Gaussian data.
Efficiently clusters data with weak assumptions, robust to contamination.
We establish the Gaussian Double-Bubble Conjecture: the least Gaussian-weighted perimeter way to decompose into three cells of prescribed (positive) Gaussian measure is to use a tripod-cluster, whose interfaces consist of three half-hyperplanes meeting along an -dimensional plane at …
CCP clusters correlated features and projects them to 1D for efficient dimensionality reduction.
Develops a new cluster validity index to find multiple optimal cluster numbers.
In the context of clustering, we consider a generative model in a Euclidean ambient space with clusters of different shapes, dimensions, sizes and densities. In an asymptotic setting where the number of points becomes large, we obtain theoretical guaranties for a few emblematic methods based on pairwise distances: a si…
We consider the problem of clustering with the longest-leg path distance (LLPD) metric, which is informative for elongated and irregularly shaped clusters. We prove finite-sample guarantees on the performance of clustering with respect to this metric when random samples are drawn from multiple intrinsically low-dimensi…
Studying the impact of climate change on precipitation is constrained by finding a way to evaluate the evolution of precipitation variability over time. Classical approaches (feature-based) have shown their limitations for this issue due to the intermittent and irregular nature of precipitation. In this study, we prese…
Bayesian nonparametric method partitions shapes using curves.
We introduce the notion of multiscale covariance tensor fields (CTF) associated with Euclidean random variables as a gateway to the shape of their distributions. Multiscale CTFs quantify variation of the data about every point in the data landscape at all spatial scales, unlike the usual covariance tensor that only qua…
The paper challenges the validity of cluster validity measures in unsupervised learning.
A new clustering method estimates non-linear boundaries and automatically selects the number of clusters.
A novel clustering method uses torque balance to group objects.
New clustering algorithm for time series data using RNN and variational Bayes.
Benchmark study evaluates 8 clustering methods on 99 UCR time series datasets.
The paper addresses uncertainties in spectral clustering of corrupted data.
Combinatorial optimization problems for clustering are known to be NP-hard. Most optimization methods are not able to find the global optimum solution for all datasets. To solve this problem, we propose a global optimal path-based clustering (GOPC) algorithm in this paper. The GOPC algorithm is based on two facts: (1) …
Recently, deep clustering, which is able to perform feature learning that favors clustering tasks via deep neural networks, has achieved remarkable performance in image clustering applications. However, the existing deep clustering algorithms generally need the number of clusters in advance, which is usually unknown in…
Spectral clustering is a fast and popular algorithm for finding clusters in networks. Recently, Chaudhuri et al. (2012) and Amini et al.(2012) proposed inspired variations on the algorithm that artificially inflate the node degrees for improved statistical performance. The current paper extends the previous statistical…