Efficient clustering in high dimensions with Quick Shift and LSH.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Clustering stocks reduces estimation error in global minimum variance portfolio.
Parameter-free clustering method using cluster catch digraphs (CCDs).
Entropy regularization improves interpretability of probabilistic clustering models.
A new clustering method uses nonparametric smoothing to estimate cluster membership functions.
In many practical applications of clustering, the objects to be clustered evolve over time, and a clustering result is desired at each time step. In such applications, evolutionary clustering typically outperforms traditional static clustering by producing clustering results that reflect long-term trends while being ro…
A major challenge in cluster analysis is that the number of data clusters is mostly unknown and it must be estimated prior to clustering the observed data. In real-world applications, the observed data is often subject to heavy tailed noise and outliers which obscure the true underlying structure of the data. Consequen…
JojoSCL improves scRNA-seq clustering by reducing intra-cluster dispersion.
We study the connections between spectral clustering and the problems of maximum margin clustering, and estimation of the components of level sets of a density function. Specifically, we obtain bounds on the eigenvectors of graph Laplacian matrices in terms of the between cluster separation, and within cluster connecti…
A new method estimates the number of clusters on spherical data.
Mean shift clustering finds the modes of the data probability density by identifying the zero points of the density gradient. Since it does not require to fix the number of clusters in advance, the mean shift has been a popular clustering algorithm in various application fields. A typical implementation of the mean shi…
Most density-based clustering methods largely rely on how well the underlying density is estimated. However, density estimation itself is also a challenging problem, especially the determination of the kernel bandwidth. A large bandwidth could lead to the over-smoothed density estimation in which the number of density …
A new clustering evaluation index based on density estimation.
New method for clustering tasks with heterogeneous data.
Two new methods improve clustering with missing data.
This paper improves robust cluster enumeration for RES data.
One challenge impeding the analysis of terabyte scale x-ray scattering data from the Linac Coherent Light Source LCLS, is determining the number of clusters required for the execution of traditional clustering algorithms. Here we demonstrate that previous work using bi-cross validation (BCV) to determine the number of …
A unified clustering approach that can estimate number of clusters and produce clustering against this number simultaneously is proposed. Average silhouette width (ASW) is a widely used standard cluster quality index. A distance based objective function that optimizes ASW for clustering is defined. The proposed algorit…
For a density on , a {\it high-density cluster} is any connected component of , for some . The set of all high-density clusters forms a hierarchy called the {\it cluster tree} of . We present two procedures for estimating the cluster tree given samples from . The first…
New estimator uses clustering to improve off-policy evaluation accuracy.
CDL index improves clustering validation for non-convex data.
Motivated by modern applications in which one constructs graphical models based on a very large number of features, this paper introduces a new class of cluster-based graphical models, in which variable clustering is applied as an initial step for reducing the dimension of the feature space. We employ model assisted cl…
We derive and analyze a generic, recursive algorithm for estimating all splits in a finite cluster tree as well as the corresponding clusters. We further investigate statistical properties of this generic clustering algorithm when it receives level set estimates from a kernel density estimator. In particular, we derive…
New RESK distributions improve robust clustering of skewed data.
Proposes a robust clustering method using the Median-of-Means estimator.
Gradient Boosted Mixed Models estimate mean and variance components for clustered data.
In addition to finding meaningful clusters, centroid-based clustering algorithms such as K-means or mean-shift should ideally find centroids that are valid patterns in the input space, representative of data in their cluster. This is challenging with data having a nonconvex or manifold structure, as with images or text…
The paper proposes a new method for density estimation using spline quasi-interpolation for clustering.
Parallel neural networks estimate TVD for merging over-clustered datasets.
One of the applications of center-based clustering algorithms such as K-Means is partitioning data points into K clusters. In some examples, the feature space relates to the underlying problem we are trying to solve, and sometimes we can obtain a suitable feature space. Nevertheless, while K-Means is one of the most ef…
Cluster analysis is widely used in the areas of machine learning and data mining. Fuzzy clustering is a particular method that considers that a data point can belong to more than one cluster. Fuzzy clustering helps obtain flexible clusters, as needed in such applications as text categorization. The performance of a clu…
Cluster analysis is a fundamental tool for pattern discovery of complex heterogeneous data. Prevalent clustering methods mainly focus on vector or matrix-variate data and are not applicable to general-order tensors, which arise frequently in modern scientific and business applications. Moreover, there is a gap between …
Method estimates number of clusters in Block Markov Chain trajectories.
Proposes a weighted conformal approach for cluster label uncertainty.
Discussing issues in robust clustering, especially with Gaussian models.
MTLRRC improves MTL by robustly clustering tasks and detecting outliers.
High density clusters can be characterized by the connected components of a level set of the underlying probability density function generating the data, at some appropriate level . The complete hierarchical clustering can be characterized by a cluster tree ${\cal T}= \bigcup_λ L(λ)…
t-NEB clusters high-dimensional data hierarchically with density paths.
Develops a new random forest method for clustered data with improved prediction and inference.
The paper addresses Qini curve estimation under clustered network interference.
Commonly-used clustering algorithms usually find ellipsoidal, spherical or other regular-structured clusters, but are more challenged when the underlying groups lack formal structure or definition. Syncytial clustering is the name that we introduce for methods that merge groups obtained from standard clustering algorit…
A new framework for clustering with uncertainty quantification.
A framework for forecasting high-dimensional time-series data using clustering.
Mode clustering is a nonparametric method for clustering that defines clusters using the basins of attraction of a density estimator's modes. We provide several enhancements to mode clustering: (i) a soft variant of cluster assignment, (ii) a measure of connectivity between clusters, (iii) a technique for choosing the …
Bayesian SAE model with spectral clustering and uncertainty quantification.
Consistent estimator for mixtures of nonparametric elliptical distributions helps cluster analysis.
Producing overlapping schemes is a major issue in clustering. Recent proposed overlapping methods relies on the search of an optimal covering and are based on different metrics, such as Euclidean distance and I-Divergence, used to measure closeness between observations. In this paper, we propose the use of another meas…
Proposes a novel method to cluster individuals based on treatment effects.