Proposes a new model for unsupervised clustering with latent variables.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In this paper, we propose a general model for plane-based clustering. The general model contains many existing plane-based clustering methods, e.g., k-plane clustering (kPC), proximal plane clustering (PPC), twin support vector clustering (TWSVC) and its extensions. Under this general model, one may obtain an appropria…
This study evaluates cluster search algorithms using Gaussian mixture models.
Integrates VAEs into EM for deep clustering and generation.
ARMED models improve deep learning interpretability and generalize better on clustered data.
The paper introduces the concept of a cluster structure to define a joint distribution of the sample size and its exchangeable random partitions. The cluster structure allows the probability distribution of the random partitions of a subset of the sample to be dependent on the sample size, a feature not presented in a …
Adaptive graph auto-encoder improves general data clustering.
Meta-learning model clusters data better than standard methods.
K-ARMA models cluster time series data robustly.
The learning of mixture models can be viewed as a clustering problem. Indeed, given data samples independently generated from a mixture of distributions, we often would like to find the {\it correct target clustering} of the samples according to which component distribution they were generated from. For a clustering pr…
A new method clusters survival data using deep variational models.
We describe a probabilistic (generative) view of affinity matrices along with inference algorithms for a subclass of problems associated with data clustering. This probabilistic view is helpful in understanding different models and algorithms that are based on affinity functions OF the data. IN particular, we show how(…
Paper presents robust clustering methods for general mixture models.
clusterBMA combines clustering results from multiple models using Bayesian model averaging.
A new framework for predictive clustering and optimization.
Graph-based clustering methods have demonstrated the effectiveness in various applications. Generally, existing graph-based clustering methods first construct a graph to represent the input data and then partition it to generate the clustering result. However, such a stepwise manner may make the constructed graph not f…
Graph clustering involves the task of dividing nodes into clusters, so that the edge density is higher within clusters as opposed to across clusters. A natural, classic and popular statistical setting for evaluating solutions to this problem is the stochastic block model, also referred to as the planted partition model…
New deep learning framework for tabular data clusters with interpretable features.
Model-based clustering defines population level clusters relative to a model that embeds notions of similarity. Algorithms tailored to such models yield estimated clusters with a clear statistical interpretation. We take this view here and introduce the class of G-block covariance models as a background model for varia…
End-to-end clustering without labels using normalized cuts.
GMMSEQ clusters AE data streams, identifying cluster onsets and growth.
CURE extracts relations without supervision by clustering similar entity pairs.
The paper proposes a method to select clusters, models, and algorithms based on quadratic discriminant scores.
The paper analyzes how clustering sensitive data can improve model generalization without revealing individual information.
In the context of clustering, we assume a generative model where each cluster is the result of sampling points in the neighborhood of an embedded smooth surface; the sample may be contaminated with outliers, which are modeled as points sampled in space away from the clusters. We consider a prototype for a higher-order …
A new clustering model for mixed datasets combines continuous and non-continuous data.
Most generative models for clustering implicitly assume that the number of data points in each cluster grows linearly with the total number of data points. Finite mixture models, Dirichlet process mixture models, and Pitman--Yor process mixture models make this assumption, as do all other infinitely exchangeable cluste…
Posterior regularization enhances Bayesian hierarchical mixture clustering by improving node separation.
We present an approach to model-based hierarchical clustering by formulating an objective function based on a Bayesian analysis. This model organizes the data into a cluster hierarchy while specifying a complex feature-set partitioning that is a key component of our model. Features can have either a unique distribution…
MMGAN stabilizes GANs for multi-modal data clustering.
The information bottleneck (IB) approach to clustering takes a joint distribution and maps the data to cluster labels which retain maximal information about (Tishby et al., 1999). This objective results in an algorithm that clusters data points based upon the similarity of their condit…
Deep generative models trained with large amounts of unlabelled data have proven to be powerful within the domain of unsupervised learning. Many real life data sets contain a small amount of labelled data points, that are typically disregarded when training generative models. We propose the Cluster-aware Generative Mod…
A new framework for clustering with uncertainty quantification.
Model selection in clustering requires (i) to specify a suitable clustering principle and (ii) to control the model order complexity by choosing an appropriate number of clusters depending on the noise level in the data. We advocate an information theoretic perspective where the uncertainty in the measurements quantize…
Generative model for hypergraph clustering improves detection of higher-order structure.
iCVI-ARTMAP accelerates clustering with adaptive resonance theory and validity indices.
We develop a model to cluster time-series data with interval censoring, improving disease phenotyping.
Algorithm identifies intended fairness constraints from expert demonstrations for fair clustering.
In this paper we formulate in general terms an approach to prove strong consistency of the Empirical Risk Minimisation inductive principle applied to the prototype or distance based clustering. This approach was motivated by the Divisive Information-Theoretic Feature Clustering model in probabilistic space with Kullbac…
Most generative models for clustering implicitly assume that the number of data points in each cluster grows linearly with the total number of data points. Finite mixture models, Dirichlet process mixture models, and Pitman--Yor process mixture models make this assumption, as do all other infinitely exchangeable cluste…
Semi-supervised clustering is the task of clustering data points into clusters where only a fraction of the points are labelled. The true number of clusters in the data is often unknown and most models require this parameter as an input. Dirichlet process mixture models are appealing as they can infer the number of clu…
The study compares clustering risk in Hidden Markov and i.i.d. models, showing the Bayes classifier is nearly optimal.
MC-GMENN improves neural networks for clustered data using Monte Carlo methods.
A hybrid method clusters and characterizes cancer data efficiently.
New spectral clustering method for multi-layer networks improves accuracy.
Randomized spectral co-clustering speeds up large-scale directed networks.
EGMM improves clustering by better handling uncertainty with evidential framework.
Study provides guarantees for kernel clustering under non-parametric mixtures.