This work studies clustering in transformer models, proving exponential convergence to a single token state.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Review of clustering methods for functional data across various fields.
A new MFG framework for evolving clusters from Gaussian mixtures.
STICC clusters geographic objects considering both spatial contiguity and attributes.
Generates infinite-depth hierarchical clusters from few examples.
The mean field methods, which entail approximating intractable probability distributions variationally with distributions from a tractable family, enjoy high efficiency, guaranteed convergence, and provide lower bounds on the true likelihood. But due to requirement for model-specific derivation of the optimization equa…
A new method uses Mean Field Games to optimize mixture models of Bernoulli and categorical distributions.
Clusters of withdrawals emerge in banks due to latent fragility.
This study addresses the issue of predicting the glaucomatous visual field loss from patient disease datasets. Our goal is to accurately predict the progress of the disease in individual patients. As very few measurements are available for each patient, it is difficult to produce good predictors for individuals. A rece…
Nowadays, data are generated massively and rapidly from scientific fields as bioinformatics, neuroscience and astronomy to business and engineering fields. Cluster analysis, as one of the major data analysis tools, is therefore more significant than ever. We propose in this work an effective Semi-supervised Divisive Cl…
Recently, variational approximations such as the mean field approximation have received much interest. We extend the standard mean field method by using an approximating distribution that factorises into cluster potentials. This includes undirected graphs, directed acyclic graphs and junction trees. We derive generaliz…
The paper develops a Galois theory for cluster algebras and Riemann surfaces.
Deep learning improves community detection in graph datasets.
Saec compresses recommendation system embeddings by clustering similar features.
Quantum theory improves counting overlapping clusters.
Survey of robust clustering methods for hotspot detection.
A new clustering method using Bayesian techniques improves robustness and interpretability.
A new clustering algorithm fuses heat diffusion and turning angle for robustness.
Mixed data comprises both numeric and categorical features, and mixed datasets occur frequently in many domains, such as health, finance, and marketing. Clustering is often applied to mixed datasets to find structures and to group similar objects for further analysis. However, clustering mixed data is challenging becau…
Model-based clustering is a popular approach for clustering multivariate data which has seen applications in numerous fields. Nowadays, high-dimensional data are more and more common and the model-based clustering approach has adapted to deal with the increasing dimensionality. In particular, the development of variabl…
SurvMixClust clusters survival data and predicts individual survival curves.
We present a large catalog of optically selected galaxy clusters from the application of a new Gaussian Mixture Brightest Cluster Galaxy (GMBCG) algorithm to SDSS Data Release 7 data. The algorithm detects clusters by identifying the red sequence plus Brightest Cluster Galaxy (BCG) feature, which is unique for galaxy c…
The problem of finding groups in data (cluster analysis) has been extensively studied by researchers from the fields of Statistics and Computer Science, among others. However, despite its popularity it is widely recognized that the investigation of some theoretical aspects of clustering has been relatively sparse. One …
Sparse elasticity reconstruction from local displacements reduces error.
An autonomous variational inference algorithm for arbitrary graphical models requires the ability to optimize variational approximations over the space of model parameters as well as over the choice of tractable families used for the variational approximation. In this paper, we present a novel combination of graph part…
The area of constrained clustering has been extensively explored by researchers and used by practitioners. Constrained clustering formulations exist for popular algorithms such as k-means, mixture models, and spectral clustering but have several limitations. A fundamental strength of deep learning is its flexibility, a…
There has been a surge in the number of large and flat data sets - data sets containing a large number of features and a relatively small number of observations - due to the growing ability to collect and store information in medical research and other fields. Hierarchical clustering is a widely used clustering tool. I…
This paper analyzes various graph clustering methods and their applications.
Clustering is ubiquitous in data analysis, including analysis of time series. It is inherently subjective: different users may prefer different clusterings for a particular dataset. Semi-supervised clustering addresses this by allowing the user to provide examples of instances that should (not) be in the same cluster. …
New method clusters matrix-variate data with outliers.
A novel clustering algorithm inspired by atomic fission.
New algorithm clusters multiple samples from hidden distributions.
Clustering is one of the most common unsupervised learning tasks in machine learning and data mining. Clustering algorithms have been used in a plethora of applications across several scientific fields. However, there has been limited research in the clustering of point patterns - sets or multi-sets of unordered elemen…
We propose a novel method for multiple clustering that assumes a co-clustering structure (partitions in both rows and columns of the data matrix) in each view. The new method is applicable to high-dimensional data. It is based on a nonparametric Bayesian approach in which the number of views and the number of feature-/…
Study recovers tree structure in noisy MRFs with support size 3 or more.
Scientists in many fields have the common and basic need of dimensionality reduction: visualizing the underlying structure of the massive multivariate data in a low-dimensional space. However, many dimensionality reduction methods confront the so-called "crowding problem" that clusters tend to overlap with each other i…
A deep clustering method for hyperspectral images improves clustering performance by constraining intra-class distances.
We present *K-means clustering algorithm and source code by expanding statistical clustering methods applied in https://ssrn.com/abstract=2802753 to quantitative finance. *K-means is statistically deterministic without specifying initial centers, etc. We apply *K-means to extracting cancer signatures from genome data w…
New algorithms improve spectral clustering for finite mixture models.
New method explains cluster assignments in neural networks.
FedCBO solves clustered federated learning by optimizing groups of users without knowing their structure.
DGC clusters data with side-information for better prediction.
New deep clustering network uses divergence measures for unlabeled data.
A new method estimates the number of clusters on spherical data.
New clustering algorithm for time series data using RNN and variational Bayes.
ALPCAHUS clusters data from multiple subspaces with varying noise.
Unified framework improves deep multi-view clustering by addressing self-supervision and contrastive alignment issues.
Lumbermark clusters data robustly, slicing limbs of mutual reachability trees.