Paper finds methods to accurately determine the number of clusters in data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Despite its popularity, it is widely recognized that the investigation of some theoretical aspects of clustering has been relatively sparse. One of the main reasons for this lack of theoretical results is surely the fact that, whereas for other statistical problems the theoretical population goal is clearly defined (as…
We advocate the use of cluster algebras and their y-variables in the study of hyperbolic 3-manifolds. We study hyperbolic structures on the mapping tori of pseudo-Anosov mapping classes of punctured surfaces, and show that cluster y-variables naturally give the solutions of the edge-gluing conditions of ideal tetrahedr…
New algorithm clusters sparse data effectively.
In spectral clustering, one defines a similarity matrix for a collection of data points, transforms the matrix to get the Laplacian matrix, finds the eigenvectors of the Laplacian matrix, and obtains a partition of the data using the leading eigenvectors. The last step is sometimes referred to as rounding, where one ne…
Solves Riemann-Hilbert problems on surface triangulations.
We use Bonahon-Wong's trace map to study character varieties of the once-punctured torus and of the 4-punctured sphere. We clarify a relationship with cluster algebra associated with ideal triangulations of surfaces, and we show that the Goldman Poisson algebra of loops on surfaces is recovered from the Poisson structu…
The problem of finding groups in data (cluster analysis) has been extensively studied by researchers from the fields of Statistics and Computer Science, among others. However, despite its popularity it is widely recognized that the investigation of some theoretical aspects of clustering has been relatively sparse. One …
We formalize the arithmetic topology, i.e. a relationship between knots and primes. Namely, using the notion of a cluster C*-algebra we construct a functor from the category of 3-dimensional manifolds M to a category of algebraic number fields K, such that the prime ideals (ideals, resp.) in the ring of integers of K c…
The paper generalizes Thurston's earthquake map to cluster algebras of finite type.
Quantum trace maps for surfaces are shown to be compatible under triangulations.
In addition to finding meaningful clusters, centroid-based clustering algorithms such as K-means or mean-shift should ideally find centroids that are valid patterns in the input space, representative of data in their cluster. This is challenging with data having a nonconvex or manifold structure, as with images or text…
Dual regularized graph Laplacian improves spectral clustering for community detection.
We propose a new description of 3d theories which do not admit conventional Lagrangians. Given a quiver and a mutation sequence on it, we define a 3d theory in such a way that the partition function of the theory coincides with the cluster partition f…
Constraint-based clustering algorithms exploit background knowledge to construct clusterings that are aligned with the interests of a particular user. This background knowledge is often obtained by allowing the clustering system to pose pairwise queries to the user: should these two elements be in the same cluster or n…
PET-TURTLE improves clustering accuracy for imbalanced data.
Discriminative clustering uses mutual information to cluster data.
Density-based clustering relies on the idea of linking groups to some specific features of the probability distribution underlying the data. The reference to a true, yet unknown, population structure allows to frame the clustering problem in a standard inferential setting, where the concept of ideal population clusteri…
New groups connect braids and 3-manifolds.
A framework for forecasting high-dimensional time-series data using clustering.
Simplified image clustering achieves competitive results without text-based embeddings.
STICC clusters geographic objects considering both spatial contiguity and attributes.
Paper constructs representations for virtual braids and flat braids.
New index improves anomaly detection in correlated time series data.
Boltzmann machines are physics informed generative models with wide applications in machine learning. They can learn the probability distribution from an input dataset and generate new samples accordingly. Applying them back to physics, the Boltzmann machines are ideal recommender systems to accelerate Monte Carlo simu…
DKLM learns adaptive kernels for robust nonlinear subspace clustering.
Two novel clustering methods improve community detection in networks.
Geometric model of unbounded sl3 laminations with tropical coordinates.
This paper connects spinors to horospheres in hyperbolic space.
A new method for community detection in networks is presented.
A low-rank transformation learning framework for subspace clustering and classification is here proposed. Many high-dimensional data, such as face images and motion sequences, approximately lie in a union of low-dimensional subspaces. The corresponding subspace clustering problem has been extensively studied in the lit…
This paper aims to develop new techniques to describe joint behavior of stocks, beyond regression and correlation. For example, we want to identify the clusters of the stocks that move together. Our work is based on applying Kernel Principal Component Analysis(KPCA) and Functional Principal Component Analysis(FPCA) to …
Product Kanerva Machines dynamically combine smaller models for better memory organization.
K-Models clusters functional data with ordinal constraints for better interpretability.
In 2006, Fock and Goncharov constructed a nice basis of the ring of regular functions on the moduli space of framed -local systems on a punctured surface . The moduli space is birational to a cluster -variety, whose positive real points recover the enhanced Teichmüller space of . Their b…
This paper presents a novel adaptive resonance theory (ART)-based modular architecture for unsupervised learning, namely the distributed dual vigilance fuzzy ART (DDVFA). DDVFA consists of a global ART system whose nodes are local fuzzy ART modules. It is equipped with the distinctive features of distributed higher-ord…
Study characterizes Uniswap v3 liquidity pools using transaction graphs and identifies ideal trading conditions.
Unified HDP and LDA models for efficient topic clustering of online course queries.
Classical collaborative filtering, and content-based filtering methods try to learn a static recommendation model given training data. These approaches are far from ideal in highly dynamic recommendation domains such as news recommendation and computational advertisement, where the set of items and users is very fluid.…
Dynamics and function of neuronal networks are determined by their synaptic connectivity. Current experimental methods to analyze synaptic network structure on the cellular level, however, cover only small fractions of functional neuronal circuits, typically without a simultaneous record of neuronal spiking activity. H…
Preserving the privacy of individuals by protecting their sensitive attributes is an important consideration during microdata release. However, it is equally important to preserve the quality or utility of the data for at least some targeted workloads. We propose a novel framework for privacy preservation based on the …
Paper explores supervised learning methods to approximate ideal observer for joint signal detection and localization.
The study proves poor ideal three-edge triangulations are minimal for certain 3-manifolds.
We give a simple method to find ideal points of the character variety of a 3-manifold from an ideal triangulation.
We address the problem of acoustic source separation in a deep learning framework we call "deep clustering." Rather than directly estimating signals or masking functions, we train a deep network to produce spectrogram embeddings that are discriminative for partition labels given in training data. Previous deep network …
Anomaly detection is challenging, especially for large datasets in high dimensions. Here we explore a general anomaly detection framework based on dimensionality reduction and unsupervised clustering. We release DRAMA, a general python package that implements the general framework with a wide range of built-in options.…
Paper provides new Alexander ideal-based obstruction to 0-concordance of knotted surfaces.
Short proof for ideal polygons with near optimal orthogeodesic decomposition.