We define transit clusters to simplify causal diagrams and preserve their essential properties.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Framework for multi-scale clustering using phase transitions.
We present a global optimization algorithm for clustering data given the ratio of likelihoods that each pair of data points is in the same cluster or in different clusters. To define a clustering solution in terms of pairwise relationships, a necessary and sufficient condition is that belonging to the same cluster sati…
Multilayer graphs are commonly used for representing different relations between entities and handling heterogeneous data processing tasks. Non-standard multilayer graph clustering methods are needed for assigning clusters to a common multilayer node set and for combining information from each layer. This paper present…
The interaction between transitivity and sparsity, two common features in empirical networks, implies that there are local regions of large sparse networks that are dense. We call this the blessing of transitivity and it has consequences for both modeling and inference. Extant research suggests that statistical inferen…
There have been several spectral bounds for the percolation transition in networks, using spectrum of matrices associated with the network such as the adjacency matrix and the non-backtracking matrix. However they are far from being tight when the network is sparse and displays clustering or transitivity, which is repr…
Multilayer graphs are commonly used for representing different relations between entities and handling heterogeneous data processing tasks. New challenges arise in multilayer graph clustering for assigning clusters to a common multilayer node set and for combining information from each layer. This paper presents a theo…
One of the longstanding open problems in spectral graph clustering (SGC) is the so-called model order selection problem: automated selection of the correct number of clusters. This is equivalent to the problem of finding the number of connected components or communities in an undirected graph. We propose automated mode…
Paper detects gradual changes in cluster structure using MC fusion.
Geometric framework links clustering accuracy to structural recovery.
Proposes ConiVAT for better cluster assessment and clustering with background knowledge.
SC-InfoNCE improves InfoNCE for feature clustering in contrastive learning.
Solves complex clustering and rotation synchronization problem.
Develops a new tensor model for clustering with degree correction.
Method estimates number of clusters in Block Markov Chain trajectories.
New method clusters directed and undirected graphs without losing directional information.
A framework clusters vehicle motion trajectories efficiently.
New method makes machine learning approximations unbiased and efficient.
Clusters of financial market states identified over 2006-2019.
New insights show stochastic initialization prevents token clustering in deep Transformers.
The paper reveals a spinning top geometry in real-world games.
Paper proposes an alternative to anchor points for learning with noisy labels.
Adaptive orthogonalization of data for clustering and visualization.
Unsupervised clustering of series using dynamic programming.
We generalise surface cluster algebras to the case of infinite surfaces where the surface contains finitely many accumulation points of boundary marked points. To connect different triangulations of an infinite surface, we consider infinite mutation sequences. We show transitivity of infinite mutation sequences on tria…
This paper studies how to find compact state embeddings from high-dimensional Markov state trajectories, where the transition kernel has a small intrinsic rank. In the spirit of diffusion map, we propose an efficient method for learning a low-dimensional state embedding and capturing the process's dynamics. This idea a…
We use a cluster ensemble to determine the number of clusters, k, in a group of data. A consensus similarity matrix is formed from the ensemble using multiple algorithms and several values for k. A random walk is induced on the graph defined by the consensus matrix and the eigenvalues of the associated transition proba…
Consider a two-class clustering problem where we observe , , . The feature vector is unknown but is presumably sparse. The class labels are also unknown and the main interest is to estimate them. We are interested …
Paper proposes a tensor model for clustering noisy multi-view data.
Unsupervised learning is a discipline of machine learning which aims at discovering patterns in big data sets or classifying the data into several categories without being trained explicitly. We show that unsupervised learning techniques can be readily used to identify phases and phases transitions of many body systems…
Characterizes optimal reconstruction error in high-dimensional Gaussian mixtures.
Clustering is inherently ill-posed: there often exist multiple valid clusterings of a single dataset, and without any additional information a clustering system has no way of knowing which clustering it should produce. This motivates the use of constraints in clustering, as they allow users to communicate their interes…
In this paper, we study the sensitivity of the spectral clustering based community detection algorithm subject to a Erdos-Renyi type random noise model. We prove phase transitions in community detectability as a function of the external edge connection probability and the noisy edge presence probability under a general…
The article detects market regimes from covariance matrices using VLSTAR and clustering models.
In this paper, we perform statistical segmentation and clustering analysis of the Dow Jones Industrial Average time series between January 1997 and August 2008. Modeling the index movements and log-index movements as stationary Gaussian processes, we find a total of 116 and 119 statistically stationary segments respect…
Diffusion maps help learn complex quantum phase transitions from data.
This work improved clustering methods by analyzing various datasets and dendrograms.
We consider the problem of Gaussian mixture clustering in the high-dimensional limit where the data consists of points in dimensions, and stays finite. Using exact but non-rigorous methods from statistical physics, we determine the critical value of and the distance between…
New algorithm clusters trajectories from multiple Markov chains with near-optimal error.
We analyze the time series of four major cryptocurrencies (Bitcoin, Ethereum, Litecoin, and Ripple) before the digital market crash at the end of 2017 - beginning 2018. We introduce a methodology that combines topological data analysis with a machine learning technique -- -means clustering -- in order to automatical…
Machine learning predicts critical points for directed percolation models.
This article explores and analyzes the unsupervised clustering of large partially observed graphs. We propose a scalable and provable randomized framework for clustering graphs generated from the stochastic block model. The clustering is first applied to a sub-matrix of the graph's adjacency matrix associated with a re…
Study quantifies performance gap between tensor and matrix-based approaches in nested matrix-tensor model.
As a model problem for clustering, we consider the densest k-disjoint-clique problem of partitioning a weighted complete graph into k disjoint subgraphs such that the sum of the densities of these subgraphs is maximized. We establish that such subgraphs can be recovered from the solution of a particular semidefinite re…
Nowadays, data are generated massively and rapidly from scientific fields as bioinformatics, neuroscience and astronomy to business and engineering fields. Cluster analysis, as one of the major data analysis tools, is therefore more significant than ever. We propose in this work an effective Semi-supervised Divisive Cl…
Linear dynamical systems are a fundamental and powerful parametric model class. However, identifying the parameters of a linear dynamical system is a venerable task, permitting provably efficient solutions only in special cases. This work shows that the eigenspectrum of unknown linear dynamics can be identified without…
Robust feature-weighted jump models for time-dependent clustering
Mathematically proves SYZ conjecture for conifold transition.