INCAD clusters and detects anomalies in streaming data without thresholds.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new clustering method reduces time and memory usage for massive datasets.
Locally adaptive clustering for tree delineation.
Study finds the cutoff for exact recovery in Gaussian mixture models.
The binary symmetric stochastic block model deals with a random graph of vertices partitioned into two equal-sized clusters, such that each pair of vertices is connected independently with probability within clusters and across clusters. In the asymptotic regime of and for fixe…
The MBO scheme for data clustering is analyzed in the large data limit, proving convergence to optimal partition problems.
Clustering explores meaningful patterns in the non-labeled data sets. Cluster Ensemble Selection (CES) is a new approach, which can combine individual clustering results for increasing the performance of the final results. Although CES can achieve better final results in comparison with individual clustering algorithms…
Resolving a conjecture of Abbe, Bandeira and Hall, the authors have recently shown that the semidefinite programming (SDP) relaxation of the maximum likelihood estimator achieves the sharp threshold for exactly recovering the community structure under the binary stochastic block model of two equal-sized clusters. The s…
Subspace clustering refers to the problem of clustering high-dimensional data points into a union of low-dimensional linear subspaces, where the number of subspaces, their dimensions and orientations are all unknown. In this paper, we propose a variation of the recently introduced thresholding-based subspace clustering…
We consider the problem of Gaussian mixture clustering in the high-dimensional limit where the data consists of points in dimensions, and stays finite. Using exact but non-rigorous methods from statistical physics, we determine the critical value of and the distance between…
We consider the problem of clustering a set of high-dimensional data points into sets of low-dimensional linear subspaces. The number of subspaces, their dimensions, and their orientations are unknown. We propose a simple and low-complexity clustering algorithm based on thresholding the correlations between the data po…
Paper explores limits of high-order clustering with planted structures.
We consider the problem of clustering noisy high-dimensional data points into a union of low-dimensional subspaces and a set of outliers. The number of subspaces, their dimensions, and their orientations are unknown. A probabilistic performance analysis of the thresholding-based subspace clustering (TSC) algorithm intr…
In this work we propose a simple and easily parallelizable algorithm for multiway graph partitioning. The algorithm alternates between three basic components: diffusing seed vertices over the graph, thresholding the diffused seeds, and then randomly reseeding the thresholded clusters. We demonstrate experimentally that…
Method selects the best deep learner for time-series prediction using Bayesian networks.
DSSP improves deep learning training speed by dynamically adjusting staleness thresholds.
This study aimed to find temporal clusters for several commodity prices using the threshold non-linear autoregressive model. It is expected that the process of determining the commodity groups that are time-dependent will advance the current knowledge about the dynamics of co-moving and coherent prices, and can serve a…
Characterizes optimal reconstruction error in high-dimensional Gaussian mixtures.
Based on the daily data of American and Chinese stock markets, the dynamic behavior of a financial network with static and dynamic thresholds is investigated. Compared with the static threshold, the dynamic threshold suppresses the large fluctuation induced by the cross-correlation of individual stock prices, and leads…
A new SOM method learns from both labeled and unlabeled data.
We apply RMT, Network and MF-DFA methods to investigate correlation, network and multifractal properties of 20 global financial indices. We compare results before and during the financial crisis of 2008 respectively. We find that the network method gives more useful information about the formation of clusters as compar…
New method clusters tensors with heteroskedastic noise.
Model-based clustering defines population level clusters relative to a model that embeds notions of similarity. Algorithms tailored to such models yield estimated clusters with a clear statistical interpretation. We take this view here and introduce the class of G-block covariance models as a background model for varia…
We consider the effects of the 2008 global financial crisis on the global stock market before, during, and after the crisis. We generate complex networks from a cross-correlation matrix such as the threshold network (TN) and the minimal spanning tree (MST). In the threshold network, we assign a threshold value by using…
We consider the effects of the global financial crisis through a local Korean financial market around the 2008 crisis. We analyze 185 individual stock prices belonging to the KOSPI (Korea Composite Stock Price Index), cosidering three time periods: the time before, during, and after the crisis. The complex networks gen…
In this paper, a similarity-driven cluster merging method is proposed for unsuper-vised fuzzy clustering. The cluster merging method is used to resolve the problem of cluster validation. Starting with an overspecified number of clusters in the data, pairs of similar clusters are merged based on the proposed similarity-…
New algorithms detect communities in sparse graphs with labeled data.
The paper examines the optimality of kernel methods in high-dimensional clustering.
This paper tackles exact recovery of clusters in a stochastic Ising model on a SBM graph.
The problem of clustering noisy and incompletely observed high-dimensional data points into a union of low-dimensional subspaces and a set of outliers is considered. The number of subspaces, their dimensions, and their orientations are assumed unknown. We propose a simple low-complexity subspace clustering algorithm, w…
New algorithm achieves strong consistency in binary non-uniform hypergraph classification.
A novel multi-resolution cluster detection (MCD) method is proposed to identify irregularly shaped clusters in space. Multi-scale test statistic on a single cell is derived based on likelihood ratio statistic for Bernoulli sequence, Poisson sequence and Normal sequence. A neighborhood variability measure is defined to …
Extends MSC for triclustering tensors, using DBSCAN to find clusters.
Using data from 92 indices of stock exchanges worldwide, I analize the cluster formation and evolution from 2007 to 2010, which includes the Subprime Mortgage Crisis of 2008, using asset graphs based on distance thresholds. I also study the survivability of connections and of clusters through time and the influence of …
Paper improves anomaly detection by using non-uniform random choices in isolation forests.
A natural approach to analyze interaction data of form "what-connects-to-what-when" is to create a time-series (or rather a sequence) of graphs through temporal discretization (bandwidth selection) and spatial discretization (vertex contraction). Such discretization together with non-negative factorization techniques c…
funLOCI identifies clusters in functional data.
New method for triclustering with reduced arbitrariness.
Using data from world stock exchange indices prior to and during periods of global financial crises, clusters and networks of indices are built for different thresholds and diverse periods of time, so that it is then possible to analyze how clusters are formed according to correlations among indices and how they evolve…
An adaptive clustering algorithm learns from evolving data without manual tuning.
The paper compares clustering techniques for personalized food kits.
New algorithm detects communities near KS threshold with optimal rate, even in noisy conditions.
Clusters of withdrawals emerge in banks due to latent fragility.
New clustering methods for binary data using combinatorial optimization.
For a certain class of distributions, we prove that the linear programming relaxation of -medoids clustering---a variant of -means clustering where means are replaced by exemplars from within the dataset---distinguishes points drawn from nonoverlapping balls with high probability once the number of points drawn a…
Framework for multi-scale clustering using phase transitions.
Study examines local extrema and crossing statistics in financial markets.
Sparse spectral decomposition identifies overlapping communities in networks.