A new clustering algorithm fuses heat diffusion and turning angle for robustness.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper addresses online identification and clustering for mixed linear regression models.
A new clustering method for vector time series using autoregressive dynamics.
Paper develops a novel approach to identify clusters of features in multivariate extremes.
This paper clusters networks with annotated time-series data using kernel-ARMA and Grassmannian geometry.
This study reviews and evaluates clustering methods for single-cell RNA-seq data.
Linear dynamical systems are a fundamental and powerful parametric model class. However, identifying the parameters of a linear dynamical system is a venerable task, permitting provably efficient solutions only in special cases. This work shows that the eigenspectrum of unknown linear dynamics can be identified without…
We focus on spectral clustering of unlabeled graphs and review some results on clustering methods which achieve weak or strong consistent identification in data generated by such models. We also present a new algorithm which appears to perform optimally both theoretically using asymptotic theory and empirically.
Develops a new framework to measure network connectedness across and within markets.
Enhances clustering for functional data, robust to outliers.
Clustering, or unsupervised classification, is a task often plagued by outliers. Yet there is a paucity of work on handling outliers in clustering. Outlier identification algorithms tend to fall into three broad categories: outlier inclusion, outlier trimming, and post hoc outlier identification methods, with the forme…
DADC algorithm improves clustering for data with varying density.
New method clusters disease subtypes from model explanations.
We present a novel algorithm, called Links, designed to perform online clustering on unit vectors in a high-dimensional Euclidean space. The algorithm is appropriate when it is necessary to cluster data efficiently as it streams in, and is to be contrasted with traditional batch clustering algorithms that have access t…
Two new algorithms improve online clustering of bandits by accelerating cluster identification without strong assumptions.
It is often the case that, within an online recommender system, multiple users share a common account. Can such shared accounts be identified solely on the basis of the userprovided ratings? Once a shared account is identified, can the different users sharing it be identified as well? Whenever such user identification …
New method identifies cluster representatives with minimal pulls.
Assessment of risk levels for existing credit accounts is important to the implementation of bank policies and offering financial products. This paper uses cluster analysis of behaviour of credit card accounts to help assess credit risk level. Account behaviour is modelled parametrically and we then implement the behav…
Unified approach for non-stationary and clustered bandits.
Proposes clustering and pruning to simplify causal data fusion models.
The task of clustering a set of objects based on multiple sources of data arises in several modern applications. We propose an integrative statistical model that permits a separate clustering of the objects for each data source. These separate clusterings adhere loosely to an overall consensus clustering, and hence the…
New method detects long-term structures with internal dynamics in time series data.
Proposes SAG-DBSCAN for clustering with self-adaptation.
We present a novel probabilistic clustering model for objects that are represented via pairwise distances and observed at different time points. The proposed method utilizes the information given by adjacent time points to find the underlying cluster structure and obtain a smooth cluster evolution. This approach allows…
Paper addresses global convergence of MLR estimation under weak data conditions.
In this work a robust clustering algorithm for stationary time series is proposed. The algorithm is based on the use of estimated spectral densities, which are considered as functional data, as the basic characteristic of stationary time series for clustering purposes. A robust algorithm for functional data is then app…
One challenge impeding the analysis of terabyte scale x-ray scattering data from the Linac Coherent Light Source LCLS, is determining the number of clusters required for the execution of traditional clustering algorithms. Here we demonstrate that previous work using bi-cross validation (BCV) to determine the number of …
We use statistically validated networks, a recently introduced method to validate links in a bipartite system, to identify clusters of investors trading in a financial market. Specifically, we investigate a special database allowing to track the trading activity of individual investors of the stock Nokia. We find that …
Cluster analysis methods are used to identify homogeneous subgroups in a data set. In biomedical applications, one frequently applies cluster analysis in order to identify biologically interesting subgroups. In particular, one may wish to identify subgroups that are associated with a particular outcome of interest. Con…
EAP clusters evolving data, promoting temporal smoothness and automatic cluster tracking.
Unified framework for DR and clustering using Gromov-Wasserstein.
A pairwise clustering approach is applied to the analysis of the Dow Jones index companies, in order to identify similar temporal behavior of the traded stock prices. To this end, the chaotic map clustering algorithm is used, where a map is associated to each company and the correlation coefficients of the financial ti…
In this paper we present deterministic conditions for success of sparse subspace clustering (SSC) under missing data, when data is assumed to come from a Union of Subspaces (UoS) model. We consider two algorithms, which are variants of SSC with entry-wise zero-filling that differ in terms of the optimization problems u…
JojoSCL improves scRNA-seq clustering by reducing intra-cluster dispersion.
M-learner estimates treatment effects in mediation models with subgroup identification.
New RL method tackles dynamic, heterogeneous data.
This study examines whether PCA can effectively identify nitrogen pollution sources in rivers.
Method detects and locates eavesdropping in optical links.
New method identifies valid IVs for bi-directional MR with invalid instruments.
We define transit clusters to simplify causal diagrams and preserve their essential properties.
Data clustering is a fundamental problem with a wide range of applications. Standard methods, eg the -means method, usually require solving a non-convex optimization problem. Recently, total variation based convex relaxation to the -means model has emerged as an attractive alternative for data clustering. However…
Vehicle recognition and classification have broad applications, ranging from traffic flow management to military target identification. We demonstrate an unsupervised method for automated identification of moving vehicles from roadside audio sensors. Using a short-time Fourier transform to decompose audio signals, we t…
Paper uses DBSCAN variation to detect ship anomalies.
We study clustering algorithms based on neighborhood graphs on a random sample of data points. The question we ask is how such a graph should be constructed in order to obtain optimal clustering results. Which type of neighborhood graph should one choose, mutual k-nearest neighbor or symmetric k-nearest neighbor? What …
This paper advocates a novel framework for segmenting a dataset in a Riemannian manifold into clusters lying around low-dimensional submanifolds of . Important examples of , for which the proposed clustering algorithm is computationally efficient, are the sphere, the set of positive definite matrices, and the…
Two new methods improve clustering with missing data.
This paper improves robust cluster enumeration for RES data.
Analyzing large X-ray diffraction (XRD) datasets is a key step in high-throughput mapping of the compositional phase diagrams of combinatorial materials libraries. Optimizing and automating this task can help accelerate the process of discovery of materials with novel and desirable properties. Here, we report a new met…