A new clustering algorithm fuses heat diffusion and turning angle for robustness.
problem Cluster similar elements in various fields.
method Combines heat diffusion and maximal turning angle for robust fission clustering.
result The SARFC algorithm outperforms other methods in clustering performance.
Paper addresses online identification and clustering for mixed linear regression models.
problem Online identification and clustering of mixed linear regression models.
method Introduces two online identification algorithms based on the EM principle, proving global convergence without i.i.d. data assumptions.
result Global convergence of the proposed algorithms for mixed linear regression models.
A new clustering method for vector time series using autoregressive dynamics.
problem Clustering of vector time series based on their dynamics is challenging.
method System identification approach using mixture autoregressive models.
result Developed a computationally manageable algorithm k-LMVAR for clustering vector time series.
This work identifies eigenvalues of unknown linear dynamics without full system identification.
problem Identifying parameters of a linear dynamical system is challenging.
method Developed a computationally efficient algorithm to estimate eigenvalues of the state-transition matrix.
result The algorithm can efficiently cluster multi-dimensional time series with temporal offsets and varying lengths.
Paper develops a novel approach to identify clusters of features in multivariate extremes.
problem Understanding the complex structure of multivariate extremes in various fields.
method Optimization-based approach to assess the dependence structure of extremes.
result Estimating clusters of features that best capture the support of extremes.
This paper clusters networks with annotated time-series data using kernel-ARMA and Grassmannian geometry.
problem Clustering networks with annotated time-series data, including state, node, and subnetwork clustering.
method Extract features from time-series data using kernel-ARMA, map onto Grassmannian, and cluster using Riemannian geometry.
result The proposed framework outperforms state-of-the-art clustering schemes on brain-network data.
This study reviews and evaluates clustering methods for single-cell RNA-seq data.
problem Identifying and characterizing novel cell types from single-cell RNA-seq data.
method Review and performance comparison of clustering methods.
result Performance comparison experiments on two datasets.
A new algorithm identifies outliers in Gaussian clustering models.
problem Handling outliers in Gaussian model-based clustering.
method OCLUST algorithm removes least plausible points based on subset log-likelihoods until they adhere to a reference distribution.
result OCLUST inherently estimates the number of outliers.
We focus on spectral clustering of unlabeled graphs and review some results on clustering methods which achieve weak or strong consistent identification in data generated by such models. We also present a new algorithm which appears to perform optimally both theoretically using asymptotic theory and empirically.
Develops a new framework to measure network connectedness across and within markets.
problem Lack of flexible methods to measure network connectedness and its evolution.
method Allows network nodes to be connected in clusters, with shocks orthogonal across clusters and correlated within clusters.
result Demonstrates the effectiveness of the new framework in a detailed empirical analysis of equity markets.
Enhances clustering for functional data, robust to outliers.
problem Challenges of clustering infinite-dimensional functional data and outlier sensitivity.
method Extends OCLUST algorithm to handle functional data, trimming outliers.
result Strong performance in clustering and outlier identification on simulated and real-world datasets.
DADC algorithm improves clustering for data with varying density.
problem Sparse cluster loss and cluster fragmentation in density peak clustering.
method Domain-adaptive density measurement, cluster center self-identification, and cluster self-ensemble.
result DADC achieves more reasonable clustering results on data with varying density.
New method clusters disease subtypes from model explanations.
problem Discovering disease subtypes in noisy, high-dimensional data.
method Train classifier, extract explanations, cluster in explanation space.
result Cluster analysis on model explanations outperforms classical methods.
We present a novel algorithm, called Links, designed to perform online clustering on unit vectors in a high-dimensional Euclidean space. The algorithm is appropriate when it is necessary to cluster data efficiently as it streams in, and is to be contrasted with traditional batch clustering algorithms that have access t…
Two new algorithms improve online clustering of bandits by accelerating cluster identification without strong assumptions.
problem Challenges in accurately identifying unknown user clusters in online bandit settings.
method Proposes UniCLUB and PhaseUniCLUB algorithms with enhanced exploration mechanisms.
result Achieves comparable regret bounds to prior work with weaker assumptions.
New method identifies cluster representatives with minimal pulls.
problem Identifying cluster representatives in multi-armed bandits.
method Fixed confidence approach using confidence intervals.
result Sample complexity matches theoretical lower bound.
It is often the case that, within an online recommender system, multiple users share a common account. Can such shared accounts be identified solely on the basis of the userprovided ratings? Once a shared account is identified, can the different users sharing it be identified as well? Whenever such user identification …
Assessment of risk levels for existing credit accounts is important to the implementation of bank policies and offering financial products. This paper uses cluster analysis of behaviour of credit card accounts to help assess credit risk level. Account behaviour is modelled parametrically and we then implement the behav…
Unified approach for non-stationary and clustered bandits.
problem Solving non-stationary and clustered bandits with overlapping solutions.
method Test of homogeneity for seamless integration of non-stationary and clustered bandits.
result Unified solution framework for change detection and cluster identification.
Proposes clustering and pruning to simplify causal data fusion models.
problem Combining observational and experimental data to identify causal effects.
method Generalizes pruning and clustering operations for multiple data sources.
result Derives conditions for inferring causal effects from simplified models.
The task of clustering a set of objects based on multiple sources of data arises in several modern applications. We propose an integrative statistical model that permits a separate clustering of the objects for each data source. These separate clusterings adhere loosely to an overall consensus clustering, and hence the…
New method detects long-term structures with internal dynamics in time series data.
problem Identifying patterns across time in data growth.
method Adaptive identification of majority overlaps between groups at different time points.
result Detection of persistent structural elements with internal dynamics.
Proposes SAG-DBSCAN for clustering with self-adaptation.
problem Clustering analysis in data mining.
method Uses grey relational matrix and DBSCAN for clustering.
result Demonstrates superior performance compared to other methods.
We present a novel probabilistic clustering model for objects that are represented via pairwise distances and observed at different time points. The proposed method utilizes the information given by adjacent time points to find the underlying cluster structure and obtain a smooth cluster evolution. This approach allows…
Paper addresses global convergence of MLR estimation under weak data conditions.
problem Learning mixed linear regression models with general data conditions.
method Two-step recursive identification algorithm using least squares and EM principles.
result Global convergence and optimal clustering performance established under general data conditions.
In this work a robust clustering algorithm for stationary time series is proposed. The algorithm is based on the use of estimated spectral densities, which are considered as functional data, as the basic characteristic of stationary time series for clustering purposes. A robust algorithm for functional data is then app…
We use statistically validated networks, a recently introduced method to validate links in a bipartite system, to identify clusters of investors trading in a financial market. Specifically, we investigate a special database allowing to track the trading activity of individual investors of the stock Nokia. We find that …
Cluster analysis methods are used to identify homogeneous subgroups in a data set. In biomedical applications, one frequently applies cluster analysis in order to identify biologically interesting subgroups. In particular, one may wish to identify subgroups that are associated with a particular outcome of interest. Con…
EAP clusters evolving data, promoting temporal smoothness and automatic cluster tracking.
problem Clustering time-evolving data with temporal smoothness and automatic cluster identification.
method Evolutionary Affinity Propagation (EAP) on a factor graph exchanging messages between adjacent data snapshots.
result EAP clusters data with temporal smoothness and automatically tracks clusters, outperforming existing methods.
BCV method helps estimate clusters and hyper-parameters in large data sets.
problem Determining the number of clusters in large-scale data.
method Bi-cross validation (BCV) for spectral clustering.
result BCV directly applies to spectral clustering for estimating clusters and hyper-parameters.
Unified framework for DR and clustering using Gromov-Wasserstein.
problem Capturing structure in high-dimensional datasets.
method Distributional reduction framework using Gromov-Wasserstein.
result Unified approach recovers DR and clustering as special cases.
A pairwise clustering approach is applied to the analysis of the Dow Jones index companies, in order to identify similar temporal behavior of the traded stock prices. To this end, the chaotic map clustering algorithm is used, where a map is associated to each company and the correlation coefficients of the financial ti…
In this paper we present deterministic conditions for success of sparse subspace clustering (SSC) under missing data, when data is assumed to come from a Union of Subspaces (UoS) model. We consider two algorithms, which are variants of SSC with entry-wise zero-filling that differ in terms of the optimization problems u…
JojoSCL improves scRNA-seq clustering by reducing intra-cluster dispersion.
problem High dimensionality and sparsity of scRNA-seq data challenge clustering models.
method Integrates shrinkage estimator and contrastive learning for improved clustering.
result JojoSCL outperforms existing methods on ten scRNA-seq datasets.
M-learner estimates treatment effects in mediation models with subgroup identification.
problem Estimating heterogeneous treatment effects in mediation models.
method Four-step procedure: compute conditional effects, construct distance matrix, apply tSNE and K-means clustering, refine clusters.
result Validates robustness and effectiveness in real-world dataset.
New RL method tackles dynamic, heterogeneous data.
problem Temporal non-stationarity and subject heterogeneity in reinforcement learning.
method Alternates between change point detection and cluster identification.
result Improves policy learning by detecting similar dynamics over time and across individuals.
This study examines whether PCA can effectively identify nitrogen pollution sources in rivers.
problem Identifying pollution sources in rivers for effective environmental management.
method Principal Component Analysis and its modifications, along with Independent Component Analysis and Factor Analysis, are applied to nitrogen pollution source identification.
result PCA and related techniques can be powerful tools for uncovering nitrogen pollution sources in rivers.
Method detects and locates eavesdropping in optical links.
problem Detect and locate eavesdropping in optical links with small power losses.
method Cluster-based approach using OPM data at receiver and in-line OPM data for localization.
result Subtle eavesdropping losses can be detected and localized using OPM data.
New method identifies valid IVs for bi-directional MR with invalid instruments.
problem Estimating causal effects from observational data with invalid instruments and unmeasured confounding.
method Theoretical investigation and cluster fusion-like method to discover valid IV sets.
result Theoretical demonstration and experimental validation of the method's effectiveness.
We define transit clusters to simplify causal diagrams and preserve their essential properties.
problem Clustering variables in causal diagrams can alter essential properties of causal effects.
method We define transit clusters and provide an algorithm to find them, ensuring they preserve causal effect identifiability.
result Transit clusters simplify causal effect identification and maintain their essential properties.
Data clustering is a fundamental problem with a wide range of applications. Standard methods, eg the k-means method, usually require solving a non-convex optimization problem. Recently, total variation based convex relaxation to the k-means model has emerged as an attractive alternative for data clustering. However…
Vehicle recognition and classification have broad applications, ranging from traffic flow management to military target identification. We demonstrate an unsupervised method for automated identification of moving vehicles from roadside audio sensors. Using a short-time Fourier transform to decompose audio signals, we t…
Paper uses DBSCAN variation to detect ship anomalies.
problem Detecting anomalous ship behavior.
method Variation of DBSCAN algorithm applied to AIS data.
result Alternative anomaly metric is more statistically informative.
We study clustering algorithms based on neighborhood graphs on a random sample of data points. The question we ask is how such a graph should be constructed in order to obtain optimal clustering results. Which type of neighborhood graph should one choose, mutual k-nearest neighbor or symmetric k-nearest neighbor? What …
This paper advocates a novel framework for segmenting a dataset in a Riemannian manifold M into clusters lying around low-dimensional submanifolds of M. Important examples of M, for which the proposed clustering algorithm is computationally efficient, are the sphere, the set of positive definite matrices, and the…
Two new methods improve clustering with missing data.
problem Handling missing data in Gaussian Mixture Models.
method Proposes two methods using Monte Carlo Expectation-Maximization (MCEM) for data augmentation.
result Proposed methods outperform multiple imputation in clustering and density estimation.
This paper improves robust cluster enumeration for RES data.
problem Challenges in determining optimal clusters in noisy data.
method Generalizes robust Bayesian cluster enumeration for RES mixtures.
result Significant robustness improvement over existing methods.
Analyzing large X-ray diffraction (XRD) datasets is a key step in high-throughput mapping of the compositional phase diagrams of combinatorial materials libraries. Optimizing and automating this task can help accelerate the process of discovery of materials with novel and desirable properties. Here, we report a new met…