The paper defines matrices related to cluster transformations and proves certain quivers have no maximal sequences.
problem Proving quivers associated with once-punctured surfaces do not have maximal green or reddening sequences.
method Defining matrices related to cluster transformations and showing their relationships to the Jacobian and C-matrix.
result Quivers associated with once-punctured surfaces do not have maximal green or reddening sequences.
nTreeClus clusters categorical sequences using tree-based learners and k-mers.
problem Challenges in clustering categorical and sequential data.
method nTreeClus uses Tree-based Learners, k-mers, and autoregressive models for categorical time series.
result nTreeClus outperformed baseline methods in various validation metrics.
Clustered attention improves transformer efficiency for large sequences.
problem Quadratic complexity of transformer attention matrix for large sequences.
method Group queries into clusters, compute attention only for centroids, and use centroids to approximate key/query dot products.
result Linear complexity with respect to sequence length for a fixed number of clusters.
Paper proposes SLINK clustering for nonparametric data sequences with improved consistency.
problem Nonparametric clustering of data sequences from unknown distributions.
method Exponentially consistent nonparametric SLINK clustering algorithm.
result SLINK clustering achieves exponential consistency under less strict conditions.
The paper clusters sequences from unknown distributions using k-medoids.
problem Clustering sequences from unknown composite distributions.
method k-medoids algorithm for sequences with composite distributions.
result Error probability decreases exponentially with increasing sample size.
A measure called relative cluster entropy distinguishes between correlated and uncorrelated sequences.
problem Distinguishing between sequences with different correlation degrees.
method Minimum relative entropy principle applied to cluster partitions of power-law correlated sequences.
result Optimal Hurst exponents are selected for market price series, indicating non-markovianity.
Infinite rank surface cluster algebras extend traditional concepts to surfaces with accumulation points.
problem Extending surface cluster algebras to infinite surfaces with accumulation points.
method Consider infinite mutation sequences and hyperbolic structures to define cluster variables as lambda lengths of arcs.
result Established transitivity of infinite mutation sequences on triangulations of infinite surfaces and provided expansion formulas for cluster variables.
Proposes a new model for clustering event sequences.
problem Clustering asynchronous event sequences with diverse triggering patterns.
method Dirichlet mixture model of Hawkes processes with variational Bayesian inference.
result Automatic learning of the number of clusters and robustness to misspecification.
SGT embeds sequence features efficiently for clustering and classification.
problem Challenges in extracting long-term dependencies from unstructured sequence data.
method Sequence Graph Transform (SGT) embeds varying amounts of short- to long-term dependencies efficiently.
result SGT features yield superior results in sequence clustering and classification.
New k-means method clusters radar image sequences using SPD matrices.
problem Clustering radar image sequences efficiently.
method Developed k-means on SPD matrices for non-Euclidean data. result Effective clustering of radar image sequences via SPD matrices.
Efficient tests detect outlying sequences without knowing their distributions.
problem Detecting outlying sequences among multiple distributions without prior knowledge.
method Distribution clustering-based tests with linear complexity and exponential consistency.
result Tests are computationally efficient and perform similarly to existing methods.
A novel multi-resolution cluster detection (MCD) method is proposed to identify irregularly shaped clusters in space. Multi-scale test statistic on a single cell is derived based on likelihood ratio statistic for Bernoulli sequence, Poisson sequence and Normal sequence. A neighborhood variability measure is defined to …
Paper clusters event sequences using a reinforcement learning approach with policy mixture model.
problem Clustering event sequences with varying temporal patterns.
method Reinforcement learning with a policy mixture model, decomposing sequences into states and actions.
result Effective clustering of event sequences into underlying policies, outperforming existing methods.
The study compares different scRNA sequencing methods using a high-dimensional dataset.
problem To identify unique characteristics of different scRNA sequencing methods.
method Quantitative comparison through clustering analysis of a high-dimensional dataset.
result Identifies unique characteristics associated with different scRNA sequencing methods.
Flexible models cluster RNA sequencing data.
problem Clustering discrete data from RNA sequencing studies.
method Finite mixtures of multivariate Poisson-log normal factor analyzers with constraints.
result Models give favorable clustering performance on real and simulated data.
New method embeds RNN Seq2Seq models to visualize spatiotemporal data.
problem Visualizing and interpreting spatiotemporal data in sequence prediction tasks.
method Embedding approach to visualize and interpret RNN Seq2Seq model representations.
result Embedding space projections of RNN Seq2Seq models capture spatiotemporal dynamics.
Forest Fire Clustering discovers cell types from single-cell data.
problem Discovering cell types from large-scale single-cell sequencing data.
method Iterative label propagation and parallelized Monte Carlo simulation.
result Forest Fire Clustering outperforms state-of-the-art methods on diverse benchmarks.
A system is presented that segments, clusters and predicts musical audio in an unsupervised manner, adjusting the number of (timbre) clusters instantaneously to the audio input. A sequence learning algorithm adapts its structure to a dynamically changing clustering tree. The flow of the system is as follows: 1) segment…
Bayesian context trees capture complex dependencies in categorical sequences.
problem Complex, long-range dependencies in categorical sequences are not well captured by simple models.
method Parsimonious Bayesian context trees with model-based agglomerative clustering for efficient inference.
result The proposed framework outperforms existing models on real-world data.
We present a large catalog of optically selected galaxy clusters from the application of a new Gaussian Mixture Brightest Cluster Galaxy (GMBCG) algorithm to SDSS Data Release 7 data. The algorithm detects clusters by identifying the red sequence plus Brightest Cluster Galaxy (BCG) feature, which is unique for galaxy c…
We suggest a novel method of clustering and exploratory analysis of temporal event sequences data (also known as categorical time series) based on three-dimensional data grid models. A data set of temporal event sequences can be represented as a data set of three-dimensional points, each point is defined by three varia…
Paper proposes a new method for sparse spectral clustering on Stiefel manifold.
problem Sparse spectral clustering on Stiefel manifold with nonsmooth and nonconvex objective.
method Proposes a manifold proximal linear method (ManPL) to solve the original SSC formulation.
result Demonstrates the advantage of ManPL over existing methods on single-cell RNA sequencing data.
JojoSCL improves scRNA-seq clustering by reducing intra-cluster dispersion.
problem High dimensionality and sparsity of scRNA-seq data challenge clustering models.
method Integrates shrinkage estimator and contrastive learning for improved clustering.
result JojoSCL outperforms existing methods on ten scRNA-seq datasets.
KT combines treelets with kernel functions for hierarchical clustering.
problem Hierarchical clustering of non-numeric data.
method Combines treelets and kernel functions to handle non-numeric data.
result KT effectively clusters non-numeric data.
Despite its popularity, it is widely recognized that the investigation of some theoretical aspects of clustering has been relatively sparse. One of the main reasons for this lack of theoretical results is surely the fact that, whereas for other statistical problems the theoretical population goal is clearly defined (as…
A deep learning approach for clustering time series of varying lengths.
problem Clustering time series with variable lengths and temporal relations.
method Recurrent Deep Divergence-based Clustering framework.
result Outperforms previous methods on benchmark datasets.
Method extracts knowledge from LSTM for sequence validation.
problem Validating sequences generated by unknown automata.
method Clustering hidden states to build a corresponding automaton.
result Automaton accurately predicts sequence validity.
IMPACC improves consensus clustering for bioinformatics data.
problem Consensus clustering's inefficiency and lack of interpretability for large-scale data.
method Ensemble minipatch co-occurrences, adaptive sampling of observations and features.
result Significantly improved accuracy and interpretability with substantial computational savings.
New framework quantifies uncertainty in flexible density-based clustering.
problem Uncertainty quantification in clustering with non-parametric density estimation.
method Martingale posterior distributions and density-based clustering.
result Efficient GPU-compatible inference on clustering structures with uncertainty.
Proposes selective inference for testing differences in means between clusters.
problem Inflated type I error rate when testing differences in means between clusters.
method Selective inference approach to control selective type I error rate.
result Controls selective type I error rate by accounting for data-driven cluster definition.
New method uses dendrograms for better mixture model selection and clustering.
problem Selecting the correct number of components in finite mixture models.
method Hierarchical clustering tree derived from overfitted latent mixing measures.
result Consistently selects the true number of mixing components and optimal convergence rate for parameter estimation.
A new smooth edit distance for easier optimization in machine learning.
problem Hard optimization of edit distance for variable-length sequences.
method Soft edit distance (SED) as a differentiable approximation.
result SED can be optimized with gradient methods and used for clustering.
ElbowSig assesses clustering structure at multiple scales.
problem Selecting optimal number of clusters in unsupervised learning.
method Formalizes elbow heuristic with a normalized discrete curvature statistic.
result Validates multiscale clustering structure over various resolutions.
We consider a problem of clustering a sequence of multinomial observations by way of a model selection criterion. We propose a form of a penalty term for the model selection procedure. Our approach subsumes both the conventional AIC and BIC criteria but also extends the conventional criteria in a way that it can be app…
This paper discusses about an R package that implements the Pattern Sequence based Forecasting (PSF) algorithm, which was developed for univariate time series forecasting. This algorithm has been successfully applied to many different fields. The PSF algorithm consists of two major parts: clustering and prediction. The…
Introduces BWMD, a new distance measure for DNA and malware clustering.
problem Shortcomings of previous compression-based distance metrics.
method Embeds sequences into a fixed-length feature vector.
result Significantly improved clustering performance on larger malware corpora.
New method improves cluster instability estimation for better k selection.
problem Selecting the optimal number of clusters in cluster analysis.
method Developed a normalized cluster instability measure to correct for cluster size distribution.
result Normalized instability measure outperforms current methods across all possible k. Proposes models to analyze irregular healthcare time series data.
problem Irregular timestamps in healthcare time series data.
method Data augmentation, temporal coarsening, MultiResolution Ensemble (MRE) model.
result Improves mAP on mortality prediction task from 51.53% to 53.92%.
New model learns human actions from video data without labels.
problem Learning human actions from unlabeled video data.
method Clustering-aware structure-constrained low-rank representation (CS-LRR) model integrating spectral clustering and hierarchical subspace clustering.
result Efficiently learns human action attributes from video data without requiring labeled data.
Scalable hybrid HMM with Gaussian Process for time-series data clustering.
problem Large number of parameters and long sequences in time-series data make HMM-GPSM training difficult.
method Stochastic Variational Inference (SVI) for long sequences and reparameterized random Fourier features (R-RFF) for large data points.
result Significant reduction in training time and improved hidden-state estimation accuracy.
We present a global optimization algorithm for clustering data given the ratio of likelihoods that each pair of data points is in the same cluster or in different clusters. To define a clustering solution in terms of pairwise relationships, a necessary and sufficient condition is that belonging to the same cluster sati…
Selective inference controls Type I error in k-means clustering tests.
problem Inflated Type I error in classical hypothesis tests for k-means clusters.
method Selective inference approach to control Type I error.
result Proposes a computable finite-sample p-value for selective inference.
We investigate an efficient context-dependent clustering technique for recommender systems based on exploration-exploitation strategies through multi-armed bandits over multiple users. Our algorithm dynamically groups users based on their observed behavioral similarity during a sequence of logged activities. In doing s…
An efficient algorithm for k-median clustering in a sequential setting without substitutions.
problem Clustering a sequence of examples without being able to substitute centers later.
method An efficient algorithm with a multiplicative approximation factor of twice the offline algorithm's factor, and an optimal offline algorithm.
result The efficient algorithm achieves a good approximation of the optimal offline solution.
New algorithms improve decision-making with limited offline data.
problem Using limited offline data to cluster users for better decision-making.
method Proposed two algorithms: Off-C2LUB and Off-CLUB to address data insufficiency.
result Both algorithms outperform existing methods under limited offline user data.
sCSC clusters data without prior assumptions, revealing natural groupings.
problem Clustering data without prior knowledge of its structure.
method sCSC performs binary splittings maximizing dissimilarity, producing a binary tree.
result Clusters emerge naturally from the binary tree, revealing data structure.
Unified framework for clustering with sparse convex combinations.
problem Challenges in subspace clustering with limited labelled data.
method Spectral-based sparse subspace representation with extensions to constrained and active learning.
result Effective and competitive clustering results on simulated and real data.
Tree-SNE combines t-SNE and hierarchical clustering for data visualization.
problem Data visualization and clustering in complex datasets.
method Stacked one-dimensional t-SNE embeddings and alpha-clustering.
result Effective hierarchical clustering and visualization of various datasets.