Extends MSC for triclustering tensors, using DBSCAN to find clusters.
problem Finding clusters in multi-slice triclustering of tensors with unknown cluster sizes.
method Extends Multi-Slice Clustering (MSC) with DBSCAN to find clusters in tensors.
result Can find clusters in tensors that are sums of multiple rank-one tensors.
New method for multiway clustering of 3rd order tensors.
problem Clustering of 3rd order tensors.
method MCAM method based on affinity matrix and clustering.
result Competitive results on synthetic and real datasets.
New method clusters tensors with heteroskedastic noise.
problem Clustering tensors with varying noise levels.
method Two-stage method: subspace estimation followed by approximate k-means. result Proves exact clustering for SNR above computational limit.
Dynamic tensor data are becoming prevalent in numerous applications. Existing tensor clustering methods either fail to account for the dynamic nature of the data, or are inapplicable to a general-order tensor. Also there is often a gap between statistical guarantee and computational efficiency for existing tensor clust…
New method for triclustering with reduced arbitrariness.
problem Need for reduced arbitrariness in specifying cluster size.
method Spectral decomposition of tensor slices and intersection of clusters.
result Effective triclustering on synthetic and real-world data.
Cluster analysis is a fundamental tool for pattern discovery of complex heterogeneous data. Prevalent clustering methods mainly focus on vector or matrix-variate data and are not applicable to general-order tensors, which arise frequently in modern scientific and business applications. Moreover, there is a gap between …
Paper develops a method to robustly cluster tensors with outliers.
problem Clustering tensors contaminated by outliers or sample-specific corruptions.
method Transformed Tensor Low-Rank Representation (OR-TLRR) method.
result Provably recovers row space of clean data and detects outliers.
Proposes a new model for clustering passenger trips considering hierarchical and multi-dimensional data.
problem Clustering passenger trips with hierarchical and multi-dimensional data, especially in large-scale transportation systems.
method Tensor Dirichlet Process Multinomial Mixture (Tensor-DPMM) model, incorporating Dirichlet Process for automatic cluster number determination and tensor representation for multi-mode data.
result Automatic determination of the number of clusters and improved clustering quality.
Paper proposes a tensor model for clustering noisy multi-view data.
problem Clustering noisy multi-view data with non-uniform variances.
method Nested matrix-tensor model for best rank-one approximation.
result Theoretical results predict the exact accuracy of clustering.
Paper proposes efficient methods for high-order clustering in tensor block models.
problem High-order clustering of multiway datasets in neuroimaging, genomics, etc.
method Tensor block model and computationally efficient algorithms (HLloyd, HSC)
result Achieves high-order exact clustering with statistical optimality and computational efficiency.
Paper explores limits of high-order clustering with planted structures.
problem Statistical and computational limits of high-order clustering with planted structures.
method Developed methods for detection and recovery of clusters, identified signal-to-noise ratio boundaries.
result Sharp boundaries of signal-to-noise ratio for statistical and computational feasibility.
MultiwayPAM clusters LLM-as-a-Judge scores to reveal evaluator bias.
problem Revealing bias in LLM evaluations of text quality.
method Tensor clustering with MultiwayPAM method.
result Medoids reveal cluster membership of questions, answerers, and evaluators.
Develops a new tensor model for clustering with degree correction.
problem Clustering with unknown degree heterogeneity in multiway data.
method Degree-corrected tensor block model with estimation guarantees.
result Demonstrates an intrinsic statistical-to-computational gap for tensors of order three or greater.
Paper proposes a new method to improve clustering ensemble performance.
problem Improving clustering ensemble performance by refining co-association matrix.
method Low-rank tensor approximation to derive coherent-link matrix and refine co-association matrix.
result The proposed method achieves breakthrough in clustering performance compared to state-of-the-art methods.
Paper improves MVSC using tensor low-rank modeling.
problem Improving multi-view spectral clustering.
method Structured tensor low-rank norm for MVSC optimization.
result Proposed method outperforms state-of-the-art methods.
Proposes a new model for clustering passenger trajectories with graphs.
problem Hierarchical trip structure, inaccurate clustering number, and lack of spatial semantic graphs.
method Tensor Dirichlet Process Multinomial Mixture model with graphs and a tensor version of Collapsed Gibbs Sampling.
result Automatic determination of the number of clusters and better cluster quality.
Study quantifies performance gap between tensor and matrix-based approaches in nested matrix-tensor model.
problem Estimating a planted signal in a nested matrix-tensor model.
method Comparing tensor-based and matrix-based approaches for best rank-one approximation of tensor data.
result Derives precise algorithmic threshold for the unfolding approach and shows BBP-type transition behavior.
Develops a tensor mixture model for high-dimensional data clustering.
problem Jointly modeling and clustering tensors in high dimensions.
method High-dimensional tensor mixture model with plausible dimension reduction assumptions. EHCMA algorithm for efficient estimation.
result The HECM algorithm converges geometrically to a neighborhood within statistical precision of the true parameter.
A tensor provides a concise way to codify the interdependence of complex data. Treating a tensor as a d-way array, each entry records the interaction between the different indices. Clustering provides a way to parse the complexity of the data into more readily understandable information. Clustering methods are heavily …
In this paper we present a method for the unsupervised clustering of high-dimensional binary data, with a special focus on electronic healthcare records. We present a robust and efficient heuristic to face this problem using tensor decomposition. We present the reasons why this approach is preferable for tasks such as …
Co-Clustering, the problem of simultaneously identifying clusters across multiple aspects of a data set, is a natural generalization of clustering to higher-order structured data. Recent convex formulations of bi-clustering and tensor co-clustering, which shrink estimated centroids together using a convex fusion penalt…
TACE unifies scalar and tensorial modeling in Cartesian space for accurate, stable, and efficient atomistic predictions.
problem Complexity and challenges in equivariant atomistic machine learning models.
method Tensor Atomic Cluster Expansion (TACE) in Cartesian space, decomposing local environments into irreducible Cartesian tensors (ICT).
result Universal invariant and equivariant embeddings, enabling explicit control at inference.
We consider the problem of identifying multiway block structure from a large noisy tensor. Such problems arise frequently in applications such as genomics, recommendation system, topic modeling, and sensor network localization. We propose a tensor block model, develop a unified least-square estimation, and obtain the t…
The performance of most the clustering methods hinges on the used pairwise affinity, which is usually denoted by a similarity matrix. However, the pairwise similarity is notoriously known for its vulnerability of noise contamination or the imbalance in samples or features, and thus hinders accurate clustering. To tackl…
Proposes a tensor Laplacian-based method for better subspace clustering of non-uniformly distributed data.
problem LRR's inability to handle non-uniform data distribution and local information loss.
method Tensor Laplacian Regularized Low-Rank Representation (TLRR) using hypergraph model and tensor Laplacian algorithm.
result Higher accuracy and precision in subspace clustering compared to state-of-the-art methods.
In many real-world applications, data are often unlabeled and comprised of different representations/views which often provide information complementary to each other. Although several multi-view clustering methods have been proposed, most of them routinely assume one weight for one view of features, and thus inter-vie…
Exact partitioning of high-order planted models achieved through convex optimization.
problem Efficiently partitioning hypergraphs generated by high-order planted models.
method Solving a computationally efficient convex optimization problem with a tensor nuclear norm constraint.
result Exact recovery of true underlying cluster structures with high probability.
Physical activity levels are an important predictor of cardiovascular health and increasingly being measured by sensors, like accelerometers. Accelerometers produce rich multivariate data that can inform important clinical decisions related to individual patients and public health. The CHAMPION study, a study of youth …
Objective Function Mismatch (OFM) occurs when the optimization of one objective has a negative impact on the optimization of another objective. In this work we study OFM in deep clustering, and find that the popular autoencoder-based approach to deep clustering can lead to both reduced clustering performance, and a sig…
Paper optimizes clustering for multi-layer networks and discrete mixtures.
problem Optimizing clustering in multi-layer networks and discrete mixtures.
method Two-stage method: tensor-based initialization and likelihood-based refinement.
result Achieves minimax optimal error rate for multi-layer networks and discrete mixtures.
Dimensionality reduction techniques play an essential role in data analytics, signal processing and machine learning. Dimensionality reduction is usually performed in a preprocessing stage that is separate from subsequent data analysis, such as clustering or classification. Finding reduced-dimension representations tha…
Paper develops a method for causal representation learning from irregular tensors.
problem Complex patterns in high-dimensional, irregular tensor data.
method Novel causal formulation and CaRTeD framework integrating temporal causal representation learning with irregular tensor decomposition.
result Framework provides theoretical guarantees and outperforms state-of-the-art techniques.
We introduce a tensor-based clustering method to extract sparse, low-dimensional structure from high-dimensional, multi-indexed datasets. This framework is designed to enable detection of clusters of data in the presence of structural requirements which we encode as algebraic constraints in a linear program. Our cluste…
Matrix factorizations and their extensions to tensor factorizations and decompositions have become prominent techniques for linear and multilinear blind source separation (BSS), especially multiway Independent Component Analysis (ICA), NonnegativeMatrix and Tensor Factorization (NMF/NTF), Smooth Component Analysis (Smo…
Paper perfect clusters sparse, diverse multilayer networks.
problem Clustering sparse, diverse multilayer networks.
method Tensor-based methodology pooling all layers' information.
result Achieves perfect clustering under sparser conditions than previous models.
MSFA clusters high-dimensional spatial data using spline-based covariance structures.
problem Clustering high-dimensional spatial data with flexible covariance structures.
method Mixture of spatial factor analyzers with spline-based covariance and matrix variate factor analyzers for dimensionality reduction.
result Proposed models accurately infer and differentiate distinct spatial patterns in tensor-variate data.
For a symplectic manifold with quantizing line bundle, a choice of almost complex structure determines a Laplacian acting on tensor powers of the bundle. For high tensor powers Guillemin-Uribe showed that there is a well-defined cluster of low-lying eigenvalues, whose distribution is described by a spectral density fun…
Study graph-based algorithms for multi-manifold clustering with sufficient conditions.
problem Clustering data from a union of manifolds with different dimensions and intersections.
method Investigate sufficient conditions for similarity graphs to capture geometric information.
result High probability error bounds for spectral approximation of tensorized Laplacian.
New approach learns mixtures of linear dynamical systems without separation conditions.
problem Learning mixtures of linear dynamical systems with better fit or understanding.
method Tensor decompositions to learn mixtures of linear dynamical systems.
result Algorithm succeeds without strong separation conditions and can compete with Bayes optimal clustering.
Study gaps and clusters in eigenvalues of magnetic Laplacian on manifolds.
problem Understanding gaps and clusters in eigenvalues of magnetic Laplacian on manifolds.
method Analyzes high tensor powers of Hermitian line bundles with non degenerate curvature, proving Riemann-Roch numbers for eigenvalue clusters and describing spectral projectors.
result Clusters and gaps in eigenvalues are described by Riemann-Roch numbers and have pointwise kernel descriptions.
The report analyzes Legendre decomposition for tensor data.
problem Finding effective lower dimensional representations of tensors.
method Theoretical analysis of dual parameters and dually flat manifold properties, followed by experimental verification and clustering.
result Parameters on submanifold cannot be directly used as low-rank representations.
New method estimates and completes tensors from ordinal data, improving accuracy and efficiency.
problem Estimating and completing tensors from incomplete, ordinal observations.
method Multi-linear cumulative link model with rank-constrained M-estimator.
result The proposed estimator achieves faster convergence and is minimax optimal.
Proposes a method to recover sparse tensors with covariate info.
problem Sparse tensor with high missing entries and many zeros.
method Covariate-assisted Sparse Tensor Completion (COSTCO) using latent components.
result 23% accuracy improvement over baseline in advertisement dataset.
In the present paper, we studied a Dynamic Stochastic Block Model (DSBM) under the assumptions that the connection probabilities, as functions of time, are smooth and that at most s nodes can switch their class memberships between two consecutive time points. We estimate the edge probability tensor by a kernel-type p…
Paper learns meaningful state and action representations from MDP trajectories.
problem Learning good state and action representations from MDP trajectories.
method Tensor decomposition, kernelization, importance sampling, low-Tucker-rank approximation.
result The learned state/action abstractions provide accurate approximations to latent block structures.
We consider the problem of decomposing a higher-order tensor with binary entries. Such data problems arise frequently in applications such as neuroimaging, recommendation system, topic modeling, and sensor network localization. We propose a multilinear Bernoulli model, develop a rank-constrained likelihood-based estima…
New model analyzes customer churn with tensor completion and binary data.
problem Analyzing the impact of interventions on customer churn.
method Tensorized latent factor block hazard model with 1-bit tensor completion.
result Effective categorization of interventions by similar impacts.
ALMA improves clustering of multilayer networks.
problem Clustering multilayer networks with distinct layers and communities.
method Alternating minimization algorithm (ALMA) for simultaneous layer partition and community estimation.
result ALMA achieves higher accuracy than TWIST in clustering multilayer networks.