Extends MSC for triclustering tensors, using DBSCAN to find clusters.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New method for multiway clustering of 3rd order tensors.
New method clusters tensors with heteroskedastic noise.
Dynamic tensor data are becoming prevalent in numerous applications. Existing tensor clustering methods either fail to account for the dynamic nature of the data, or are inapplicable to a general-order tensor. Also there is often a gap between statistical guarantee and computational efficiency for existing tensor clust…
New method for triclustering with reduced arbitrariness.
Cluster analysis is a fundamental tool for pattern discovery of complex heterogeneous data. Prevalent clustering methods mainly focus on vector or matrix-variate data and are not applicable to general-order tensors, which arise frequently in modern scientific and business applications. Moreover, there is a gap between …
Paper develops a method to robustly cluster tensors with outliers.
Proposes a new model for clustering passenger trips considering hierarchical and multi-dimensional data.
Paper proposes a tensor model for clustering noisy multi-view data.
Paper proposes efficient methods for high-order clustering in tensor block models.
Paper explores limits of high-order clustering with planted structures.
MultiwayPAM clusters LLM-as-a-Judge scores to reveal evaluator bias.
Develops a new tensor model for clustering with degree correction.
Paper proposes a new method to improve clustering ensemble performance.
Paper improves MVSC using tensor low-rank modeling.
Proposes a new model for clustering passenger trajectories with graphs.
Study quantifies performance gap between tensor and matrix-based approaches in nested matrix-tensor model.
Develops a tensor mixture model for high-dimensional data clustering.
A tensor provides a concise way to codify the interdependence of complex data. Treating a tensor as a d-way array, each entry records the interaction between the different indices. Clustering provides a way to parse the complexity of the data into more readily understandable information. Clustering methods are heavily …
In this paper we present a method for the unsupervised clustering of high-dimensional binary data, with a special focus on electronic healthcare records. We present a robust and efficient heuristic to face this problem using tensor decomposition. We present the reasons why this approach is preferable for tasks such as …
Co-Clustering, the problem of simultaneously identifying clusters across multiple aspects of a data set, is a natural generalization of clustering to higher-order structured data. Recent convex formulations of bi-clustering and tensor co-clustering, which shrink estimated centroids together using a convex fusion penalt…
TACE unifies scalar and tensorial modeling in Cartesian space for accurate, stable, and efficient atomistic predictions.
We consider the problem of identifying multiway block structure from a large noisy tensor. Such problems arise frequently in applications such as genomics, recommendation system, topic modeling, and sensor network localization. We propose a tensor block model, develop a unified least-square estimation, and obtain the t…
The performance of most the clustering methods hinges on the used pairwise affinity, which is usually denoted by a similarity matrix. However, the pairwise similarity is notoriously known for its vulnerability of noise contamination or the imbalance in samples or features, and thus hinders accurate clustering. To tackl…
Proposes a tensor Laplacian-based method for better subspace clustering of non-uniformly distributed data.
In many real-world applications, data are often unlabeled and comprised of different representations/views which often provide information complementary to each other. Although several multi-view clustering methods have been proposed, most of them routinely assume one weight for one view of features, and thus inter-vie…
Exact partitioning of high-order planted models achieved through convex optimization.
Physical activity levels are an important predictor of cardiovascular health and increasingly being measured by sensors, like accelerometers. Accelerometers produce rich multivariate data that can inform important clinical decisions related to individual patients and public health. The CHAMPION study, a study of youth …
Objective Function Mismatch (OFM) occurs when the optimization of one objective has a negative impact on the optimization of another objective. In this work we study OFM in deep clustering, and find that the popular autoencoder-based approach to deep clustering can lead to both reduced clustering performance, and a sig…
Paper optimizes clustering for multi-layer networks and discrete mixtures.
Dimensionality reduction techniques play an essential role in data analytics, signal processing and machine learning. Dimensionality reduction is usually performed in a preprocessing stage that is separate from subsequent data analysis, such as clustering or classification. Finding reduced-dimension representations tha…
Paper develops a method for causal representation learning from irregular tensors.
We introduce a tensor-based clustering method to extract sparse, low-dimensional structure from high-dimensional, multi-indexed datasets. This framework is designed to enable detection of clusters of data in the presence of structural requirements which we encode as algebraic constraints in a linear program. Our cluste…
Matrix factorizations and their extensions to tensor factorizations and decompositions have become prominent techniques for linear and multilinear blind source separation (BSS), especially multiway Independent Component Analysis (ICA), NonnegativeMatrix and Tensor Factorization (NMF/NTF), Smooth Component Analysis (Smo…
Paper perfect clusters sparse, diverse multilayer networks.
MSFA clusters high-dimensional spatial data using spline-based covariance structures.
For a symplectic manifold with quantizing line bundle, a choice of almost complex structure determines a Laplacian acting on tensor powers of the bundle. For high tensor powers Guillemin-Uribe showed that there is a well-defined cluster of low-lying eigenvalues, whose distribution is described by a spectral density fun…
Study graph-based algorithms for multi-manifold clustering with sufficient conditions.
New approach learns mixtures of linear dynamical systems without separation conditions.
Study gaps and clusters in eigenvalues of magnetic Laplacian on manifolds.
The report analyzes Legendre decomposition for tensor data.
Proposes a method to recover sparse tensors with covariate info.
In the present paper, we studied a Dynamic Stochastic Block Model (DSBM) under the assumptions that the connection probabilities, as functions of time, are smooth and that at most nodes can switch their class memberships between two consecutive time points. We estimate the edge probability tensor by a kernel-type p…
Paper learns meaningful state and action representations from MDP trajectories.
Higher-order tensors arise frequently in applications such as neuroimaging, recommendation system, social network analysis, and psychological studies. We consider the problem of low-rank tensor estimation from possibly incomplete, ordinal-valued observations. Two related problems are studied, one on tensor denoising an…
We consider the problem of decomposing a higher-order tensor with binary entries. Such data problems arise frequently in applications such as neuroimaging, recommendation system, topic modeling, and sensor network localization. We propose a multilinear Bernoulli model, develop a rank-constrained likelihood-based estima…
New model analyzes customer churn with tensor completion and binary data.
ALMA improves clustering of multilayer networks.