VC-PCR improves prediction by clustering correlated variables.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study explores K-means clustering of variables and its relation to PCA.
With any non necessarily orientable unpunctured marked surface (S,M) we associate a commutative algebra, called quasi-cluster algebra, equipped with a distinguished set of generators, called quasi-cluster variables, in bijection with the set of arcs and one-sided simple closed curves in (S,M). Quasi-cluster variables a…
New method clusters variables using robust nodewise regression.
Model-based clustering defines population level clusters relative to a model that embeds notions of similarity. Algorithms tailored to such models yield estimated clusters with a clear statistical interpretation. We take this view here and introduce the class of G-block covariance models as a background model for varia…
Model-based clustering is a popular approach for clustering multivariate data which has seen applications in numerous fields. Nowadays, high-dimensional data are more and more common and the model-based clustering approach has adapted to deal with the increasing dimensionality. In particular, the development of variabl…
Study connects Morse theory with cluster variables for wall-crossing in Cerf diagrams.
Clustering analysis is one of the most widely used statistical tools in many emerging areas such as microarray data analysis. For microarray and other high-dimensional data, the presence of many noise variables may mask underlying clustering structures. Hence removing noise variables via variable selection is necessary…
The mixture models have become widely used in clustering, given its probabilistic framework in which its based, however, for modern databases that are characterized by their large size, these models behave disappointingly in setting out the model, making essential the selection of relevant variables for this type of cl…
In this paper, we introduce Adaptive Cluster Lasso(ACL) method for variable selection in high dimensional sparse regression models with strongly correlated variables. To handle correlated variables, the concept of clustering or grouping variables and then pursuing model fitting is widely accepted. When the dimension is…
A new method clusters mixed-type data efficiently.
We construct geometric realization for non-exceptional mutation-finite cluster algebras by extending the theory of Fomin and Thurston to skew-symmetrizable case. Cluster variables for these algebras are renormalized lambda lengths on certain hyperbolic orbifolds. We also compute growth rate of these cluster algebras, p…
New AI-block models for clustering high-dimensional variables based on maxima of random processes.
Discrete random variables are natural components of probabilistic clustering models. A number of VAE variants with discrete latent variables have been developed. Training such methods requires marginalizing over the discrete latent variables, causing training time complexity to be linear in the number clusters. By appl…
Clustering is an essential technique for discovering patterns in data. The steady increase in amount and complexity of data over the years led to improvements and development of new clustering algorithms. However, algorithms that can cluster data with mixed variable types (continuous and categorical) remain limited, de…
This paper addresses the problem of unsupervised clustering which remains one of the most fundamental challenges in machine learning and artificial intelligence. We propose the clustered generator model for clustering which contains both continuous and discrete latent variables. Discrete latent variables model the clus…
MultiDendrograms is a Java-written application that computes agglomerative hierarchical clusterings of data. Starting from a distances (or weights) matrix, MultiDendrograms is able to calculate its dendrograms using the most common agglomerative hierarchical clustering methods. The application implements a variable-gro…
In this paper, we present a new R package COREclust dedicated to the detection of representative variables in high dimensional spaces with a potentially limited number of observations. Variable sets detection is based on an original graph clustering strategy denoted CORE-clustering algorithm that detects CORE-clusters,…
A new method clusters mixed-type data tables effectively.
A new method selects important variables for clustering from dependency networks.
Unified framework for variable selection in model-based clustering with missing data.
This paper studies ordered weighted L1 (OWL) norm regularization for sparse estimation problems with strongly correlated variables. We prove sufficient conditions for clustering based on the correlation/colinearity of variables using the OWL norm, of which the so-called OSCAR is a particular case. Our results extend pr…
Proposes ARSK for robust and sparse clustering.
Generative Adversarial networks (GANs) have obtained remarkable success in many unsupervised learning tasks and unarguably, clustering is an important unsupervised learning problem. While one can potentially exploit the latent-space back-projection in GANs to cluster, we demonstrate that the cluster structure is not re…
Distributional (or distribution-valued) data are a new type of data arising from several sources and are considered as realizations of distributional variables. A new set of fuzzy c-means algorithms for data described by distributional variables is proposed. The algorithms use the Wasserstein distance between dist…
The task of clustering unlabeled time series and sequences entails a particular set of challenges, namely to adequately model temporal relations and variable sequence lengths. If these challenges are not properly handled, the resulting clusters might be of suboptimal quality. As a key solution, we present a joint clust…
We define transit clusters to simplify causal diagrams and preserve their essential properties.
New method clusters matrix-valued data by latent variables.
A new clustering model for mixed datasets combines continuous and non-continuous data.
Framework achieves fairness in predictions using partially known causal graph over clusters of variables.
Variable clustering is important for explanatory analysis. However, only few dedicated methods for variable clustering with the Gaussian graphical model have been proposed. Even more severe, small insignificant partial correlations due to noise can dramatically change the clustering result when evaluating for example w…
New solutions to 3D integrability equations using quantum cluster algebras.
VICatMix clusters categorical biomedical data efficiently and selects relevant variables.
Q-learning with cSMART data assesses cAI tailoring variables.
We advocate the use of cluster algebras and their y-variables in the study of hyperbolic 3-manifolds. We study hyperbolic structures on the mapping tori of pseudo-Anosov mapping classes of punctured surfaces, and show that cluster y-variables naturally give the solutions of the edge-gluing conditions of ideal tetrahedr…
The variability of the clusters generated by clustering techniques in the domain of latitude and longitude variables of fatal crash data are significantly unpredictable. This unpredictability, caused by the randomness of fatal crash incidents, reduces the accuracy of crash frequency (i.e., counts of fatal crashes per c…
In this paper, we show that Alexander polynomials for any 2-bridge knots are specializations of cluster variables. A key tool is an ancestral triangle which appeared in both quantum topology and hyperbolic geometry in different ways.
Fixed points found in cluster modular groups under specific conditions.
Finite mixture model is an important branch of clustering methods and can be applied on data sets with mixed types of variables. However, challenges exist in its applications. First, it typically relies on the EM algorithm which could be sensitive to the choice of initial values. Second, biomarkers subject to limits of…
In clustering we normally output one cluster variable for each datapoint. However it is not necessarily the case that there is only one way to partition a given dataset into cluster components. For example, one could cluster objects by their colour, or by their type. Different attributes form a hierarchy, and we could …
We propose a method for estimating coefficients in multivariate regression when there is a clustering structure to the response variables. The proposed method includes a fusion penalty, to shrink the difference in fitted values from responses in the same cluster, and an L1 penalty for simultaneous variable selection an…
We generalise surface cluster algebras to the case of infinite surfaces where the surface contains finitely many accumulation points of boundary marked points. To connect different triangulations of an infinite surface, we consider infinite mutation sequences. We show transitivity of infinite mutation sequences on tria…
Motivated by modern applications in which one constructs graphical models based on a very large number of features, this paper introduces a new class of cluster-based graphical models, in which variable clustering is applied as an initial step for reducing the dimension of the feature space. We employ model assisted cl…
New model clusters cells and individuals, revealing genetic influences on cell types.
Method selects valid IVs from a large set using clustering and test of overidentifying restrictions.
Cluster analysis methods seek to partition a data set into homogeneous subgroups. It is useful in a wide variety of applications, including document processing and modern genetics. Conventional clustering methods are unsupervised, meaning that there is no outcome variable nor is anything known about the relationship be…
Proposes a two-stage method for selecting correlated predictors in high-dimensional data.
We consider the task of estimating a Gaussian graphical model in the high-dimensional setting. The graphical lasso, which involves maximizing the Gaussian log likelihood subject to an l1 penalty, is a well-studied approach for this task. We begin by introducing a surprising connection between the graphical lasso and hi…