Proposes a hierarchical clustering method for positive and negative dissimilarities.
problem Clustering dissimilarities, especially positive and negative.
method Hierarchical correlation clustering followed by tree preserving embedding.
result Performance on various datasets.
Clusters of highly correlated stocks are identified for better asset selection.
problem Identifying a small set of stocks to approximate the diversification of the whole stock universe.
method Data-driven correlation blockmodel clustering approach.
result The algorithm effectively detects clusters of highly correlated stocks.
The paper tackles fair correlation clustering with new algorithms and analysis.
problem Fair variants of correlation clustering under various constraints.
method Introducing a novel combinatorial optimization problem for fairlet decomposition.
result Approximation algorithms for fair correlation clustering under multiple fairness constraints.
CCP clusters correlated features and projects them to 1D for efficient dimensionality reduction.
problem Efficiency in handling large datasets with high intrinsic dimensions.
method CCP partitions features into correlated clusters and projects them to 1D based on sample correlations.
result CCP achieves efficient dimensionality reduction without matrix diagonalization.
CSTS benchmarks time series clustering by evaluating correlation structures.
problem Lack of validated ground truth for objectively assessing clustering quality.
method Synthetic benchmark CSTS for evaluating correlation structures in multivariate time series data.
result CSTS enables precise diagnosis of methodological limitations in correlation-based time series clustering.
A new algorithm removes unexpected correlations in biased data for better clustering.
problem Clustering with selection bias in data.
method Decorrelation regularized K-Means (DCKM) algorithm.
result DCKM achieves significant performance gains on real-world datasets.
Active learning optimizes correlation clustering by querying the most informative pairwise comparisons.
problem Efficiently clustering data with limited pairwise similarity information.
method Developed principled active learning approach using information-theoretic acquisition functions.
result Significantly outperforms existing baselines in clustering accuracy and query efficiency.
DynMSA detects market clusters for better portfolio allocation.
problem Identifying stable market clusters for effective portfolio management.
method Combining Random Matrix Theory with modularity optimization and spectral clustering.
result DynMSA outperforms baseline models in intra- and inter-cluster correlation differences.
The paper tackles fair correlation clustering with fairness constraints.
problem Minimizing disagreements while adhering to fairness constraints for clustering.
method Two variants of fairness constraints are considered: equal distribution and relative bounds. Approximation algorithms are developed for these constraints.
result Approximation algorithms for fair correlation clustering with theoretical guarantees and empirical validation.
This paper analyzes correlations in patterns of trading of different members of the London Stock Exchange. The collection of strategies associated with a member institution is defined by the sequence of signs of net volume traded by that institution in hour intervals. Using several methods we show that there are signif…
Paper uses HPCA for better stock correlation modeling.
problem Challenges in modeling cross-sectional correlations between thousands of stocks.
method Hierarchical Principal Component Analysis (HPCA) and statistical clustering.
result HPCA provides better cross-sectional correlations than classic PCA.
Clusters of financial market states identified over 2006-2019.
problem Understanding the statistical properties of financial markets.
method Clustering analysis of correlation matrices constructed from sliding epochs.
result Financial markets can be classified into distinct states with transitions indicating precursors to catastrophic events.
The study uses DCC for financial market analysis, revealing hidden correlations.
problem Identifying hidden nonlinear correlations in financial markets.
method Agglomerative hierarchical clustering with distance correlation coefficient.
result DCC reveals more information than Pearson correlation for financial data.
New RDPC dissimilarity measure improves time series clustering.
problem Improving time series clustering methods for diverse data.
method Combining weighted Pearson correlation with largest element-wise differences.
result RDPC outperforms existing methods in complex datasets.
Paper develops active learning for clustering unknown pairwise similarities.
problem Learning positive and negative pairwise similarities efficiently.
method Generic active learning framework for correlation clustering.
result Demonstrates effectiveness of query strategies in clustering.
This paper optimizes cryptocurrency portfolios by clustering price correlations and improving risk-return profiles.
problem Volatility and regulatory uncertainty in cryptocurrency markets make portfolio construction challenging.
method The paper combines network analysis, price forecasting, and portfolio theory to identify stable groups of correlated cryptocurrencies.
result Predictive consensus-clustering portfolios maintain positive and stable performance up to a 14-day horizon, with favourable gain-loss asymmetry and tighter tail-risk control.
Cluster GARCH model improves multivariate GARCH for high-dimensional asset returns.
problem Modeling high-dimensional asset returns with flexible tail dependencies and cluster structures.
method Introduced a novel multivariate GARCH model with flexible convolution-t distributions, tractable likelihood and derivatives for dynamic correlation structure.
result Cluster GARCH model outperforms existing models in daily returns of 100 assets, both in-sample and out-of-sample.
A measure called relative cluster entropy distinguishes between correlated and uncorrelated sequences.
problem Distinguishing between sequences with different correlation degrees.
method Minimum relative entropy principle applied to cluster partitions of power-law correlated sequences.
result Optimal Hurst exponents are selected for market price series, indicating non-markovianity.
We describe a new optimization scheme for finding high-quality correlation clusterings in planar graphs that uses weighted perfect matching as a subroutine. Our method provides lower-bounds on the energy of the optimal correlation clustering that are typically fast to compute and tight in practice. We demonstrate our a…
Quantum GBS boosts asset clustering for robust statistical arbitrage portfolios.
problem Identifying co-moving assets from correlation matrices for statistical arbitrage.
method Mapping S&P 500 correlation data to GBS-compatible adjacency matrices, benchmarking classical and quantum clustering algorithms.
result Quantum GBS generates superior alpha during high volatility periods, persisting under low-loss conditions.
In correlation clustering, we are given n objects together with a binary similarity score between each pair of them. The goal is to partition the objects into clusters so to minimise the disagreements with the scores. In this work we investigate correlation clustering as an active learning problem: each similarity sc…
Cluster stability selection improves feature selection in correlated data.
problem Feature selection stability in correlated data.
method Cluster stability selection exploiting known cluster structure.
result Better predictive performance than lasso alone and stability selection.
This paper studies ordered weighted L1 (OWL) norm regularization for sparse estimation problems with strongly correlated variables. We prove sufficient conditions for clustering based on the correlation/colinearity of variables using the OWL norm, of which the so-called OSCAR is a particular case. Our results extend pr…
Data de-duplication is the task of detecting multiple records that correspond to the same real-world entity in a database. In this work, we view de-duplication as a clustering problem where the goal is to put records corresponding to the same physical entity in the same cluster and putting records corresponding to diff…
This article investigates the correlation structure of the global crude oil market using the daily returns of 71 oil price time series across the world from 1992 to 2012. We identify from the correlation matrix six clusters of time series exhibiting evident geographical traits, which supports Weiner's (1991) regionaliz…
End-to-end deep learning for multi-view clustering improves accuracy across various data types.
problem Limited multi-view clustering methods for general data types and suboptimal two-stage process.
method Permutation-based canonical correlation objective for fused representations; pseudo-labels for clustering; theoretical error bound.
result Proposed model provides meaningful fused representations and effective clustering across multiple views.
Given a similarity graph between items, correlation clustering (CC) groups similar items together and dissimilar ones apart. One of the most popular CC algorithms is KwikCluster: an algorithm that serially clusters neighborhoods of vertices, and obtains a 3-approximation ratio. Unfortunately, KwikCluster in practice re…
VC-PCR improves prediction by clustering correlated variables.
problem Decreased prediction accuracy due to cluster structure in predictor variables.
method Supervised variable selection and clustering to integrate cluster information into a sparse modeling process.
result VC-PCR achieves better prediction, variable selection, and clustering performance.
Motivated by social balance theory, we develop a theory of link classification in signed networks using the correlation clustering index as measure of label regularity. We derive learning bounds in terms of correlation clustering within three fundamental transductive learning settings: online, batch and active. Our mai…
Spatially relaxed inference tackles high-dimensional linear models with correlated covariates.
problem Accurate inference is challenging in high-dimensional settings with spatially correlated covariates.
method Proposes ensembled clustered inference algorithms that control the δ-FWER under standard assumptions. result Ensembled clustered inference algorithms control the δ-FWER and achieve decent power. In this work, the possibility of clustering correlated random variables was examined, both because of their mutual similarity and because of their similarity to the principal components. The k-means algorithm and spectral algorithms were used for clustering. For spectral methods, the similarity matrix was both the matr…
New method clusters evolving networks using spatio-temporal graph Laplacian.
problem Clustering communities in time-varying graphs.
method Extends spectral clustering to dynamic graphs using CCA and spatio-temporal graph Laplacian.
result The spatio-temporal graph Laplacian clearly interprets cluster evolution over time.
FASC clusters data with latent factors, improving on naive methods.
problem Clustering high-dimensional data with correlated variables.
method Factor Adjusted Spectral Clustering (FASC) algorithm.
result FASC achieves an exponentially low mislabeling rate under general assumptions.
This study uses moving average cluster entropy to analyze financial market dynamics.
problem Understanding long-range dependence in financial markets.
method Moving average cluster entropy approach applied to ARFIMA and FBM processes.
result Long-range positive correlation in financial markets is linked to the cluster entropy behavior.
The cluster analysis methods are used in order to perform a comparative study of 15 EU countries in relation with the fluctuations of some basic macroeconomic indicators. The statistical distances between countries are calculated for various moving time windows, and the time variation of the mean statistical distance i…
In this paper, we introduce Adaptive Cluster Lasso(ACL) method for variable selection in high dimensional sparse regression models with strongly correlated variables. To handle correlated variables, the concept of clustering or grouping variables and then pursuing model fitting is widely accepted. When the dimension is…
We study the structure of locational marginal prices in day-ahead and real-time wholesale electricity markets. In particular, we consider the case of two North American markets and show that the price correlations contain information on the locational structure of the grid. We study various clustering methods and intro…
Researchers have used from 30 days to several years of daily returns as source data for clustering financial time series based on their correlations. This paper sets up a statistical framework to study the validity of such practices. We first show that clustering correlated random variables from their observed values i…
Geometric QHD tests improve hub detection in correlated data.
problem Detecting hubs in correlated data with evolving correlations.
method Geometric QHD tests combining QCD and QHD, clustering.
result Improved hub detection in correlated data.
Clusters cryptocurrency market states via cross correlation analysis.
problem Analyse cryptocurrency market dynamics.
method Cross correlation structure analysis over 5 years.
result Cryptocurrency market clusters into 4 states.
Variable clustering is important for explanatory analysis. However, only few dedicated methods for variable clustering with the Gaussian graphical model have been proposed. Even more severe, small insignificant partial correlations due to noise can dramatically change the clustering result when evaluating for example w…
This paper presents a novel application of a clustering algorithm developed for constructing a phylogenetic network to the correlation matrix for 126 stocks listed on the Shanghai A Stock Market. We show that by visualizing the correlation matrix using a Neighbor-Net network and using the circular ordering produced dur…
A new method captures higher-order interactions in data clusters.
problem Accurately characterizing complex higher-order variable interactions.
method Local Correlation Explanation (CorEx) method: clustering and total correlation.
result Captures higher-order interactions at a local scale.
Develops a new random forest method for clustered data with improved prediction and inference.
problem Improving prediction and inference accuracy for clustered data with within-cluster dependence.
method Clustered Random Forests, using weighted least squares estimators for leaf predictions.
result Optimal prediction and inference weights vary under covariate shift, necessitating user-chosen weights.
We review the state of the art of clustering financial time series and the study of their correlations alongside other interaction networks. The aim of this review is to gather in one place the relevant material from different fields, e.g. machine learning, information geometry, econophysics, statistical physics, econo…
The paper explores how macroeconomic variables' correlation structure changes over time and under different scenarios.
problem Understanding the changing correlation structure of macroeconomic variables.
method The paper uses a principal component based algorithm to perform unsupervised clustering on macroeconomic variables.
result The correlation structure of macroeconomic variables changes significantly during financial crises and under hypothetical scenarios.
Cluster jackknife improves inference for staggered DID methods.
problem Over-rejection of CSDID in small clusters or treated clusters.
method Cluster jackknife for CSDID inference.
result Cluster jackknife greatly improves inference for CSDID.
New method separates market motion from stock correlations.
problem Understanding the dynamics of stock correlations relative to market motion.
method Cluster reduced-rank correlation matrices by subtracting the largest eigenvalue.
result Extracted market states are quasi-stationary over long periods.