Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

316192122 · Jun 202019922001200920182026
48 results for Consensus Clustering

IMPACC improves consensus clustering for bioinformatics data.

problem Consensus clustering's inefficiency and lack of interpretability for large-scale data.
method Ensemble minipatch co-occurrences, adaptive sampling of observations and features.
result Significantly improved accuracy and interpretability with substantial computational savings.

We use a cluster ensemble to determine the number of clusters, k, in a group of data. A consensus similarity matrix is formed from the ensemble using multiple algorithms and several values for k. A random walk is induced on the graph defined by the consensus matrix and the eigenvalues of the associated transition proba…

2014-08-05abs ↗pdf ↗

A novel framework for consensus clustering is presented which has the ability to determine both the number of clusters and a final solution using multiple algorithms. A consensus similarity matrix is formed from an ensemble using multiple algorithms and several values for k. A variety of dimension reduction techniques …

2014-08-05abs ↗pdf ↗

Theoretical extension of Condorcet's Jury Theorem for consensus clustering.

problem Quality of consensus clustering depends on the diversity of sample partitions.
method Extending Condorcet's Jury Theorem to mean partition approach under specific assumptions.
result Limiting the diversity of mean partitions is necessary for controlling the quality of consensus clustering.

The paper proposes a method to evaluate clustering models using ensemble techniques.

problem Lack of robust cluster validity scores for unsupervised learning.
method Cluster ensemble aggregation techniques and normalized mutual information.
result The method can highlight standout clustering and hyperparameter configurations in an ensemble.

A new framework for clustering high-dimensional data using vertical shards.

problem Clustering high-dimensional data with the curse of dimensionality.
method Vertical Consensus Inference (VCI) that splits data into vertical shards for posterior inference.
result VCI can approximate inference on random partitions for high-dimensional data.

The study shows mean partitions are consistent and asymptotically normal.

problem Lack of knowledge about the consistency of mean partitions in consensus clustering.
method Represented partitions as points in orbit space, used Fréchet means and stochastic programming, and analyzed continuous extensions of cluster criteria.
result Mean partitions are consistent and asymptotically normal under normal assumptions.

The task of clustering a set of objects based on multiple sources of data arises in several modern applications. We propose an integrative statistical model that permits a separate clustering of the objects for each data source. These separate clusterings adhere loosely to an overall consensus clustering, and hence the…

2013-02-28abs ↗pdf ↗

Clustering ensemble is one of the most recent advances in unsupervised learning. It aims to combine the clustering results obtained using different algorithms or from different runs of the same clustering algorithm for the same data set, this is accomplished using on a consensus function, the efficiency and accuracy of…

2012-08-20abs ↗pdf ↗

We consider grouping as a general characterization for problems such as clustering, community detection in networks, and multiple parametric model estimation. We are interested in merging solutions from different grouping algorithms, distilling all their good qualities into a consensus solution. In this paper, we propo…

2014-04-30abs ↗pdf ↗

A new clustering framework optimizes customer search data for personalized travel recommendations.

problem Personalized travel recommendations based on customer search data.
method Multi-objective optimization-based clustering ensemble framework.
result Optimizes diversity in clustering ensemble search space and automatically determines the number of clusters.

A method for learning embeddings from multi-view data using Gromov-Wasserstein.

problem Challenges in learning low-dimensional representations from multi-view relational data with differing geometries.
method Bary-GWMDS and Mean-GWMDS-C, Gromov-Wasserstein-based methods operating on distance matrices.
result Stable and geometrically meaningful embeddings learned from synthetic and real-world datasets.

This paper improves multi-view spectral clustering by preserving local manifold structures.

problem Agreeing on data object grouping across multiple views with noisy data.
method Iterative Low-Rank based Structured Optimization method using multi-graph Laplacians.
result Improved spectral clustering by preserving local manifold structures.

FedCBO solves clustered federated learning by optimizing groups of users without knowing their structure.

problem Training models for multiple users with privacy and communication constraints, especially in clustered settings.
method FedCBO uses a particle system approach inspired by consensus-based optimization to train models for each user group.
result FedCBO outperforms other methods in training models for clustered federated learning.

This paper optimizes cryptocurrency portfolios by clustering price correlations and improving risk-return profiles.

problem Volatility and regulatory uncertainty in cryptocurrency markets make portfolio construction challenging.
method The paper combines network analysis, price forecasting, and portfolio theory to identify stable groups of correlated cryptocurrencies.
result Predictive consensus-clustering portfolios maintain positive and stable performance up to a 14-day horizon, with favourable gain-loss asymmetry and tighter tail-risk control.

A novel multi-clustering method based on boosting improves hierarchical clustering quality.

problem Improving hierarchical clustering quality in flat clustering problems.
method A boosting iteration with weighted random sampling of elements from the original dataset, followed by hierarchical clustering on each subsample and consensus combination.
result The proposed method provides superior quality solutions compared to standard hierarchical clustering methods.

Proposes a method to learn a low-rank kernel matrix for graph-based clustering.

problem Challenges in learning an optimal kernel matrix for graph-based clustering.
method Unified framework for graph construction and kernel learning, focusing on a low-rank kernel matrix.
result Efficacy of the proposed method validated through extensive experiments.

Enhanced ensemble clustering via fast propagation of cluster-wise similarities.

problem Challenges in exploring higher-level granularity and multi-scale indirect relationships in ensemble clustering.
method A novel ensemble clustering approach based on fast propagation of cluster-wise similarities via random walks.
result Proposes a new cluster-wise similarity matrix and consensus functions to achieve enhanced co-association and consensus clustering.

This paper addresses clustering with missing data using Rubin's rules.

problem How to pool partitions and assess instability when data are incomplete after multiple imputation.
method Consensus clustering and bootstrap theory are used to address the problem.
result New rules for pooling partitions and assessing instability are proposed and validated.

Proposes a new ensemble clustering method to handle uncertain links and incorporate global information.

problem Uncertain links and lack of global information in ensemble clustering.
method Sparse graph representation, elite neighbor selection, random walk process, and dense similarity measure.
result Significantly better consensus results compared to using all graph links.

A new MKL framework improves graph-based clustering and semi-supervised classification.

problem MKL methods often fail to improve performance over single kernels.
method Proposes a new MKL framework based on consensus kernels and automatic weight assignment.
result The proposed method outperforms existing MKL methods on multiple benchmark datasets.

A new method clusters multi-view data by sharing a common trace-norm of coefficient matrices.

problem Insufficient exploitation of multi-view data due to uniform coefficient matrices.
method Imposes bilinear factorization with orthonormality and low-rank constraints on coefficient matrices.
result The proposed CBF-MSC method effectively clusters multi-view data more comprehensively.

FCMSC combines multi-view data through feature concatenation for improved clustering.

problem Clustering multi-view data with diverse and sometimes incompatible views.
method FCMSC concatenates multi-view data, integrates l2,1l_{2,1}-norm, and uses graph regularization to explore consensus and complementary information.
result FCMSC outperforms state-of-the-art multi-view clustering methods on six real-world datasets.

A new model clusters multi-faceted data with uncertainty quantification.

problem Uncertainty in multi-view clustering of high-dimensional data.
method Approximate Bayes approach, treating similarity matrices as rough estimates, refining with low-rank matrix.
result Each simplex coordinate encodes cluster assignment uncertainty.

Stochastic Sparse Subspace Clustering improves subspace clustering by reducing over-segmentation through dropout.

problem Over-segmentation in subspace clustering.
method Introducing dropout regularization to enforce denser connections between points from the same subspace.
result Stochastic Sparse Subspace Clustering effectively handles large datasets and reduces over-segmentation.

The paper proposes a consensus algorithm to improve deep neural network interpretability and accuracy in mortality prediction.

problem The black-box nature and overgeneralization of deep neural networks in healthcare applications.
method An (nn, kk) consensus algorithm that is insensitive to adversarial examples and can reliably reject out-of-distribution samples.
result The consensus algorithm improves both prediction accuracy and interpretability of deep neural network models in mortality prediction.

A new procedure aggregates models to predict data from multiple clusters.

problem Predicting data from multiple clusters with different underlying models.
method Three-step procedure: clustering, model fitting, and aggregation.
result The method outperforms existing models in various prediction problems.

DEMVC improves multi-view clustering with collaborative training and deep autoencoders.

problem Existing multi-view clustering methods have high computation and space complexities or lack representation capability.
method DEMVC learns embedded representations of multiple views individually using deep autoencoders and collaboratively trains all views.
result DEMVC achieves significant improvements over state-of-the-art methods on multi-view datasets.

EAP clusters evolving data, promoting temporal smoothness and automatic cluster tracking.

problem Clustering time-evolving data with temporal smoothness and automatic cluster identification.
method Evolutionary Affinity Propagation (EAP) on a factor graph exchanging messages between adjacent data snapshots.
result EAP clusters data with temporal smoothness and automatically tracks clusters, outperforming existing methods.

Anomaly detection identifies unusual malaria transmission patterns in Ghana.

problem Identifying atypical malaria transmission patterns in Ghana's spatiotemporal surveillance data.
method Consensus-based anomaly detection framework applied to monthly malaria surveillance data.
result High-burden areas are not necessarily those with the most frequent anomalous transmission.

A new method clusters data from multiple sources using a mixture of multilayer SBMs.

problem Aggregating multiple clustering results from different data sources.
method Uses a mixture of multilayer Stochastic Block Models (SBM) to group co-membership matrices.
result Identifies and clusters observations based on their specificities within components.