Proportional centroid clustering aims to fairly group points without prior protected subsets.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The method integrates survival constraints into NMF for identifying survival-associated gene clusters.
Study examines how cluster number affects short-text clustering, introducing a stability metric.
Proposes FMC for fair clustering with independent parameters.
Understanding proper distance measures between distributions is at the core of several learning tasks such as generative models, domain adaptation, clustering, etc. In this work, we focus on mixture distributions that arise naturally in several application domains where the data contains different sub-populations. For …
Study compares clustering methods for mixed-type data.
A fair clustering method for multiple sensitive attributes is proposed.
A neural-network model clusters subjects based on their lifetime distributions.
The joint optimization of representation learning and clustering in the embedding space has experienced a breakthrough in recent years. In spite of the advance, clustering with representation learning has been limited to flat-level categories, which often involves cohesive clustering with a focus on instance relations.…
FCA improves fair clustering by optimizing utility and fairness.
The paper investigates how irrelevant features affect clustering performance.
FBC clusters data fairly without needing cluster count.
The binary symmetric stochastic block model deals with a random graph of vertices partitioned into two equal-sized clusters, such that each pair of vertices is connected independently with probability within clusters and across clusters. In the asymptotic regime of and for fixe…
This paper tackles fair clustering with multiple types and provides scalable coreset solutions.
Paper detects gradual changes in cluster structure using MC fusion.
FONT clusters patients across health systems with privacy and efficiency.
We consider the problem of community detection or clustering in the labeled Stochastic Block Model (LSBM) with a finite number of clusters of sizes linearly growing with the global population of items . Every pair of items is labeled independently at random, and label appears with probability $p(i,j,\ell)…
The paper introduces group-representative clustering to ensure fair representation of different groups in clusters.
Resolving a conjecture of Abbe, Bandeira and Hall, the authors have recently shown that the semidefinite programming (SDP) relaxation of the maximum likelihood estimator achieves the sharp threshold for exactly recovering the community structure under the binary stochastic block model of two equal-sized clusters. The s…
Preschool attendance correlates with lower developmental vulnerabilities in Queensland, Australia.
Random column sampling is not guaranteed to yield data sketches that preserve the underlying structures of the data and may not sample sufficiently from less-populated data clusters. Also, adaptive sampling can often provide accurate low rank approximations, yet may fall short of producing descriptive data sketches, es…
Mixtures of multivariate contaminated shifted asymmetric Laplace distributions are developed for handling asymmetric clusters in the presence of outliers (also referred to as bad points herein). In addition to the parameters of the related non-contaminated mixture, for each (asymmetric) cluster, our model has one param…
Clustering high-dimensional data often requires some form of dimensionality reduction, where clustered variables are separated from "noise-looking" variables. We cast this problem as finding a low-dimensional projection of the data which is well-clustered. This yields a one-dimensional projection in the simplest situat…
New method for clustering tasks with heterogeneous data.
Modeling price clustering in financial markets using discrete distributions.
Under the framework of spectral clustering, the key of subspace clustering is building a similarity graph which describes the neighborhood relations among data points. Some recent works build the graph using sparse, low-rank, and -norm-based representation, and have achieved state-of-the-art performance. Howeve…
New algorithm speeds up fair clustering by 12x.
Exact hierarchical clustering algorithms for data analysis.
The paper analyzes how clustering sensitive data can improve model generalization without revealing individual information.
We study the optimal design problems where the goal is to choose a set of linear measurements to obtain the most accurate estimate of an unknown vector in dimensions. We study the -optimal design variant where the objective is to minimize the average variance of the error in the maximum likelihood estimate of th…
New system constructs cell-type taxonomy across multiple samples.
The in-game economies of massively multi-player online games (MMOGs) are complex systems that have to be carefully designed and managed. This paper presents the results of an analysis of auction house data from the MMOG Glitch, across a 14 month time period, the entire lifetime of the game. The data comprise almost 3 m…
A new LDA model with covariates for mixed-membership clusters.
Bayesian SAE model with spectral clustering and uncertainty quantification.
Paper addresses data heterogeneity in federated learning for CoxPH models in healthcare.
A new method for efficient inference and model selection in SBMs using OT.
In many situations where the interest lies in identifying clusters one might expect that not all available variables carry information about these groups. Furthermore, data quality (e.g. outliers or missing entries) might present a serious and sometimes hard-to-assess problem for large and complex datasets. In this pap…
Multilayer bootstrap network builds a gradually narrowed multilayer nonlinear network from bottom up for unsupervised nonlinear dimensionality reduction. Each layer of the network is a nonparametric density estimator. It consists of a group of k-centroids clusterings. Each clustering randomly selects data points with r…
The Lloyd-Max algorithm is a classical approach to perform K-means clustering. Unfortunately, its cost becomes prohibitive as the training dataset grows large. We propose a compressive version of K-means (CKM), that estimates cluster centers from a sketch, i.e. from a drastically compressed representation of the traini…
In this paper, we explore the relationship between one of the most elementary and important properties of graphs, the presence and relative frequency of triangles, and a combinatorial notion of Ricci curvature. We employ a definition of generalized Ricci curvature proposed by Ollivier in a general framework of Markov p…
We propose Rademacher complexity bounds for multiclass classifiers trained with a two-step semi-supervised model. In the first step, the algorithm partitions the partially labeled data and then identifies dense clusters containing predominant classes using the labeled training examples such that the proportion of t…
DC-NAS improves neural architecture search by clustering and evaluating sub-networks.
Unlike common cancers, such as those of the prostate and breast, tumor grading in rare cancers is difficult and largely undefined because of small sample sizes, the sheer volume of time needed to undertake on such a task, and the inherent difficulty of extracting human-observed patterns. One of the most challenging exa…
Several algorithms have been proposed to filter information on a complete graph of correlations across stocks to build a stock-correlation network. Among them the planar maximally filtered graph (PMFG) algorithm uses edges to build a graph whose features include a high frequency of small cliques and a good clust…
Spectral dimensionality reduction methods enable linear separations of complex data with high-dimensional features in a reduced space. However, these methods do not always give the desired results due to irregularities or uncertainties of the data. Thus, we consider aggressively modifying the scales of the features to …
Improved neural network model for predicting latent budgets in compositional data.
A new framework improves fairness in clustering and Wasserstein Barycenter problems.
The use of mutual information as a similarity measure in agglomerative hierarchical clustering (AHC) raises an important issue: some correction needs to be applied for the dimensionality of variables. In this work, we formulate the decision of merging dependent multivariate normal variables in an AHC procedure as a Bay…