Ideas from the image processing literature have recently motivated a new set of clustering algorithms that rely on the concept of total variation. While these algorithms perform well for bi-partitioning tasks, their recursive extensions yield unimpressive results for multiclass clustering tasks. This paper presents a g…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A main task in data analysis is to organize data points into coherent groups or clusters. The stochastic block model is a probabilistic model for the cluster structure. This model prescribes different probabilities for the presence of edges within a cluster and between different clusters. We assume that the cluster ass…
Network Lasso clusters sparse graph clusters efficiently.
Data clustering is a fundamental problem with a wide range of applications. Standard methods, eg the -means method, usually require solving a non-convex optimization problem. Recently, total variation based convex relaxation to the -means model has emerged as an attractive alternative for data clustering. However…
While it is believed that denoising is not always necessary in many big data applications, we show in this paper that denoising is helpful in urban traffic analysis by applying the method of bounded total variation denoising to the urban road traffic prediction and clustering problem. We propose two easy-to-implement m…
Study clusters distributions with known or unknown clusters using distribution testing.
We consider the initial situation where a dataset has been over-partitioned into clusters and seek a domain independent way to merge those initial clusters. We identify the total variation distance (TVD) as suitable for this goal. By exploiting the relation of the TVD to the Bayes accuracy we show how neural networ…
New framework estimates staged tree models using hierarchical clustering on the probability simplex.
Robustly clusters mixtures of Gaussians even with outliers.
Generalizes underlap coefficient for multivariate group separation.
Mixture models and topic models generate each observation from a single cluster, but standard variational posteriors for each observation assign positive probability to all possible clusters. This requires dense storage and runtime costs that scale with the total number of clusters, even though typically only a few clu…
Unified federated learning via GTV minimization.
We propose and analyze a method for semi-supervised learning from partially-labeled network-structured data. Our approach is based on a graph signal recovery interpretation under a clustering hypothesis that labels of data points belonging to the same well-connected subset (cluster) are similar valued. This lends natur…
We consider point clouds obtained as random samples of a measure on a Euclidean domain. A graph representing the point cloud is obtained by assigning weights to edges based on the distance between the points they connect. Our goal is to develop mathematical tools needed to study the consistency, as the number of availa…
We present a convex approach to probabilistic segmentation and modeling of time series data. Our approach builds upon recent advances in multivariate total variation regularization, and seeks to learn a separate set of parameters for the distribution over the observations at each time point, but with an additional pena…
New method clusters non-spherical Gaussian mixtures with fewer samples and time.
This paper analyzes the connection between innovation activities of companies -- implemented before crisis -- and their performance -- measured at time of crisis. The companies listed in the STAR Market Segment of the Italian Stock Exchange are analyzed. Innovation is measured through the level of investments in total …
New model clusters cells and individuals, revealing genetic influences on cell types.
The paper tackles fair correlation clustering with fairness constraints.
New method clusters matrix-variate data with outliers.
In a variety of research areas, the weighted bag of vectors and the histogram are widely used descriptors for complex objects. Both can be expressed as discrete distributions. D2-clustering pursues the minimum total within-cluster variation for a set of discrete distributions subject to the Kantorovich-Wasserstein metr…
The paper studies curves in Riemannian manifolds using total variation flow.
The paper analyzes how companies' investments before crises affect their performance after crises.
Optimal pre-processing reduces disparate impact by minimizing total variation distance.
New clustering algorithm for time series data using RNN and variational Bayes.
NeuralFLoC unifies registration and clustering of functional data, overcoming phase variation challenges.
The paper connects neural collapse and low-rank bias in networks with L2 regularization.
Proposes variational Wasserstein barycenters for geometric clustering.
New method learns low-dimensional representations of nonlinear time series without supervision.
DIVA clusters dynamic data without needing cluster count, outperforming baselines.
MFCVAE clusters data over multiple facets, improving disentanglement and generation.
Locally isoperimetric partitions minimize perimeter in space.
We show a very simple and general total second variation formula for Perelman's -functional at arbitrary points in the space of Riemannian metrics. Moreover we perform a study of the properties of the variations of Kähler structures. We deduce a quite simple and general total second variation formula for P…
One iteration of standard -means (i.e., Lloyd's algorithm) or standard EM for Gaussian mixture models (GMMs) scales linearly with the number of clusters , data points , and data dimensionality . In this study, we explore whether one iteration of -means or EM for GMMs can scale sublinearly with at run…
We give a microscopic representation of the stock-market in which the microscopic agents are the individual traders and their capital. Their basic dynamics consists in the auto-catalysis of the individual capital and in the global competition/cooperation between the agents mediated by the total wealth invested in the s…
A new method clusters survival data using deep variational models.
Most generative models for clustering implicitly assume that the number of data points in each cluster grows linearly with the total number of data points. Finite mixture models, Dirichlet process mixture models, and Pitman--Yor process mixture models make this assumption, as do all other infinitely exchangeable cluste…
Most generative models for clustering implicitly assume that the number of data points in each cluster grows linearly with the total number of data points. Finite mixture models, Dirichlet process mixture models, and Pitman--Yor process mixture models make this assumption, as do all other infinitely exchangeable cluste…
We derive variational formulas for the total Q-prime curvature under the deformation of strictly pseudoconvex domains in a complex manifold. We also show that the total Q-prime curvature agrees with the renormalized volume of such domains with respect to the complete Einstein-Kähler metric. In the appendix, by Rod Gove…
Algorithm tackles clustered contextual bandits with resource constraints.
Paper proposes MMC to avoid high-density bias in clustering.
We propose a general variational framework of fair clustering, which integrates an original Kullback-Leibler (KL) fairness term with a large class of clustering objectives, including prototype or graph based. Fundamentally different from the existing combinatorial and spectral solutions, our variational multi-term appr…
Proposes a new model for clustering multiplex networks with compositional data.
We consider the problem of estimating a function defined over locations on a -dimensional grid (having all side lengths equal to ). When the function is constrained to have discrete total variation bounded by , we derive the minimax optimal (squared) estimation error rate, parametrized by …
A new method captures higher-order interactions in data clusters.
This paper represents a preliminary (pre-reviewing) version of a sublinear variational algorithm for isotropic Gaussian mixture models (GMMs). Further developments of the algorithm for GMMs with diagonal covariance matrices (instead of isotropic clusters) and their corresponding benchmarking results have been published…
Proposes VCLANC for attributed network clustering using node and attribute embeddings.
Paper presents a reparameterized DP-DLGMM for clustering.