Forest Fire Clustering discovers cell types from single-cell data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New system constructs cell-type taxonomy across multiple samples.
In recent years, the advances in single-cell RNA-seq techniques have enabled us to perform large-scale transcriptomic profiling at single-cell resolution in a high-throughput manner. Unsupervised learning such as data clustering has become the central component to identify and characterize novel cell types and gene exp…
New model clusters cells and individuals, revealing genetic influences on cell types.
JojoSCL improves scRNA-seq clustering by reducing intra-cluster dispersion.
scICML integrates multi-omics data from single cells using co-clustering.
Proposes CCCVAE for better single-cell clustering with cell-cell communication.
New method improves clustering accuracy in noisy single-cell data.
Motivation: Single cell transcriptome sequencing (scRNA-Seq) has become a revolutionary tool to study cellular and molecular processes at single cell resolution. Among existing technologies, the recently developed droplet-based platform enables efficient parallel processing of thousands of single cells with direct coun…
Elastic co-clustering improves clustering of single-cell genomic data.
SimCD simultaneously clusters cells and identifies differential gene expression in scRNA-seq data.
Improved GPLVM model for single-cell RNA-seq data.
Multi-cell cooperative processing with limited backhaul traffic is studied for cellular uplinks. Aiming at reduced backhaul overhead, a sparsity-regularized multi-cell receive-filter design problem is formulated. Both unstructured distributed cooperation as well as clustered cooperation, in which base station groups ar…
IMPACC improves consensus clustering for bioinformatics data.
t-SNE and hierarchical clustering are popular methods of exploratory data analysis, particularly in biology. Building on recent advances in speeding up t-SNE and obtaining finer-grained structure, we combine the two to create tree-SNE, a hierarchical clustering and visualization algorithm based on stacked one-dimension…
We present a Bayesian hierarchical multi-view mixture model termed Symphony that simultaneously learns clusters of cells representing cell types and their underlying gene regulatory networks by integrating data from two views: single-cell gene expression data and paired epigenetic data, which is informative of gene-gen…
We introduce a tensor-based clustering method to extract sparse, low-dimensional structure from high-dimensional, multi-indexed datasets. This framework is designed to enable detection of clusters of data in the presence of structural requirements which we encode as algebraic constraints in a linear program. Our cluste…
New hypergraph method improves scRNA-seq clustering.
With ongoing developments and innovations in single-cell RNA sequencing methods, advancements in sequencing performance could empower significant discoveries as well as new emerging possibilities to address biological and medical investigations. In the study, we will be using the dataset collected by the authors of Sys…
A novel criterion selects optimal distance metrics for cell profile analysis.
We establish the Gaussian Multi-Bubble Conjecture: the least Gaussian-weighted perimeter way to decompose into cells of prescribed (positive) Gaussian measure when , is to use a "simplicial cluster", obtained from the Voronoi cells of equidistant points. Moreover, we prove that…
Paper proposes a new method for sparse spectral clustering on Stiefel manifold.
Proposes a model for identifying 4G cells with network throughput problems.
Motivation: With the development of droplet based systems, massive single cell transcriptome data has become available, which enables analysis of cellular and molecular processes at single cell resolution and is instrumental to understanding many biological processes. While state-of-the-art clustering methods have been…
TransST improves spatial transcriptomics data analysis by identifying cell clusters and biomarkers.
Extracting an understanding of the underlying system from high dimensional data is a growing problem in science. Discovering informative and meaningful features is crucial for clustering, classification, and low dimensional data embedding. Here we propose to construct features based on their ability to discriminate bet…
BanditPAM clusters data faster than traditional methods.
In future cell-free (or cell-less) wireless networks, a large number of devices in a geographical area will be served simultaneously in non-orthogonal multiple access scenarios by a large number of distributed access points (APs), which coordinate with a centralized processing pool. For such a centralized cell-free net…
New methods detect continuous variation in single-cell data.
Cluster LOCO: A model-agnostic feature importance score for interpreting cluster outputs
Proposes selective inference for testing differences in means between clusters.
New method uses dendrograms for better mixture model selection and clustering.
Generates infinite-depth hierarchical clusters from few examples.
New methods for visualizing multi-view data improve clustering accuracy.
Study of generalized double Bruhat cells and their integrations.
Flow cytometry is a high-throughput technology used to quantify multiple surface and intracellular markers at the level of a single cell. This enables to identify cell sub-types, and to determine their relative proportions. Improvements of this technology allow to describe millions of individual cells from a blood samp…
CTEF fits ellipsoids to noisy data in any dimension.
There are various algorithms and methodologies used for automated screening of cervical cancer by segmenting and classifying cervical cancer cells into different categories. This study presents a critical review of different research papers published that integrated AI methods in screening cervical cancer via different…
Clustering high-dimensional data, such as images or biological measurements, is a long-standingproblem and has been studied extensively. Recently, Deep Clustering has gained popularity due toits flexibility in fitting the specific peculiarities of complex data. Here we introduce the Mixture-of-Experts Similarity Variat…
We consider the use of the Joint Clustering and Matching (JCM) procedure for the supervised classification of a flow cytometric sample with respect to a number of predefined classes of such samples. The JCM procedure has been proposed as a method for the unsupervised classification of cells within a sample into a numbe…
Selective inference controls Type I error in k-means clustering tests.
A novel multi-resolution cluster detection (MCD) method is proposed to identify irregularly shaped clusters in space. Multi-scale test statistic on a single cell is derived based on likelihood ratio statistic for Bernoulli sequence, Poisson sequence and Normal sequence. A neighborhood variability measure is defined to …
We establish the Gaussian Double-Bubble Conjecture: the least Gaussian-weighted perimeter way to decompose into three cells of prescribed (positive) Gaussian measure is to use a tripod-cluster, whose interfaces consist of three half-hyperplanes meeting along an -dimensional plane at …
This paper deals with unsupervised clustering with feature selection. The problem is to estimate both labels and a sparse projection matrix of weights. To address this combinatorial non-convex problem maintaining a strict control on the sparsity of the matrix of weights, we propose an alternating minimization of the Fr…
sgdGMF efficiently estimates generalized matrix factorization models for single-cell RNA sequencing data.
We present *K-means clustering algorithm and source code by expanding statistical clustering methods applied in https://ssrn.com/abstract=2802753 to quantitative finance. *K-means is statistically deterministic without specifying initial centers, etc. We apply *K-means to extracting cancer signatures from genome data w…
Unsupervised clustering of curves according to their shapes is an important problem with broad scientific applications. The existing model-based clustering techniques either rely on simple probability models (e.g., Gaussian) that are not generally valid for shape analysis or assume the number of clusters. We develop an…
VampPrior Mixture Model improves clustering in DLVMs.