In recent years, the advances in single-cell RNA-seq techniques have enabled us to perform large-scale transcriptomic profiling at single-cell resolution in a high-throughput manner. Unsupervised learning such as data clustering has become the central component to identify and characterize novel cell types and gene exp…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Improved GPLVM model for single-cell RNA-seq data.
HSSE framework embeds single-cell RNA-seq data at multiple scales.
MarkerMap selects key genes for cell type analysis in single-cell RNA-seq.
DiffKnock improves feature selection in neural networks with complex dependencies and non-linear associations.
Super-OT combines GANs and optimal transport for lineage tracing.
Sources of variability in experimentally derived data include measurement error in addition to the physical phenomena of interest. This measurement error is a combination of systematic components, originating from the measuring instrument, and random measurement errors. Several novel biological technologies, such as ma…
The paper improves Fisher-Pitman tests for Poisson mixtures, detecting autism-related genes.
Until recently, transcriptomics was limited to bulk RNA sequencing, obscuring the underlying expression patterns of individual cells in favor of a global average. Thanks to technological advances, we can now profile gene expression across thousands or millions of individual cells in parallel. This new type of data has …
spex-LVM infers interpretable latent factors from biomedical data.
Causal methods for GRN inference from single-cell data often fail in real-world benchmarks.
DISCoVeR learns disentangled representations by separating shared and condition-specific factors.
Random small feature subsets outperform FS in diverse datasets.
Single-cell RNA sequencing (scRNA-seq) is a fast growing approach to measure the genome-wide transcriptome of many individual cells in parallel, but results in noisy data with many dropout events. Existing methods to learn molecular signatures from bulk transcriptomic data may therefore not be adapted to scRNA-seq data…
Computing the medoid of a large number of points in high-dimensional space is an increasingly common operation in many data science problems. We present an algorithm Med-dit which uses O(n log n) distance evaluations to compute the medoid with high probability. Med-dit is based on a connection with the multi-armed band…
Large datasets represented by multidimensional data point clouds often possess non-trivial distributions with branching trajectories and excluded regions, with the recent single-cell transcriptomic studies of developing embryo being notable examples. Reducing the complexity and producing compact and interpretable repre…
Proposes GFMMD for comparing signals on graphs.
Scalable GPLVM reduces complexity in scRNA-seq data, accounting for technical and biological confounders.
We introduce principal differences analysis (PDA) for analyzing differences between high-dimensional distributions. The method operates by finding the projection that maximizes the Wasserstein divergence between the resulting univariate populations. Relying on the Cramer-Wold device, it requires no assumptions about th…
Learning to align multiple datasets is an important problem with many applications, and it is especially useful when we need to integrate multiple experiments or correct for confounding. Optimal transport (OT) is a principled approach to align datasets, but a key challenge in applying OT is that we need to specify a tr…
New model clusters cells and individuals, revealing genetic influences on cell types.
Extracting an understanding of the underlying system from high dimensional data is a growing problem in science. Discovering informative and meaningful features is crucial for clustering, classification, and low dimensional data embedding. Here we propose to construct features based on their ability to discriminate bet…
Recent advances in high-throughput cDNA sequencing (RNA-Seq) technology have revolutionized transcriptome studies. A major motivation for RNA-Seq is to map the structure of expressed transcripts at nucleotide resolution. With accurate computational tools for transcript reconstruction, this technology may also become us…
EB-PCA reduces noise in high-dimensional PCA by estimating a joint prior distribution.
New model detects communities in networks with signed, continuous weights.
Long non-coding RNAs (lncRNAs) are a class of non-coding RNAs which play a significant role in several biological processes. RNA-seq based transcriptome sequencing has been extensively used for identification of lncRNAs. However, accurate identification of lncRNAs in RNA-seq datasets is crucial for exploring their char…
In this work we propose a method to compute continuous embeddings for kmers from raw RNA-seq data, without the need for alignment to a reference genome. The approach uses an RNN to transform kmers of the RNA-seq reads into a 2 dimensional representation that is used to predict abundance of each kmer. We report that our…
Unified framework for large-scale hypothesis testing with confounders.
Kernel testing compares cell states in single-cell data.
SMAI framework tests and integrates single-cell data alignability.
Study compares single vs ensemble feature selection for cancer diagnosis.
LMI approximates mutual information in high dimensions using learned low-dimensional representations.
Forest Fire Clustering discovers cell types from single-cell data.
NESS improves neighbor embedding for smooth cell-state transitions in single-cell data.
Proposes CCCVAE for better single-cell clustering with cell-cell communication.
New model generates realistic single-cell gene expression data.
Motivation: Single cell transcriptome sequencing (scRNA-Seq) has become a revolutionary tool to study cellular and molecular processes at single cell resolution. Among existing technologies, the recently developed droplet-based platform enables efficient parallel processing of thousands of single cells with direct coun…
ChemCPA predicts cellular responses to novel drugs using transfer learning.
Tutorial on using neural networks for single cell data analysis.
With ongoing developments and innovations in single-cell RNA sequencing methods, advancements in sequencing performance could empower significant discoveries as well as new emerging possibilities to address biological and medical investigations. In the study, we will be using the dataset collected by the authors of Sys…
scICML integrates multi-omics data from single cells using co-clustering.
Single-cell gene expression data provide invaluable resources for systematic characterization of cellular hierarchy in multi-cellular organisms. However, cell lineage reconstruction is still often associated with significant uncertainty due to technological constraints. Such uncertainties have not been taken into accou…
Clustering with variable selection is a challenging yet critical task for modern small-n-large-p data. Existing methods based on sparse Gaussian mixture models or sparse K-means provide solutions to continuous data. With the prevalence of RNA-seq technology and lack of count data modeling for clustering, the current pr…
Elastic co-clustering improves clustering of single-cell genomic data.
New method improves clustering accuracy in noisy single-cell data.
New methods improve analysis of single cell RNA sequencing data.
sgdGMF efficiently estimates generalized matrix factorization models for single-cell RNA sequencing data.
Single-cell RNA sequencing (scRNA-seq) has revolutionized biological discovery, providing an unbiased picture of cellular heterogeneity in tissues. While scRNA-seq has been used extensively to provide insight into both healthy systems and diseases, it has not been used for disease prediction or diagnostics. Graph Atten…