Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

77155232309 · Jun 202019922001200920182026
48 results for differential gene expression

DeepDiff predicts differential gene expression from histone modifications using deep learning.

problem Predicting differential gene expression from histone modification signals, capturing combinatorial effects.
method Attention-based deep learning architecture with multiple LSTM modules and attention mechanisms.
result DeepDiff significantly outperforms state-of-the-art baselines for differential gene expression prediction.

New methods detect continuous variation in single-cell data.

problem Continuous variation within and between cell types not detected by discrete analyses.
method Three topologically motivated mathematical methods for unsupervised feature selection.
result Detect additional biologically meaningful genes with coherent expression patterns.

SimCD simultaneously clusters cells and identifies differential gene expression in scRNA-seq data.

problem Separate clustering and differential expression analysis for scRNA-seq data leads to suboptimal results.
method Develops SimCD, a unified hierarchical gamma-negative binomial model for simultaneous cell clustering and differential expression analysis.
result SimCD outperforms existing methods in discovering cell clusters and capturing dynamic expression changes.

Concrete autoencoder selects key features for efficient data reconstruction.

problem Efficiently identifying and selecting important features for data reconstruction.
method Concrete selector layer with temperature-controlled selection during training, followed by reconstruction using a standard neural network.
result Concrete autoencoder selects a small subset of genes that can reconstruct the remaining gene expression levels, improving on existing methods.

The paper proposes a method to infer differentiation trees from RNA velocity data.

problem Reconstructing dynamic cellular processes from sequencing data.
method Defining varifold distances between RNA velocity curves to approximate shortest-path distances in a tree.
result The varifold distance method approximates the shortest-path distance in a tree isomorphic to the target differentiation tree.

Improved differentially private drug sensitivity prediction using compact representations.

problem Challenges in differentially private machine learning with genomic data.
method Representation learning using variational autoencoders, PCA, and random projection.
result Variational autoencoders provide the most accurate predictions for differentially private drug sensitivity prediction.

Next-generation sequencing technologies provide a revolutionary tool for generating gene expression data. Starting with a fixed RNA sample, they construct a library of millions of differentially abundant short sequence tags or "reads", which constitute a fundamentally discrete measure of the level of gene expression. A…

2013-01-17abs ↗pdf ↗

Bayesian model learns cell types and gene networks from two data views.

problem Estimating cell types and their regulatory networks from single-cell gene expression and epigenetic data.
method Symphony Bayesian hierarchical multi-view mixture model with Variational EM inference.
result Symphony outperforms other methods in learning cell types and regulatory networks.

A comprehensive benchmark of 15 scRNA-seq imputation methods across various datasets and analyses.

problem Imputation of single-cell RNA sequencing data to recover latent transcriptional signals.
method Evaluation of 15 imputation methods across 30 datasets and 6 downstream analyses.
result Traditional methods generally outperform DL-based methods in scRNA-seq data analysis.

This paper compares feature selection methods for biomarker discovery in toxicant-treated fish.

problem Choosing the most suitable method for biomarker discovery in toxicant exposure studies.
method Three feature selection methods: SAM, mRMR, and GeoDE are compared.
result Different methods perform better in different cases, requiring dataset-specific decisions.

Collaborative filtering predicts drug responses from gene expression data.

problem Predicting drug responses from large gene expression datasets with limited samples.
method Low-rank matrix factorization and latent linear regression.
result The proposed method outperforms state-of-the-art methods in predicting drug-gene associations.

New model generates realistic single-cell gene expression data.

problem Generating realistic single-cell gene expression profiles is challenging.
method scLDM, a latent diffusion model using Diffusion Transformers and linear interpolants.
result Superior performance in generating realistic single-cell gene expression data.

Autoencoder identifies cancer cells from normal ones using gene expression data.

problem Distinguishing between normal and cancer cells using gene expression profiles.
method Autoencoder trained on large tumor dataset to capture latent representations, using HPC toolkit for efficiency.
result Autoencoder node saliency identifies key differentiating features between normal and cancer cells.

Stem uses diffusion models to infer gene expression from H&E images.

problem Inference of gene expression from H&E stained images is time-consuming and expensive.
method Conditional diffusion generative model to infer gene expression.
result Stem achieves state-of-the-art performance in spatial gene expression prediction.

A novel method selects genes for high-dimensional gene expression data with class imbalance.

problem Class imbalance in gene expression datasets.
method Synthetic data balancing, greedy search, weighted robust score.
result The proposed method outperforms existing feature selection procedures.

Most network-based protein (or gene) function prediction methods are based on the assumption that the labels of two adjacent proteins in the network are likely to be the same. However, assuming the pairwise relationship between proteins or genes is not complete, the information a group of genes that show very similar p…

2012-12-03abs ↗pdf ↗

Graphs improve deep learning for gene expression data.

problem Challenges in applying deep learning to gene expression data due to non-linear signal and low sample sizes.
method Graph Convolutional Neural Networks (GCNN) with dropout and gene embeddings.
result GCNN with graph information provides an advantage in a low data regime.

Network Elastic Net identifies smoking-specific gene expression for lung cancer prognosis.

problem Identifying smoking-specific gene expression biomarkers in lung cancer prognosis.
method Introduces Network Elastic Net, a method that clusters and regresses on graphs based on smoking behavior.
result Shows efficacy of clusters in identifying cancer stages using gene expression and smoking behavior.

Proposes a method to identify relevant genes in autism-related diseases using auxiliary information.

problem Identifying relevant genes in autism-related diseases from diverse data sources.
method Uses logistic regression to filter irrelevant genes and clusters relevant genes into cohesive groups using adjacency matrix.
result Superior performance and robustness in finite samples observed in simulation studies.

The method integrates survival constraints into NMF for identifying survival-associated gene clusters.

problem Understanding and interpreting high-dimensional biological data for disease markers.
method Cox proportional hazards regression integrated with NMF via proportional hazards non-negative matrix factorization.
result The method can uncover survival-associated gene clusters in cancer gene expression data.

LogGENE uses log-cosh loss for deep learning in gene expression datasets, improving accuracy and interpretability.

problem Mining large gene expression datasets for reliable deep learning predictions.
method Develops a smooth alternative to check loss (log-cosh) for quantile regression in gene expression datasets.
result Achieves state-of-the-art performance in accuracy and provides robust uncertainty estimates.

REP predicts drug response at every stage of treatment using time-course gene expression data.

problem Lack of dynamic drug response prediction from time-course gene expression data.
method REP framework that predicts drug response values at every stage of a long-term treatment using recursive structure and tensor completion.
result REP can estimate drug response at any stage of a given treatment from initial gene expression levels.

A model to fill in missing gene data from spatial studies and scRNA-seq.

problem Imputing missing gene expression measurements from spatial transcriptomics.
method A deep generative model (gimVI) for integrating spatial transcriptomic and scRNA-seq data.
result gimVI outperforms existing methods in imputing missing genes.

New gene selection method improves tumor classification accuracy.

problem Efficiently selecting relevant genes from high-dimensional tumor gene expression data.
method Fuzzy-Rough Set Theory for feature dependency analysis.
result The proposed method outperforms state-of-the-art techniques in tumor classification.

A new method for joint eQTL mapping and gene network estimation.

problem Discovering SNP-gene relationships and gene-gene relationships in gene expression regulation.
method L1-2 regularized multi-task graphical lasso (L1-2 GLasso).
result Competitive performance on capturing true sparse structures of eQTL mapping and gene network.

Bayesian method discovers local causal relationships among genes from gene expression data.

problem Discovering gene regulatory relationships from gene expression data.
method Bayesian approach scoring covariance structures for triplets of normally distributed variables, incorporating background knowledge as priors.
result Stable and conservative posterior probability estimates of local causal structures.