InfoSEM infers gene regulatory networks without GT labels, improving performance.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Background: Predictive, stable and interpretable gene signatures are generally seen as an important step towards a better personalized medicine. During the last decade various methods have been proposed for that purpose. However, one important obstacle for making gene signatures a standard tool in clinics is the typica…
Molecular profiling data (e.g., gene expression) has been used for clinical risk prediction and biomarker discovery. However, it is necessary to integrate other prior knowledge like biological pathways or gene interaction networks to improve the predictive ability and biological interpretability of biomarkers. Here, we…
New datasets support supervised learning for fungal BGC discovery.
Stem uses diffusion models to infer gene expression from H&E images.
Estimates target GGM using auxiliary studies with false discovery rate control.
ASCEND discovers causal relationships in multi-omics data by leveraging known hierarchical structure.
Univariate and multivariate feature selection methods can be used for biomarker discovery in analysis of toxicant exposure. Among the univariate methods, differential expression analysis (DEA) is often applied for its simplicity and interpretability. A characteristic of methods for DEA is that they treat genes individu…
Paper introduces SUEL model for integrating predictors without labeled data.
Unsupervised two-view learning, or detection of dependencies between two paired data sets, is typically done by some variant of canonical correlation analysis (CCA). CCA searches for a linear projection for each view, such that the correlations between the projections are maximized. The solution is invariant to any lin…
Motivation: Algorithms that discover variables which are causally related to a target may inform the design of experiments. With observational gene expression data, many methods discover causal variables by measuring each variable's degree of statistical dependence with the target using dependence measures (DMs). Howev…
Unsupervised method selects genes for tumor subtype discovery.
New method for causal discovery using peeling algorithms for various data types.
AcceleratedLiNGAM speeds up causal discovery methods for large datasets.
A very important topic in systems biology is developing statistical methods that automatically find causal relations in gene regulatory networks with no prior knowledge of causal connectivity. Many methods have been developed for time series data. However, discovery methods based on steady-state data are often necessar…
Method uses network biology to construct gene expression models for cancer.
New method infers causal factors from large-scale data without full graph reconstruction.
BioBO optimizes gene perturbation design using Bayesian optimization with biological priors.
Motivation : Molecular signatures for diagnosis or prognosis estimated from large-scale gene expression data often lack robustness and stability, rendering their biological interpretation challenging. Increasing the signature's interpretability and stability across perturbations of a given dataset and, if possible, acr…
BaCaDI discovers causal structures from unknown interventions.
We study the performance of Local Causal Discovery (LCD), a simple and efficient constraint-based method for causal discovery, in predicting causal effects in large-scale gene expression data. We construct practical estimators specific to the high-dimensional regime. Inspired by the ICP algorithm, we use an optional pr…
JojoSCL improves scRNA-seq clustering by reducing intra-cluster dispersion.
DCCD-CONF discovers causal graphs with unmeasured confounders.
In most gene expression data, the number of training samples is very small compared to the large number of genes involved in the experiments. However, among the large amount of genes, only a small fraction is effective for performing a certain task. Furthermore, a small subset of genes is desirable in developing gene e…
MetaCaDI learns causal graphs and unknown interventions from few data instances.
DASH simplifies neural networks for gene regulatory dynamics using domain knowledge.
Synthetic lethality (SL) is a promising concept for novel discovery of anti-cancer drug targets. However, wet-lab experiments for detecting SLs are faced with various challenges, such as high cost, low consistency across platforms or cell lines. Therefore, computational prediction methods are needed to address these is…
Study improves cancer classification using gene selection and projection methods.
Expands experimental design for causal discovery from limited data.
Motivation: Cell-biological processes are regulated through a complex network of interactions between genes and their products. The processes, their activating conditions, and the associated transcriptional responses are often unknown. Organism-wide modeling of network activation can reveal unique and shared mechanisms…
This work tackles causal graph discovery with stochastic interventions to minimize the number of interventions.
Develops a new method to discover causal relationships from nonstationary time series data.
New algorithm reduces conditional independence tests needed for causal discovery.
RobKMR improves robustness in multi-omics data analysis for osteoporosis biomarker discovery.
Genome-wide association studies (GWA studies or GWAS) investigate the relationships between genetic variants such as single-nucleotide polymorphisms (SNPs) and individual traits. Recently, incorporating biological priors together with machine learning methods in GWA studies has attracted increasing attention. However, …
Gaussian Graphical Models (GGMs) are popular tools for studying network structures. However, many modern applications such as gene network discovery and social interactions analysis often involve high-dimensional noisy data with outliers or heavier tails than the Gaussian distribution. In this paper, we propose the Tri…
Nonparametric IPSS selects features with false discovery control.
Novel framework predicts cell responses to perturbations using GRNs.
Model predicts anti-cancer drug responses using gene and molecular data.
Clustering analysis is one of the most widely used statistical tools in many emerging areas such as microarray data analysis. For microarray and other high-dimensional data, the presence of many noise variables may mask underlying clustering structures. Hence removing noise variables via variable selection is necessary…
Complex systems may contain heterogeneous types of variables that interact in a multi-level and multi-scale manner. In this context, high-level layers may considered as groups of variables interacting in lower-level layers. This is particularly true in biology, where, for example, genes are grouped in pathways and two …
Motivation: Identifying interaction clusters of large gene regulatory networks (GRNs) is critical for its further investigation, while this task is very challenging, attributed to data noise in experiment data, large scale of GRNs, and inconsistency between gene expression profiles and function modules, etc. It is prom…
ENN method uses expectile regression for genetic data analysis of complex diseases.
New method controls false discoveries in structured hypothesis spaces.
A model learns causal graphs from summary statistics of synthetic data.
Deep neural networks (DNNs) are famous for their high prediction accuracy, but they are also known for their black-box nature and poor interpretability. We consider the problem of variable selection, that is, selecting the input variables that have significant predictive power on the output, in DNNs. We propose a backw…
A fundamental question in data analysis, machine learning and signal processing is how to compare between data points. The choice of the distance metric is specifically challenging for high-dimensional data sets, where the problem of meaningfulness is more prominent (e.g. the Euclidean distance between images). In this…
Estimates effect sizes and power from a pilot experiment.