MicroRNAs (miRNAs) are small RNA molecules composed of 19-22 nt, which play important regulatory roles in post-transcriptional gene regulation by inhibiting the translation of the mRNA into proteins or otherwise cleaving the target mRNA. Inferring miRNA targets provides useful information for understanding the roles of…
Quantitative modeling of post-transcriptional regulation process is a challenging problem in systems biology. A mechanical model of the regulatory process needs to be able to describe the available spatio-temporal protein concentration and mRNA expression data and recover the continuous spatio-temporal fields. Rigorous…
Physically-inspired Gaussian process models study post-transcriptional regulation in Drosophila.
problem Understanding spatiotemporal interactions between mRNAs and gap proteins in post-transcriptional regulation.
method Two physically-inspired Gaussian process models based on reaction-diffusion equations, tested with mRNA expression data.
result Novel GP model requires only kernel function differentiation, simplifying spatial discretisation.
Deep learning models improve cancer detection and typing classification from gene expression data.
problem Challenges in establishing specificity for cancer diagnosis using gene expression data.
method Developed deep learning models using mRNA datasets for cancer detection and typing classification.
result Achieved 98% accuracy in cancer detection and 18 out of 32 cancer-typing classifications over 90% accuracy.
Despite great advances, molecular cancer pathology is often limited to the use of a small number of biomarkers rather than the whole transcriptome, partly due to computational challenges. Here, we introduce a novel architecture of Deep Neural Networks (DNNs) that is capable of simultaneous inference of various properti…
Proposes PSCCA for estimating correlations and canonical correlations in sparse count data.
problem Estimating correlations and canonical correlations in sparse count data from next-generation sequencing.
method Probabilistic approach for sparse count data sets (PSCCA).
result PSCCA outperforms other methods in estimating true correlations and canonical correlations at the natural parameter level.
With the wealth of high-throughput sequencing data generated by recent large-scale consortia, predictive gene expression modelling has become an important tool for integrative analysis of transcriptomic and epigenetic data. However, sequencing data-sets are characteristically large, and previously modelling frameworks …
We derive generalized estimators for a number of spatial statistics that have been used in the analysis of spatially resolved omics data, such as Ripley's K, H and L functions, clustering index, and degree of clustering, which allow these statistics to be calculated on data modelled by arbitrary random measures (RMs). …
To survive environmental conditions, cells transcribe their response activities into encoded mRNA sequences in order to produce certain amounts of protein concentrations. The external conditions are mapped into the cell through the activation of special proteins called transcription factors (TFs). Due to the difficult …
Gene expression levels in a population vary extensively across tissues. Such heterogeneity is caused by genetic variability and environmental factors, and is expected to be linked to disease development. The abundance of experimental data now enables the identification of features of gene expression profiles that are s…
Optimizing over the set of orthogonal matrices is a central component in problems like sparse-PCA or tensor decomposition. Unfortunately, such optimization is hard since simple operations on orthogonal matrices easily break orthogonality, and correcting orthogonality usually costs a large amount of computation. Here we…
Deep learning predicts RNA degradation from crowdsourced data.
problem Predicting RNA degradation to improve thermostability.
method Crowdsourced machine learning competition on Kaggle.
result 41% of predictions matched experimental data, and models generalized to longer RNA molecules.
Improved SBI with neural networks for complex models.
problem Accurate inference for complex models with intractable likelihood.
method Structured mixtures of probability distributions for likelihood and posterior approximation.
result Accurate posterior inference with smaller computational footprint.
With different genomes available, unsupervised learning algorithms are essential in learning genome-wide biological insights. Especially, the functional characterization of different genomes is essential for us to understand lives. In this book chapter, we review the state-of-the-art unsupervised learning algorithms fo…
We introduce a new approach to variable selection, called Predictive Correlation Screening, for predictor design. Predictive Correlation Screening (PCS) implements false positive control on the selected variables, is well suited to small sample sizes, and is scalable to high dimensions. We establish asymptotic bounds f…
MOTGNN integrates multi-omics data for disease classification with improved accuracy and interpretability.
problem Challenges in integrating multi-omics data due to high dimensionality, heterogeneity, and lack of reliable interaction networks.
method MOTGNN uses XGBoost for graph construction, modality-specific GNNs for representation learning, and a deep feedforward network for cross-omics integration.
result MOTGNN outperforms state-of-the-art baselines by 5-10% in accuracy, ROC-AUC, and F1-score across three real-world disease datasets.
Omics-GAN uses GANs to generate synthetic multi-omics data for improved disease prediction.
problem Limited sample sizes, noise, and heterogeneity in multi-omics data reduce predictive power.
method Omics-GAN is a GAN-based framework that generates high-quality synthetic multi-omics profiles.
result Synthetic datasets consistently improved prediction accuracy compared to original omics profiles.
A scalable Bayesian inference method for mixed-effects models in systems biology.
problem Scalable Bayesian inference for complex hierarchical mixed-effects models in systems biology.
method Constructing amortized approximations of likelihood and posterior distributions, refined for each individual dataset.
result Our method is both fast and competitive in statistical accuracy compared to exact pseudomarginal Bayesian inference.
BIDIFAC integrates multi-platform, multi-cohort data for shared and unique patterns.
problem Integration of multi-platform, multi-cohort data for shared and unique patterns.
method BIDIFAC integrates bidimensionally linked matrices into four components: globally shared, row-shared, column-shared, and single-matrix structural components.
result BIDIFAC reveals shared and unique patterns of variability in multi-platform, multi-cohort data.
SnapMMD forecasts cell differentiation outcomes from snapshot data.
problem Forecasting cell differentiation outcomes from limited snapshot data.
method SnapMMD learns dynamics by directly fitting the joint distribution of state measurements and observation time with MMD loss, allowing for unknown and state-dependent volatilities.
result SnapMMD delivers accurate forecasts and an R2-style statistic for diagnosing fit.
Proposes a novel classification criterion for high-dimensional data with few samples.
problem Challenges in classifying high-dimensional data with limited samples.
method Tolerance similarity criterion and No-separated Data Maximum Dispersion classifier (NPDMD).
result NPDMD outperforms state-of-the-art methods in various real-world applications.
Improves convex biclustering for high-dimensional data.
problem Discovering meaningful biclusters in high-dimensional data.
method Biconvex modification with adaptive feature weighting.
result Consistently recovers biclusters and selects features appropriately.
The medical research facilitates to acquire a diverse type of data from the same individual for particular cancer. Recent studies show that utilizing such diverse data results in more accurate predictions. The major challenge faced is how to utilize such diverse data sets in an effective way. In this paper, we introduc…
New method detects RNA modifications without prior training, revealing novel sites.
problem Detecting RNA modifications with high accuracy and sensitivity.
method Anomaly detection using nanopore raw ionic current signals and nearest neighbor comparison.
result Detects diverse RNA modifications without prior training, including a novel 2'-O-methylated site in DENV.
We solve the vector embedding problem by minimizing total distortion under constraints.
problem Assigning representative vectors to items with similarity and dissimilarity constraints.
method Projected quasi-Newton method for MDE problems, scalable to large data sets.
result Our method provides principled ways to validate embeddings and scales to millions of items.
Bias and flexibility trade off in learning algorithms.
problem Understanding the trade-off between bias and expressivity in learning algorithms.
method Measuring expressivity using entropy on algorithm outcome distributions, and deriving bounds on bias and expressivity.
result There is a necessary trade-off between bias and expressivity in learning algorithms.
Polynomial neural networks explore thresholds for maximum expressiveness.
problem Understanding the limits of polynomial neural networks' expressiveness.
method Introducing activation degree threshold to measure network expressiveness and proving its existence and upper bounds.
result Polynomial neural networks with equi-width architectures achieve the maximum expressiveness.
This work explores the relationship between expressivity and generalization in GNNs.
problem Understanding the trade-off between expressivity and generalization in GNNs.
method Introducing a novel framework that connects GNN generalization to the variance in graph structures they can capture.
result Theoretical findings align with empirical results, offering a deeper understanding of how expressivity enhances GNN generalization.
We present a technique to characterize differentially expressed genes in terms of their position in a high-dimensional co-expression network. The set-up of Gaussian graphical models is used to construct representations of the co-expression network in such a way that redundancy and the propagation of spurious informatio…
We survey results on neural network expressivity described in "On the Expressive Power of Deep Neural Networks". The paper motivates and develops three natural measures of expressiveness, which all display an exponential dependence on the depth of the network. In fact, all of these measures are related to a fourth quan…
Express improves causal attention guarantees for language models.
problem Improving causal attention guarantees for language models.
method Introducing Express, a tool for converting non-causal attention into causal with matching guarantees.
result Express improves causal attention guarantees to log^(3/2)(n)/s with minimal memory and compression overhead.
New method infers co-expression networks robustly from multiple studies.
problem Challenges in inferring co-expression networks from transcriptome data.
method Robust method based on multivariate t-distribution with shared precision matrix.
result Identifies co-expression matrix up to scaling factor.
New model generates realistic single-cell gene expression data.
problem Generating realistic single-cell gene expression profiles is challenging.
method scLDM, a latent diffusion model using Diffusion Transformers and linear interpolants.
result Superior performance in generating realistic single-cell gene expression data.
PNCs balance tractability and expressiveness in probabilistic modeling.
problem Balancing tractability and expressiveness in probabilistic models.
method Introduce probabilistic neural circuits (PNCs) as a mix of Bayesian networks and neural networks.
result PNCs are powerful function approximators.
Obtaining a non-parametric expression for an interventional distribution is one of the most fundamental tasks in causal inference. Such an expression can be obtained for an identifiable causal effect by an algorithm or by manual application of do-calculus. Often we are left with a complicated expression which can lead …
This paper explores how enforcing equivariance constraints limits neural network expressivity and proposes compensatory model size increases.
problem The impact of enforcing equivariance constraints on the expressive power of neural networks.
method Examined 2-layer ReLU networks, analyzed boundary hyperplanes and channel vectors, and constructed upper bounds on model size required for compensation.
result Enforcing equivariance constraints reduces the expressive power of neural networks, but this can be compensated by increasing model size.
A new deterministic method for symbolic regression finds mathematical expressions from data.
problem Finding mathematical expressions from datasets efficiently and reliably.
method Deterministic growth of simple expressions until they fit the data.
result Results are as good as other Machine Learning methods but in lower computational time.
New architectures improve topological deep learning's ability to capture complex data features.
problem Current TDL architectures struggle with fundamental topological and metric invariants.
method Developed multi-cellular networks (MCN) and scalable MCN (SMCN) to enhance expressivity.
result SMCN outperforms HOMP and expressive graph methods in learning topological properties.
Paper explores how GNNs can learn graph biconnectivity, finding ESAN is the only known expressive framework.
problem Understanding the expressive power of GNNs beyond the WL test.
method Introduces a novel class of expressivity metrics via graph biconnectivity and develops the GD-WL approach.
result GD-WL consistently outperforms prior GNN architectures in learning biconnectivity metrics.
Expressive quantum circuits are harder to train due to flatter cost landscapes.
problem Designing quantum circuits that are both expressive and trainable.
method Deriving a relationship between expressibility and gradient magnitude, extending barren plateau phenomenon.
result Highly expressive ansätze exhibit flatter cost landscapes, making them harder to train.
Identifying latent structure in large data matrices is essential for exploring biological processes. Here, we consider recovering gene co-expression networks from gene expression data, where each network encodes relationships between genes that are locally co-regulated by shared biological mechanisms. To do this, we de…
The paper introduces closed-form expressions for interpreting Tsetlin Machines.
problem Interpreting complex Tsetlin Machines with a large number of clauses.
method Developed closed-form expressions for local and global interpretability of Tsetlin Machines.
result The expressions enable real-time feature importance assessment and data clustering.
Transformers struggle to approximate smooth functions, relying on piecewise constant approximations.
problem Understanding the expressivity of Transformers for function approximation.
method Theoretical analysis and experimental validation of Transformer's ability to approximate smooth functions.
result Transformers cannot reliably approximate smooth functions, relying on piecewise constant approximations.
This paper addresses the problem of inferring a regular expression from a given set of strings that resembles, as closely as possible, the regular expression that a human expert would have written to identify the language. This is motivated by our goal of automating the task of postmasters of an email service who use r…
Collaborative filtering predicts drug responses from gene expression data.
problem Predicting drug responses from large gene expression datasets with limited samples.
method Low-rank matrix factorization and latent linear regression.
result The proposed method outperforms state-of-the-art methods in predicting drug-gene associations.
Enhanced text-to-speech synthesizes expressive speech from a single example.
problem Creating a new expressive speech style from a single example of speech.
method Combines VAE and Normalizing Flows to improve disentanglement and naturalness.
result Reduces KL-divergence by 22% and improves perceptual metrics.
The paper analyzes deep neural networks' expressivity and training, revealing critical expressivity issues.
problem Critical expressivity issues in deep neural networks.
method Quantitative analysis using Hilbert space and Hermite polynomials for feature mapping and activation function design.
result Deep neural networks evolve to the edge of chaos, but expressivity depends on overcoming convergence.
Weisfeiler-Leman struggles with graph isomorphism; enhanced architectures improve generalization.
problem Graph isomorphism problem and limited expressivity of 1-WL. method Augmenting 1-WL and MPNNs with subgraph information, employing margin theory, and introducing provable generalization kernels. result Increased expressivity of graph neural networks and kernels does not necessarily correlate with improved generalization performance.