Machine learning integrates diverse biological data to understand complex phenomena.
problem Combining multiple data types to understand biological and medical phenomena.
method Developing effective models to integrate heterogeneous biological data.
result Machine learning can identify important features and predict outcomes from diverse biological data.
Graph auto-encoder predicts unobserved node features from biological networks and omics data.
problem Integrating biological networks and continuous node features for better prediction.
method Graph neural networks and feature auto-encoders trained on feature reconstruction.
result Graph feature auto-encoder outperforms auto-encoders trained on graph reconstruction for predicting unobserved node features.
scICML integrates multi-omics data from single cells using co-clustering.
problem High noise and sparsity in multi-omics data from single cells.
method Information-theoretic co-clustering-based multi-view learning.
result Improves clustering performance and provides biological insights.
Unified framework improves gene prioritization in disease studies.
problem Identifying genes involved in diseases using heterogeneous biological data.
method Network propagation-based gene prioritization with integrated biological information.
result Significant improvements in prioritizing genes not identified by traditional methods.
Automated digital twin discovery from biological data improves drug discovery and personalized medicine.
problem Developing reliable digital twins from noisy, incomplete biological data.
method Symbolic and sparse regression, Bayesian frameworks, deep learning, and large language models.
result Sparse regression generally outperforms symbolic regression, especially with Bayesian frameworks.
Scalable GPLVM reduces complexity in scRNA-seq data, accounting for technical and biological confounders.
problem Complexity and confounders in scRNA-seq data hamper interpretation.
method Extended Gaussian process latent variable model (GPLVM) to handle large datasets.
result Framework reconstructs latent signatures and captures disease-specific gene expression.
Unified platform for statistical and machine learning in bioinformatics.
problem Workflow inefficiencies in using multiple tools for data analysis.
method Automated hyperparameter optimization, feature importance analysis, statistical tests.
result Accelerates biological discovery workflows with methodological soundness.
Paper proposes scalable method for analyzing multi-omic data.
problem Integrating high-dimensional multi-omic data for cancer subtyping.
method Mixed graphical model approach using Birth-Death MCMC algorithm.
result Our method outperforms LASSO and standard BDMCMC in computational efficiency and model selection accuracy.
SMAI framework tests and integrates single-cell data alignability.
problem Lack of a rigorous statistical test for alignability and distortion during alignment.
method Spectral manifold alignment and inference (SMAI) framework.
result SMAI outperforms existing methods in alignability testing and integration.
Biological neural network mimics CCA for multi-channel data.
problem Implementing CCA in a biologically plausible neural network.
method Derive an online CCA algorithm with local synaptic updates for multi-compartmental neurons.
result The derived neural network architecture and synaptic updates resemble cortical pyramidal neuron behavior.
Complex biological systems have been successfully modeled by biochemical and genetic interaction networks, typically gathered from high-throughput (HTP) data. These networks can be used to infer functional relationships between genes or proteins. Using the intuition that the topological role of a gene in a network rela…
MOTGNN integrates multi-omics data for disease classification with improved accuracy and interpretability.
problem Challenges in integrating multi-omics data due to high dimensionality, heterogeneity, and lack of reliable interaction networks.
method MOTGNN uses XGBoost for graph construction, modality-specific GNNs for representation learning, and a deep feedforward network for cross-omics integration.
result MOTGNN outperforms state-of-the-art baselines by 5-10% in accuracy, ROC-AUC, and F1-score across three real-world disease datasets.
PKB framework boosts genomic data analysis by integrating pathway knowledge.
problem Boosting discovery power and connecting new findings with biological mechanisms in genomic data.
method Pathway-based Kernel Boosting (PKB) framework integrating clinical and pathway information for prediction of various outcomes.
result PKB substantially outperforms other methods in predicting drug response and cancer survival.
Automated tests detect interactions in unstructured data.
problem Detecting interactions between latent variables in low-dimensional systems.
method Derive two interaction tests based on pairwise interventions and integrate them into an active learning pipeline.
result Tests can identify more known biological interactions than random search and standard active learning baselines.
Identifying measurable genetic indicators (or biomarkers) of a specific condition of a biological system is a key element of precision medicine. Indeed it allows to tailor diagnostic, prognostic and treatment choice to individual characteristics of a patient. In machine learning terms, biomarker discovery can be framed…
engGNN combines external and generated graphs to improve disease classification and biomarker discovery.
problem Challenges in integrating omics data due to high dimensionality and small sample sizes.
method Dual-graph framework that integrates external biological networks with data-driven generated graphs.
result engGNN outperforms state-of-the-art methods in disease classification and biomarker discovery.
InfoSEM infers gene regulatory networks without GT labels, improving performance.
problem Inferring GRNs from gene expression data with high accuracy and avoiding biases.
method InfoSEM uses deep generative models with informative priors (textual gene embeddings).
result InfoSEM outperforms existing models by 38.5% across four datasets.
Optimizes biomanufacturing processes with a new digital twin calibration method.
problem Lack of interpretability and sample efficiency in traditional DoE methods.
method Developed a computational approach to calibrate Bio-SoS digital twin model.
result Guides sample-efficient and interpretable DoEs by quantifying sub-model parameter estimation errors.
Deep learning applied to biological data mining.
problem Mining complex biological data from diverse sources.
method Artificial neural networks, deep learning architectures.
result Deep learning techniques improve pattern recognition in biological data.
Shallow networks with local learning rules can match deep learning performance.
problem Training deep neural networks is biologically implausible; the goal is to achieve similar performance with shallow networks.
method Investigated shallow networks with one hidden layer and a single readout layer, using various local learning rules for the hidden layer and supervised learning for the readout layer.
result Shallow networks can achieve test accuracy comparable to deep learning models, suggesting the use of different datasets for testing.
GeneDisco benchmarks experimental design for drug discovery.
problem Vast experimental design space in drug discovery.
method Machine learning for optimal experimental design.
result Standardised benchmark suite for active learning.
New methods learn causal networks from high-dimensional time series data.
problem Understanding complex dynamic systems with limited data and knowledge.
method Combination of Gaussian processes and abstract network modeling.
result Learn causal relations and synthesize causal networks from high-dimensional time series data.
BioBO optimizes gene perturbation design using Bayesian optimization with biological priors.
problem Efficient design of genomic perturbation experiments in drug discovery.
method Integrates Bayesian optimization with multimodal gene embeddings and enrichment analysis.
result Improves labeling efficiency by 25-40% and identifies top-performing perturbations more effectively.
Biological networks are a very convenient modelling and visualisation tool to discover knowledge from modern high-throughput genomics and postgenomics data sets. Indeed, biological entities are not isolated, but are components of complex multi-level systems. We go one step further and advocate for the consideration of …
The method integrates survival constraints into NMF for identifying survival-associated gene clusters.
problem Understanding and interpreting high-dimensional biological data for disease markers.
method Cox proportional hazards regression integrated with NMF via proportional hazards non-negative matrix factorization.
result The method can uncover survival-associated gene clusters in cancer gene expression data.
Framework integrates multi-omic data with network constraints for disease prediction.
problem Challenges in training models from multi-omic data with small sample size.
method Multi-view Factorization AutoEncoder with network constraints.
result Framework predicts disease progression-free interval and patient overall survival.
This thesis tackles data fusion issues across different biological scales and types.
problem Heterogeneity in data types and scales in systems biology.
method Developed statistical methods to fuse heterogeneous data sets.
result Advantages of proposed methods assessed through simulations and real data analysis.
Paper proposes MLPCD for protein community detection in large PPI networks.
problem Identifying reliable protein communities from large-scale PPI networks.
method Integrates Gene Expression Data and uses Multi-source Learning with cloud computing.
result Demonstrates superior performance compared to existing methods.
We introduce a novel Bayesian hybrid matrix factorisation model (HMF) for data integration, based on combining multiple matrix factorisation methods, that can be used for in- and out-of-matrix prediction of missing values. The model is very general and can be used to integrate many datasets across different entity type…
MIK improves t-SNE's local structure preservation in biological sequence data.
problem Efficiently preserving local structure in high-dimensional biological sequence data.
method Modified Isolation Kernel (MIK) using adaptive density estimation.
result MIK preserves local and global structure better than Gaussian and isolation kernels.
The paper proposes a method to integrate prior information into penalized regression.
problem Improving predictive performance in high-dimensional tasks with prior information.
method Integrating multiple sources of prior information into penalized regression.
result The method improves predictive performance, as shown by simulations and applications.
New method integrates sparse parametric and nonparametric techniques for complex system modeling.
problem Lack of accurate modeling for complex biological systems due to nonlinearities.
method Sparse nonparametric estimation framework combining parametric and nonparametric techniques.
result Accurately captures nonlinearities in complex systems without prior information.
Researchers develop flexible kernels for biological sequences with guaranteed reliability.
problem Challenges in applying machine learning to biological sequences, including unreliable methods.
method Theoretical analysis and development of modified kernels to ensure reliability and accuracy.
result Developed kernels that are universal, characteristic, and metrize the space of distributions for biological sequences.
Improves prediction performance on biological data by controlling confounding factors.
problem Challenges in statistical learning due to confounding variables in biological data.
method ONION for removing confounding covariates and DANN for penalizing confounder information.
result Significant improvements in generalization performance on simulated and empirical patient data.
Pipeline integrates cross-sectional and longitudinal multi-omics data for IBD research.
problem Integrating diverse data types from the same individuals for disease understanding.
method Statistical and deep learning methods for variable selection, feature extraction, and joint integration.
result Identified microbial pathways, metabolites, and genes discriminating IBD status.
DBNs improve accuracy of biological ODE models with missing data.
problem Uncertainty in biological ODE models with missing data.
method Converted ODE models to DBNs and used Particle Filtering for parameter estimation.
result DBNs can accurately infer model variables with missing data.
fiBAG integrates multiplatform genomic data to identify disease markers.
problem Understanding complex mechanisms underlying human diseases from multiplatform genomic data.
method fiBAG uses Gaussian process models and Bayes factors to identify functional evidence and guide variable selection.
result fiBAG improves detection of disease-related markers compared to non-integrative methods.
Omics-GAN uses GANs to generate synthetic multi-omics data for improved disease prediction.
problem Limited sample sizes, noise, and heterogeneity in multi-omics data reduce predictive power.
method Omics-GAN is a GAN-based framework that generates high-quality synthetic multi-omics profiles.
result Synthetic datasets consistently improved prediction accuracy compared to original omics profiles.
A new neural model evolves to learn at the synaptic level.
problem Lack of biologically realistic neural models in deep learning.
method Evolve individual neuron and synaptic models using ENUs.
result Evolved neural network learns complex tasks like a T-maze.
A new theory explains large associative memory with biological plausibility.
problem Large associative memory in neurobiology and machine learning.
method Microscopic theory with hidden neurons and two-body interactions.
result Valid model of large associative memory with biological plausibility.
AIME embeds multi-omics data to adjust confounders and find related features.
problem Extracting meaningful relationships between complex omics data types while accounting for confounders.
method Autoencoder-based deep learning approach that incorporates clinical confounders.
result AIME effectively adjusts for confounders and extracts biologically relevant features.
METCC learns distances to control confounders in high-dimensional data.
problem Technical and biological confounders in high-dimensional biological data.
method Contrastive metric learning using a non-linear triplet network.
result METCC outperforms linear methods in classifying biological samples.
FsNet selects features for high-dimensional biological data efficiently.
problem Efficient feature selection for high-dimensional biological data.
method FsNet combines selection and reconstruction layers with tiny networks for weight prediction.
result FsNet outperforms standard DNNs on high-dimensional biological datasets.
Develops a measure-theoretic framework for complex co-occurrence data.
problem Modeling and interpreting complex co-occurrences in high-dimensional data.
method Introduces measure-theoretic probability and conditional probability, investigates E-integrals.
result Establishes a rigorous measure-theoretic foundation for co-occurrence modeling.
Quantitatively predicting phenotype variables by the expression changes in a set of candidate genes is of great interest in molecular biology but it is also a challenging task for several reasons. First, the collected biological observations might be heterogeneous and correspond to different biological mechanisms. Seco…
Approximate Bayesian computation (ABC) using a sequential Monte Carlo method provides a comprehensive platform for parameter estimation, model selection and sensitivity analysis in differential equations. However, this method, like other Monte Carlo methods, incurs a significant computational cost as it requires explic…
Study builds complex network from multimodal physiological data.
problem Understanding dynamic interactions in biological systems.
method Network-based multimodal data fusion using recurrence plots and temporal metrics.
result Model accurately characterizes emotional states through physiological responses.
ProJIVE integrates multiple data types to explain joint and individual variation.
problem Integrating multiple types of data on the same subjects.
method Probabilistic EM algorithm for JIVE framework.
result ProJIVE learns biologically meaningful courses of variation and improves accuracy.