3DGCN predicts molecular properties and biochemical activities using 3D molecular graph.
problem Predicting molecular properties and biochemical activities from 3D molecular graphs.
method Unified graph convolution with learning operations to handle spatial information, distinguishing 3D rotations.
result Significantly higher performance on various molecular tasks compared to other deep-learning models.
Geometric modeling for human food and chemical sensitivities.
problem Modeling biochemical processes in humans with sensitivities.
method Geometric approach to biochemical modeling.
result Geometric models improve understanding of sensitivities.
siRF identifies transcription factor binding near enhancers in flies.
problem Identifying functional transcription factor binding near enhancers.
method Signed iterative random forests (siRF) for machine learning.
result Infers regulatory interactions among transcription factors and enhancers.
PUMA interprets metabolomics data to predict pathway activity and assign chemical identities.
problem Interpreting metabolomics data to determine biochemical pathway activities.
method Generative probabilistic modeling using stochastic sampling.
result PUMA predicts pathway activity and assigns chemical identities to metabolites.
Extracts biological context from biomedical texts to associate with events.
problem Identifying biological context and associating it with biochemical events in texts.
method Analyzed an annotated corpus and developed classifiers using syntactic, distance, and frequency features.
result Developed and evaluated classifiers for context-event association.
Bayesian inference for biochemical reaction networks using jump-diffusion approximations.
problem Estimating hidden quantities in poorly characterized biochemical processes.
method Developed a Bayesian inference algorithm based on Markov chain Monte Carlo and sequential Monte Carlo methods.
result Numerical evaluation of the algorithm for a partially observed multi-scale birth-death process.
Enhances drug discovery models by understanding human language.
problem Low predictive quality of activity prediction models in drug discovery.
method Proposes a novel architecture with separate chemical and natural language input modules and a contrastive pre-training objective.
result Improves predictive performance on few-shot and zero-shot learning benchmarks.
Novel parallel GNN predicts protein-ligand interactions with high accuracy.
problem Accurate prediction of protein-ligand interactions for drug design.
method Parallel Graph Neural Networks (GNN) integrating 3D structural data.
result GNN achieves high accuracy in predicting binary interactions and activity.
We propose the time-dependent generalization of an `ordinary' autonomous human biomechanics, in which total mechanical + biochemical energy is not conserved. We introduce a general framework for time-dependent biomechanics in terms of jet manifolds derived from the extended musculo-skeletal configuration manifold. The …
We propose a method for learning cyclic causal models from a combination of observational and interventional equilibrium data. Novel aspects of the proposed method are its ability to work with continuous data (without assuming linearity) and to deal with feedback loops. Within the context of biochemical reactions, we a…
NLP techniques improve drug discovery by analyzing chemical and protein text.
problem Improving drug discovery through better analysis of chemical and protein text.
method Natural language processing techniques applied to biochemical entities.
result Enhanced prediction of molecular properties and design of novel molecules.
New model uses pretrained biochemical language models to generate drug compounds.
problem Developing novel compounds targeting specific proteins.
method Exploits pretrained language models to initialize and fine-tune targeted molecule generation models.
result Warm-started models outperform baseline models, with one-stage strategy showing better generalization.
Neural network model improves leaf spectral reflectance prediction for grapevines.
problem Inaccurate modeling of grapevine leaf spectral reflectance from traits.
method Multi-head attention neural network trained on grapevine-specific data.
result Model achieved high accuracy (R^2=0.84, NRMSE=1.52%) and outperformed PROSPECT-PRO.
MEP-Net uses MEP to generate solutions from limited data.
problem Generating solutions to scientific problems with incomplete information.
method Combines MEP with neural networks to learn complex distributions from moment constraints.
result Demonstrates MEP-Net's effectiveness in modeling biochemical reaction networks and generating complex distributions.
A new GNN module learns geometric scattering features for better graph classification and feature exploration.
problem Learning long-range graph relations and extracting meaningful features from graphs.
method Proposes a learnable geometric scattering (LEGS) module in graph neural networks (GNNs), incorporating wavelet filters.
result LEGS-based GNNs outperform existing methods in graph classification and feature extraction tasks.
MoleculeSTM learns from molecule structures and texts for better drug design.
problem Lack of integration between chemical structures and textual knowledge in AI drug discovery.
method Jointly learns chemical structures and texts via contrastive learning, using a large dataset.
result MoleculeSTM achieves state-of-the-art performance in zero-shot tasks like structure-text retrieval and molecule editing.
Chemical networks outperform spiking neural networks in classification tasks.
problem Learning tasks with spiking neural networks require hidden layers, which are computationally expensive.
method Used deterministic mass-action kinetics to prove chemical reaction networks without hidden layers can solve tasks previously solved by spiking neural networks.
result A chemical reaction network without hidden layers outperforms a spiking neural network with hidden layers in a handwritten digit classification task.
BSBO optimizes constraints for high-throughput experiments.
problem Optimizing high-throughput experiments with combinatorial constraints.
method Stochastic Bayesian optimization with submodular decomposition.
result BSBO outperforms heuristics in real-world protein datasets.
In this paper we propose the time-dependent generalization of an `ordinary' autonomous human biomechanics, in which total mechanical + biochemical energy is not conserved. We introduce a general framework for time-dependent biomechanics in terms of jet manifolds associated to the extended musculo-skeletal configuration…
nUDEs use neural networks to model biology without negative values.
problem Unrealistic negative values in hybrid models of biology.
method Developed non-negative UDEs (nUDEs) with regularization techniques.
result nUDEs provide realistic solutions for biological models.
SparseChem speeds up ML for small molecules.
problem Training fast and accurate ML models for high-dimensional data.
method Supports millions of features and compounds, trains various models.
result Fast and accurate machine learning models for biochemical applications.
A fundamental aspect of biological information processing is the ubiquity of sequence-function relationships -- functions that map the sequence of DNA, RNA, or protein to a biochemically relevant activity. Most sequence-function relationships in biology are quantitative, but only recently have experimental techniques f…
Complex biological systems have been successfully modeled by biochemical and genetic interaction networks, typically gathered from high-throughput (HTP) data. These networks can be used to infer functional relationships between genes or proteins. Using the intuition that the topological role of a gene in a network rela…
Enhances drug discovery by optimizing molecular structures.
problem Accelerate drug discovery through better optimization of precursor molecules.
method Integrates substructure components with atom-level encoding in a fully autoregressive graph decoder.
result Significantly outperforms previous state-of-the-art baselines on molecular optimization tasks.
AMP0 predicts antimicrobial peptides targeting specific microbes.
problem Low-throughput screening of antimicrobial peptides.
method Zero-shot and few-shot machine learning.
result AMP0 can predict antimicrobial activity against specific microbes.
Design of experiments improves validation of biomolecular networks.
problem Efficiently validate non-machine learning designed biomolecular networks.
method Use Gaussian processes and Bayesian optimization to select experimental points.
result Developed a stopping criterion based on discrepancy metric and uncertainty.
Local search improves GFlowNets' ability to generate high-reward samples.
problem GFlowNets struggle with over-exploration in high-reward space.
method Local search focusing on high-reward samples via backtracking and reconstruction.
result Significant performance improvement in biochemical tasks.
Boosts GNN performance on molecular graphs.
problem Current GNNs struggle with training set and scalability.
method Proposes an auxiliary module to enhance GNNs.
result Improves GNN performance on molecular datasets.
This thesis tackles data fusion issues across different biological scales and types.
problem Heterogeneity in data types and scales in systems biology.
method Developed statistical methods to fuse heterogeneous data sets.
result Advantages of proposed methods assessed through simulations and real data analysis.
Study shows annealing with adaptive schedule reduces mode collapse in NFs for parameter estimation.
problem Mode collapse in normalizing flows for multimodal distributions.
method Annealing with an adaptive schedule based on effective sample size (ESS).
result Our approach reduces mode collapse and converges marginal likelihood faster than MCMC methods.
New algorithm improves model generalization in structured biomedical domains.
problem Improving model generalization in structured biomedical domains.
method Proposes a new regret minimization (RGM) algorithm and its structured extension for better performance in diverse environments.
result Significantly outperforms previous state-of-the-art baselines on molecular property prediction, protein homology, and stability prediction.
Deep neural networks correct Mie scattering in FTIR spectra of biological samples.
problem Mie scattering obscures biochemically relevant spectral information in FTIR spectra of biological samples.
method Deep neural networks to approximate the preprocessing function that removes Mie scattering.
result The model is faster and more generalizable across different tissue types.
During the past decade, with the significant progress of computational power as well as ever-rising data availability, deep learning techniques became increasingly popular due to their excellent performance on computer vision problems. The size of the Protein Data Bank has increased more than 15 fold since 1999, which …
New method infers dynamical systems from population data.
problem Inferring dynamical systems from population data.
method Deducing and estimating Fokker-Planck equation, projecting to test functions, sparse inference.
result Induces driving forces of dynamical systems.
Framework predicts mortality risk in MAFLD subjects.
problem Lack of mortality prediction methods for MAFLD subjects.
method Artificial Intelligence-based framework MAFUS using ML algorithms.
result Support Vector Machines (SVM) is the best model for mortality prediction.
Deviance-style normalization for sparse, jointly overdispersed count matrices
problem Jointly overdispersed count matrices
method Dirichlet-multinomial deviance residualization
result Preserves exact sparsity, evaluates in constant time, recovers multinomial residual
Method learns drug-disease representations for repositioning opportunities.
problem Identifying new uses for existing drugs.
method Multi-relation unsupervised graph embedding model.
result Superior prediction performance in repositioning opportunities.
Networks have in recent years emerged as an invaluable tool for describing and quantifying complex systems in many branches of science. Recent studies suggest that networks often exhibit hierarchical organization, where vertices divide into groups that further subdivide into groups of groups, and so forth over multiple…
Graph Beta Diffusion (GBD) generates graphs with mixed discrete and continuous components.
problem Generating graphs with mixed discrete and continuous components.
method Introduces Graph Beta Diffusion (GBD) using a beta diffusion process.
result Competes strongly with existing models across graph benchmarks.
Proposes using entity embedding vectors to improve Gaussian Process models for knowledge transfer across cell lines.
problem Lack of reuse of experimental data for predicting novel processes.
method Hybrid Gaussian Process models with entity embedding vectors to represent product identity.
result Improved performance in predicting novel processes compared to traditional methods.
Adaptive teacher improves sample efficiency and mode coverage in sampling tasks.
problem Efficient exploration and mode coverage in sampling tasks with amortized inference.
method An adaptive training distribution (teacher) guides the training of the primary sampler (student).
result Improves sample efficiency and mode coverage across various tasks.
Understanding the adaptation process of plants to drought stress is essential in improving management practices, breeding strategies as well as engineering viable crops for a sustainable agriculture in the coming decades. Hyper-spectral imaging provides a particularly promising approach to gain such understanding since…
BBRT improves molecular properties through iterative translation.
problem Optimizing molecular structures for improved biochemical properties.
method Iterative translation of molecules using a black box approach.
result Improvement in molecular properties with each iteration of the translator.
Proteins are commonly used by biochemical industry for numerous processes. Refining these proteins' properties via mutations causes stability effects as well. Accurate computational method to predict how mutations affect protein stability are necessary to facilitate efficient protein design. However, accuracy of predic…
New method optimises learning via surrogate PAC-Bayes bounds.
problem Computational intractability of optimising generalisation bounds.
method Iteratively optimising surrogate training objectives derived from PAC-Bayes bounds.
result Iteratively optimising surrogates implies optimising original generalisation bounds.
Many activation functions have been proposed in the past, but selecting an adequate one requires trial and error. We propose a new methodology of designing activation functions within a neural network at each layer. We call this technique an "activation ensemble" because it allows the use of multiple activation functio…
This paper studies activation sparsity in large language models, finding key trends and implications.
problem Activation sparsity in large language models (LLMs) can be improved for efficiency and interpretability.
method Proposes PPL-p% sparsity, analyzes trends with training data, width-depth ratio, and parameter scale. result ReLU is more efficient for sparsity than SiLU, and deeper architectures can improve sparsity.
Study uses HMM for real-time activity recognition from sensor data.
problem Real-time activity recognition from streaming sensor data.
method Online hierarchical hidden Markov model.
result Improved activity recognition accuracy compared to existing methods.