Optimal algorithm selects biological models without prior info.
problem Determining the correct biological model without prior knowledge.
method Systems biology models and likelihood-free inference.
result Improved model selection performance over conventional methods.
TDA improves accuracy of machine learning models for repeated measurements.
problem Limited accuracy of machine learning models for repeated measurements.
method Samples from data space, builds network graph based on data topology.
result TDA classifier achieves high accuracy (up to 96.8%) in repeated measurement datasets.
Unified platform for statistical and machine learning in bioinformatics.
problem Workflow inefficiencies in using multiple tools for data analysis.
method Automated hyperparameter optimization, feature importance analysis, statistical tests.
result Accelerates biological discovery workflows with methodological soundness.
Classic economic science is reaching the limits of its explanatory powers. Complexity science uses an increasingly larger set of different methods to analyze physical, biological, cultural, social, and economic factors, providing a broader understanding of the socio-economic dynamics involved in the development of nati…
New comparison shows differences in how value is incorporated in AIF and CAI.
problem Clarifying the relationship between Active Inference and Control-as-Inference.
method Formal comparison of AIF and CAI frameworks.
result Primary difference is how value is incorporated into generative models.
DASH simplifies neural networks for gene regulatory dynamics using domain knowledge.
problem Pruning neural networks for gene regulatory dynamics lacks biologically meaningful structure learning.
method DASH uses domain-specific structural information to guide network pruning, leading to sparser, better interpretable models.
result DASH outperforms general pruning methods in gene regulatory network inference, yielding deeper insights.
Enzyme sequences and structures are routinely used in the biological sciences as queries to search for functionally related enzymes in online databases. To this end, one usually departs from some notion of similarity, comparing two enzymes by looking for correspondences in their sequences, structures or surfaces. For a…
Paper explores why compositionality is not the right approach for emergent communication.
problem Lack of interaction between biological and computational models of emergent communication.
method Exploring the claim that compositionality is the wrong target for explaining natural language emergence.
result Suggests reflexivity as a better target for explaining natural language emergence.
Deep active inference agents learn complex environments using Monte-Carlo methods.
problem Understanding and modeling biological intelligence in complex, continuous state-spaces.
method Neural architecture for deep active inference agents using multiple forms of Monte-Carlo sampling.
result Deep active inference agents can learn environmental dynamics and plan future actions.
A new model for latent class analysis with weighted responses.
problem Limitation of latent class model for real-world data with continuous or negative responses.
method Proposed a novel generative model, the weighted latent class model (WLCM).
result The proposed WLCM is more realistic and general than the latent class model.
A main challenge of data-driven sciences is how to make maximal use of the progressively expanding databases of experimental datasets in order to keep research cumulative. We introduce the idea of a modeling-based dataset retrieval engine designed for relating a researcher's experimental dataset to earlier work in the …
ART automates synthetic biology design with machine learning.
problem Long development times in synthetic biology due to ad-hoc engineering.
method Machine learning and probabilistic modeling for systematic design.
result ART provides optimized strain recommendations and production levels.
New framework for network regression models accounting for community structure.
problem Inaccurate modeling of residual dependencies in network regression models.
method Modeling errors as community-based and exploiting exchangeability properties.
result Parsimonious standard errors for regression parameters.
Paper tackles attribute pattern learning in high-dimensional SLAMs.
problem Learning significant attribute patterns from high-dimensional SLAMs.
method Proposes a penalized likelihood method for selecting attribute patterns.
result Establishes selection consistency in overfitted SLAMs.
Two new minor minimal intrinsically chiral graphs identified.
problem Identifying intrinsically chiral graphs in molecular structures.
method Analyzing graph symmetry and embedding properties.
result Found two new minor minimal intrinsically chiral graphs Γ7 and Γ8. BioHash improves similarity search performance using sparse high-dimensional hash codes.
problem Improving similarity search performance in high-dimensional data.
method BioHash produces sparse high-dimensional hash codes through a data-driven approach based on synaptic plasticity.
result BioHash outperforms previous hashing methods in various similarity search tasks.
Study on detecting and recovering hidden dense cycles in random graphs.
problem Detecting and recovering hidden dense cycles in random graphs.
method Information-theoretic analysis of thresholds for detection and recovery.
result Characterization of information-theoretic thresholds for detection and recovery.
Bayesian optimization simplifies bioprocess engineering experiments.
problem Complex biological systems and experimental uncertainty.
method Adapts classical Bayesian optimization for bioprocess engineering.
result Provides accessible introduction to Bayesian optimization for practitioners.
Genomic models learn DNA sequences to predict functions.
problem Understanding complex genetic interactions.
method Training LLMs on DNA sequences to predict functions.
result gLMs can predict functions of DNA elements.
New algorithm predicts geolocation of fungi samples with high accuracy.
problem Identifying the origin of biological material at crime scenes.
method Ensemble of deep neural network classifiers trained on Voronoi partitions.
result More than half of geolocation errors under 100 kilometers for continental analysis and nearly 90% accuracy for global analysis.
CRF model improves protein secondary structure prediction.
problem Improving secondary structure prediction of proteins.
method Applied Conditional Random Fields (CRF) to protein classification.
result CRF model leads to extremely accurate protein secondary structure predictions.
Deep learning detects bifurcations in dynamical systems.
problem Predicting catastrophic changes in dynamical systems across sciences.
method Data-driven, physically-informed deep-learning framework for classifying dynamical regimes and characterizing bifurcation boundaries.
result Extracts topologically invariant features to detect bifurcation boundaries in unseen systems.
Researchers develop flexible kernels for biological sequences with guaranteed reliability.
problem Challenges in applying machine learning to biological sequences, including unreliable methods.
method Theoretical analysis and development of modified kernels to ensure reliability and accuracy.
result Developed kernels that are universal, characteristic, and metrize the space of distributions for biological sequences.
Skip connections improve biologically-inspired learning rules.
problem Biologically-inspired learning rules often underperform compared to backpropagation.
method Introduced skip connections between intermediate layers in biologically-motivated learning rules.
result Skip connections can match the performance of backpropagation and are robust to hyper-parameters.
SAMS-VAE models cellular perturbations using sparse additive mechanisms.
problem Modeling effects of diverse interventions on cells.
method Sparse Additive Mechanism Shift Variational Autoencoder (SAMS-VAE).
result SAMS-VAE identifies disentangled, perturbation-specific latent subspaces.
New learning algorithm mimics biological neural networks.
problem Biologically implausible backpropagation for directed neural networks.
method Introduces new neuronal dynamics and learning rule for arbitrary architectures, sparsity-inducing pruning method, and dynamical-systems characterization.
result Prunes irrelevant connections and improves learning efficiency.
Spatially positioned neurons in neural networks mimic biological systems.
problem Creating neural networks that can perform multiple tasks efficiently.
method Added spatial positions and proximity penalties to artificial neurons.
result Neurons naturally cluster, each responsible for a specific task.
Extracts biological context from biomedical texts to associate with events.
problem Identifying biological context and associating it with biochemical events in texts.
method Analyzed an annotated corpus and developed classifiers using syntactic, distance, and frequency features.
result Developed and evaluated classifiers for context-event association.
SENA-discrepancy-VAE interprets latent causal factors in biological pathways.
problem Interpreting latent causal factors in biological pathways.
method SENA-discrepancy-VAE, a model based on discrepancy-VAE, that produces interpretable latent causal factors.
result Sena-discrepancy-VAE achieves comparable predictive performance with non-interpretable counterparts while providing biologically meaningful causal factors.
The paper models retirement spending using biological age instead of chronological age.
problem Retirement spending varies at the same chronological age.
method Developed a stochastic mortality model to adjust for biological age.
result Optimal consumption rates derived using biological age.
Deep learning applied to biological data mining.
problem Mining complex biological data from diverse sources.
method Artificial neural networks, deep learning architectures.
result Deep learning techniques improve pattern recognition in biological data.
Paper explores stability, regularization, and gradient flows for stochastic inverse problems.
problem Recovering random probability distributions from measurements.
method Direct inversion, variational formulation with regularization, and optimization via gradient flows.
result The choice of metric impacts stability and properties of the optimizer.
Biological networks are a very convenient modelling and visualisation tool to discover knowledge from modern high-throughput genomics and postgenomics data sets. Indeed, biological entities are not isolated, but are components of complex multi-level systems. We go one step further and advocate for the consideration of …
Scalable GPLVM reduces complexity in scRNA-seq data, accounting for technical and biological confounders.
problem Complexity and confounders in scRNA-seq data hamper interpretation.
method Extended Gaussian process latent variable model (GPLVM) to handle large datasets.
result Framework reconstructs latent signatures and captures disease-specific gene expression.
DeepSIBA predicts biological effects of chemical structures using graph neural networks.
problem Predicting biological effects of chemical structures for drug discovery.
method Siamese Graph Convolutional Neural Networks for structure-biological effect mapping.
result Highly accurate predictions of biological effects for structurally dissimilar compounds.
Community detection in networks is a key exploratory tool with applications in a diverse set of areas, ranging from finding communities in social and biological networks to identifying link farms in the World Wide Web. The problem of finding communities or clusters in a network has received much attention from statisti…
New methods improve uncertainty quantification in dynamic biological systems.
problem Uncertainty in dynamic biological models due to nonlinearity and parameter sensitivity.
method Conformal inference methods for non-asymptotic guarantees.
result Enhanced robustness and scalability for diverse biological data structures.
Symmetry augmentation speeds up learning in robotics tasks.
problem Learning efficiency in robotics tasks with limited data.
method Data augmentation using symmetry in the quadruped domain of DeepMind control suite.
result Agent learns faster with augmented symmetry experiences.
We propose an adaptive sampling approach for multiple testing which aims to maximize statistical power while ensuring anytime false discovery control. We consider n distributions whose means are partitioned by whether they are below or equal to a baseline (nulls), versus above the baseline (actual positives). In addi…
Proposes a novel network-based neighborhood regression for biological systems.
problem Lack of comprehensive analysis on biological modules using both global and local network data.
method Develops a community-wise least square optimization approach to analyze gene modules and their regulatory strength.
result Achieves exact minimax optimality and linear consistency in identifying gene module associations.
AR algorithm simplifies backpropagation with improved scalability and biological plausibility.
problem Improving backpropagation algorithms for complex neural networks and biological plausibility.
method Introducing learnable backwards weights and avoiding nonlinear derivative computations; relaxing frozen feedforward pass assumption.
result Simplified AR algorithm maintains performance on complex CNN architectures and challenging datasets.
We develop a novel methodology based on the marriage between the Bhattacharyya distance, a measure of similarity across distributions of random variables, and the Johnson-Lindenstrauss Lemma, a technique for dimension reduction. The resulting technique is a simple yet powerful tool that allows comparisons between data-…
Neural likelihood approximates integer time series data efficiently.
problem Inference of parameters for integer-valued stochastic processes is challenging.
method Constructs a neural likelihood approximation for inference of parameters from time series data.
result Accurately approximates the true posterior with significant computational speed-ups.
Algorithm optimizes biological sequences using bootstrapped training with a score-conditioned generator.
problem Optimizing biological sequences for a black-box score function.
method Bootstrapped training of score-conditioned generator (BootGen) algorithm.
result Our method outperforms competitive baselines on biological sequential design tasks.
Unified framework connects physical laws and machine learning.
problem Combining physical laws and machine learning for scientific applications.
method Universal Differential Equations (UDEs) as a unifying framework.
result Wide variety of applications can be efficiently handled through UDE formalism.
Fault-tolerant neural networks inspired by biological error correction codes.
problem Achieving reliable computation with unreliable neurons.
method Using biological error correction codes from grid cells in the mammalian cortex to develop a fault-tolerant neural network.
result Noisy biological neurons operate below a fault-tolerance threshold, suggesting a mechanism for reliable computation in the brain.
This thesis tackles data fusion issues across different biological scales and types.
problem Heterogeneity in data types and scales in systems biology.
method Developed statistical methods to fuse heterogeneous data sets.
result Advantages of proposed methods assessed through simulations and real data analysis.
The process of transforming observed data into predictive mathematical models of the physical world has always been paramount in science and engineering. Although data is currently being collected at an ever-increasing pace, devising meaningful models out of such observations in an automated fashion still remains an op…