Reanalysis of bioactivity prediction models suggests SVM performance is competitive with deep learning.
problem Benchmarking and validation of machine learning models in drug discovery.
method Reanalysis of a large-scale comparison of machine learning models for bioactivity prediction, using numerical experiments to question ROC curve relevance and suggest precision-recall curve.
result Support vector machines show competitive performance with deep learning methods in bioactivity prediction.
Deep convolutional neural networks comprise a subclass of deep neural networks (DNN) with a constrained architecture that leverages the spatial and temporal structure of the domain they model. Convolutional networks achieve the best predictive performance in areas such as speech and image recognition by hierarchically …
KANEL combines models for early hit enrichment in virtual screening.
problem Assessing model accuracy in chemical bioactivity predictions.
method Ensemble workflow using Kolmogorov-Arnold Networks (KANs) and other models.
result Improves early hit enrichment metrics like PPV@N.
Quantitative structure-activity relationship (QSAR) modelling is effective 'bridge' to search the reliable relationship related bioactivity to molecular structure. A QSAR classification model contains a lager number of redundant, noisy and irrelevant descriptors. To address this problem, various of methods have been pr…
Graph Informer improves molecular prediction tasks with expressive route-based attention.
problem Challenges in applying machine learning to molecular graphs due to information bottlenecks.
method Introduces a multi-attention mechanism that attends to nodes several steps away in molecular graphs.
result Improves 13C NMR spectra prediction with MAE of 1.35 ppm and drug bioactivity/toxicity prediction.
Empirical scoring functions based on either molecular force fields or cheminformatics descriptors are widely used, in conjunction with molecular docking, during the early stages of drug discovery to predict potency and binding affinity of a drug-like molecule to a given target. These models require expert-level knowled…
Q-SAVI model improves drug discovery accuracy with prior knowledge of chemical space.
problem Challenges in drug discovery due to covariate shift and limited labeled data.
method Probabilistic model with domain-informed prior distributions over functions.
result Q-SAVI outperforms state-of-the-art techniques in predictive accuracy and calibration.
Predicting bioactivity and physical properties of small molecules is a central challenge in drug discovery. Deep learning is becoming the method of choice but studies to date focus on mean accuracy as the main metric. However, to replace costly and mission-critical experiments by models, a high mean accuracy is not eno…
Without any means of interpretation, neural networks that predict molecular properties and bioactivities are merely black boxes. We will unravel these black boxes and will demonstrate approaches to understand the learned representations which are hidden inside these models. We show how single neurons can be interpreted…
This work improves online fine-tuning of diffusion models for specific properties.
problem Efficiently fine-tuning diffusion models to maximize specific properties.
method A novel reinforcement learning procedure that efficiently explores feasible samples.
result The method provides a regret guarantee and empirical validation across multiple domains.
Deep learning model integrates SMILES and molecular descriptors for EGFR inhibitor prediction.
problem Improving drug discovery by integrating structural and property data.
method Attention-based deep learning architecture trained on SMILES and molecular descriptors.
result Max MCC 0.58 and AUC 90% on EGFR inhibitors dataset, outperforming reference model.
Background: Pharmacokinetic evaluation is one of the key processes in drug discovery and development. However, current absorption, distribution, metabolism, excretion prediction models still have limited accuracy. Aim: This study aims to construct an integrated transfer learning and multitask learning approach for deve…
While the use of deep learning in drug discovery is gaining increasing attention, the lack of methods to compute reliable errors in prediction for Neural Networks prevents their application to guide decision making in domains where identifying unreliable predictions is essential, e.g. precision medicine. Here, we prese…
Unified model learns from proteins and ligands for drug design.
problem Disjoint data sources and modeling assumptions limit joint use of structure- and ligand-based drug design.
method Contrastive Geometric Learning for Unified Computational Drug Design (ConGLUDe)
result Unified model achieves competitive zero-shot virtual screening performance and state-of-the-art ligand-conditioned pocket selection.
New datasets support supervised learning for fungal BGC discovery.
problem Lack of labeled data for fungal BGCs.
method Developed new publicly available datasets for supervised learning.
result Supervised learning outperforms data-driven methods in fungal BGC prediction.
IGNN improves GNNs by maximizing edge-state transform mutual information.
problem Optimizing GNNs for better relational information.
method Variational information maximization to learn optimal transform parameters.
result IGNN achieves state-of-the-art performance on molecular graph tasks.
We tackle missing data in SBI methods and introduce a neural process approach.
problem Missing data in SBI methods can bias parameter estimation.
method We introduce a neural process approach to jointly learn imputation and inference.
result Our method provides robust inference outcomes compared to baselines.
Fuzzy prediction sets generalize binary predictions to include elements at varying confidence levels.
problem Binary prediction sets are limited; fuzzy prediction sets offer richer guarantees.
method Generalize prediction sets to fuzzy sets, showing they are e-values with merging properties.
result Optimal e-values lead to optimal fuzzy prediction sets, including optimal conformal prediction.
Paper defines predictive multiplicity and measures its severity in classification problems.
problem Challenges in machine learning due to competing models with conflicting predictions.
method Formal measures and integer programming tools for linear classification problems.
result Real-world datasets may admit competing models with wildly conflicting predictions.
This paper re-examines conformal e-prediction and its advantages over conformal prediction.
problem The relationship between conformal prediction and conformal e-prediction.
method Systematic re-examination of conformal prediction and conformal e-prediction from a modern perspective.
result Conformal e-prediction has advantages such as ease of designing conditional predictors and guaranteed validity of cross-predictors.
Two new methods improve efficiency of conformal predictive systems.
problem Efficiency of conformal predictive systems in regression problems.
method Split conformal predictive systems and cross-conformal predictive systems.
result Cross-conformal predictive systems are more efficient but not guaranteed valid.
Self-calibrating conformal prediction improves interval efficiency and offers a practical alternative.
problem Improving the reliability and uncertainty quantification of machine learning predictions.
method Combines Venn-Abers calibration and conformal prediction for binary and regression problems.
result Improves interval efficiency through model calibration and offers practical alternatives.
Study uses deep learning to predict asset prices, finds complex target processes lead to meaningless predictions.
problem Complexity of successful price prediction models hinders understanding.
method Deep learning models for high-frequency price prediction, focusing on volatility and directional prediction.
result Inadequately defined target price process renders predictions meaningless.
The paper emphasizes the importance of joint predictions over marginal predictions for decision-making.
problem The need for accurate joint predictions in decision-making problems.
method The paper analyzes combinatorial decision problems, sequential predictions, and multi-armed bandits, introducing an approximate Thompson sampling algorithm and new regret bounds.
result Accurate joint predictions are essential for good performance in decision-making problems.
Behavior modification improves prediction accuracy by nudging user behavior.
problem Improving prediction accuracy using behavior modification techniques.
method Combining prediction and behavior modification with reinforcement learning algorithms.
result Behavior modification can make predictions more certain but may not generalize.
Predictions can shape outcomes, study helps predict these effects.
problem Understanding how predictions influence real-world outcomes.
method Causal identifiability analysis of prediction-covariate-outcome relationships.
result Standard supervised learning can identify transferable relationships from predictions.
Proposes feature conformal prediction for broader application in semantic feature spaces.
problem Establishing valid prediction intervals in semantic feature spaces.
method Extends conformal prediction to semantic feature spaces using deep representation learning.
result Feature conformal prediction outperforms regular conformal prediction under mild assumptions.
Acute kidney injury (AKI) commonly occurs in hospitalized patients and can lead to serious medical complications. In order to optimally predict AKI before it develops at any time during a hospital stay, we present a novel framework in which AKI is continually predicted automatically from EHR data over the entire hospit…
AutoCP automates the construction of accurate prediction intervals.
problem Creating valid and accurate prediction intervals for machine learning models.
method AutoML framework that optimizes prediction interval length for better accuracy and less conservatism.
result AutoCP significantly outperforms benchmark algorithms in constructing accurate prediction intervals.
Proposes a method to apply conformal prediction to probabilistic time series forecasting models.
problem Obtaining accurate prediction regions for multi-step time series forecasting with probabilistic models.
method Conformalises conditional normalising flows to generate potentially disjoint prediction regions.
result Improves predictive efficiency in time series forecasting with multimodal distributions.
ICP improves prediction intervals for continuous outcomes at lower computational cost.
problem Systematic bias in point predictions that undermines their use in decision-making.
method Develops Isotonic Conformal Prediction (ICP) framework to decouple calibration from prediction-set construction.
result SICP and TICP procedures match SC-CP coverage at lower computational cost.
Optimizes predictions for specific tasks using parametrized decision analysis.
problem Optimizing predictions for specific decision tasks of interest.
method Designs a class of parametrized actions for Bayesian decision analysis.
result Derives efficient and interpretable solutions for various action parametrizations and loss functions.
New measures for prediction validity and consonant plausibility introduced.
problem Challenges in predicting future observations and quantifying prediction uncertainty.
method Introducing Type-2 validity and using consonant plausibility measures and conformal prediction.
result Achieving both Type-1 and Type-2 validity through consonant plausibility measures and conformal prediction.
RFpredInterval package builds prediction intervals for random forests and boosted forests.
problem Quantifying uncertainty in random forest and boosted forest point predictions.
method 16 methods to build prediction intervals with random forests and boosted forests.
result The proposed method outperforms existing methods in building prediction intervals.
FPPI selectively uses predictions to improve inference efficiency.
problem Improving statistical inference with limited labeled data and heterogeneous prediction quality.
method Filtered Prediction-Powered Inference (FPPI) framework.
result FPPI achieves strictly improved asymptotic efficiency compared to existing methods.
COP improves online conformal prediction by incorporating data patterns, leading to tighter prediction sets.
problem Overly conservative prediction sets in online conformal prediction methods when data distribution shifts.
method Conformal Optimistic Prediction (COP) incorporating estimated cumulative distribution function of non-conformity scores.
result COP produces tighter prediction sets with valid coverage guarantees, outperforming other methods.
Proposes a model to update industrial data predictions based on temporal changes.
problem Improving prediction accuracy in industrial data analytics by addressing changing conditions over time.
method Integrates similarity and loss functions to estimate and update prediction models adaptively.
result The data renewal model enhances prediction accuracy by identifying and updating model changes.
Unified framework for estimating random forest prediction errors.
problem Estimating prediction errors for random forests.
method Novel estimator of conditional prediction error distribution function.
result Proposed estimators enable competitive prediction intervals.
Predicts food ingredient amounts from images.
problem Predicting relative amounts of ingredients from food images.
method Proposes two deep learning models for sparse and dense predictions, with semi-automatic data pre-processing.
result Encouraging experimental results on a recipe dataset.
Most existing examples of full conformal predictive systems, split-conformal predictive systems, and cross-conformal predictive systems impose severe restrictions on the adaptation of predictive distributions to the test object at hand. In this paper we develop split-conformal and cross-conformal predictive systems tha…
This paper studies trade-offs in private prediction methods.
problem Leakage of training data information in machine learning predictions.
method Private training and private prediction methods with trade-offs.
result Private training methods outperform private prediction methods in various settings.
Conformal prediction helps quantify uncertainty but its use by humans is unclear.
problem Uncertainty quantification in predictions for human decision making.
method Decision theoretic framework for evaluating predictive uncertainty.
result Conformal prediction sets and human decision making goals are in tension.
New tree splitting criteria improve probabilistic predictions.
problem Improving tree-based nonparametric predictive distributions.
method Using proper scoring rules for tree splitting criteria.
result Trees with new splitting criteria produce better predictive distributions.
Prediction is arguably one of the most basic functions of an intelligent system. In general, the problem of predicting events in the future or between two waypoints is exceedingly difficult. However, most phenomena naturally pass through relatively predictable bottlenecks---while we cannot predict the precise trajector…
New method improves predictive systems with better theoretical guarantees.
problem Constructing predictive systems with out-of-sample calibration guarantees.
method Residual Distribution Predictive Systems (RDPs) that nest conformal predictive systems and offer flexibility.
result Empirically, RDPs perform competitively with conformal predictive systems and can be implemented with various regression methods.
Paper introduces PCP for efficient, reliable predictive inference.
problem Developing reliable predictive inference methods for target variables.
method Probabilistic conformal prediction using conditional random samples.
result PCP provides sharper predictive sets compared to existing methods.
A framework for private prediction sets using conformal prediction and differential privacy.
problem Jointly addressing reliability and privacy in machine learning predictions.
method Split conformal prediction with privatized quantile subroutine.
result Private prediction sets can be generated from privately-trained models.
Variational Prediction simplifies Bayesian inference without test time costs.
problem Bayesian inference's computational costs and posterior predictive distribution marginalization.
method Variational Prediction learns a variational approximation to the posterior predictive distribution using a variational bound.
result Directly learns a variational approximation to the posterior predictive distribution without test time marginalization costs.