Calibrated Boosting-Forest improves ranking and probability calibration in classification tasks.
problem Need for superior ranking power and well-calibrated probability estimates in classification tasks.
method Ensemble of gradient boosting machines supporting both continuous and binary labels.
result Calibrated Boosting-Forest achieves significant improvements in ranking and probability calibration compared to state-of-the-art models.
AVE measures redundancy in ligand-based benchmarks, revealing overfitting.
problem Overfitting in ligand-based classification benchmarks.
method AVE (Training-Validation Redundancy Measure) for ligand-based classification problems.
result Performance of ligand-based methods correlates with AVE bias, not generalization.
Unified model learns from proteins and ligands for drug design.
problem Disjoint data sources and modeling assumptions limit joint use of structure- and ligand-based drug design.
method Contrastive Geometric Learning for Unified Computational Drug Design (ConGLUDe)
result Unified model achieves competitive zero-shot virtual screening performance and state-of-the-art ligand-conditioned pocket selection.
Graph convolutions improve molecular modeling by leveraging graph structure.
problem Limited effectiveness of fingerprint representations in capturing molecular structure.
method Molecular graph convolutions, using graph structure encoding for machine learning.
result Graph convolutions enhance model's ability to utilize molecular graph information.
Deep neural network identifies potential SARS-CoV-2 inhibitors.
problem Finding novel therapies for SARS-CoV-2.
method Used ChemAI, a deep neural network trained on 220M data points, to screen and rank one billion molecules from the ZINC database.
result Identified 30,000 top-ranked compounds for further bioassays.
Multitask neural networks improve performance on industrial ADMET datasets.
problem Improving drug discovery with deep learning methods.
method Comparison of neural networks to baseline models, analysis of multitask learning effects.
result Multitask learning provides modest benefits over single-task models, especially for smaller datasets.
New method uses Riemannian geometry to describe molecular shapes.
problem Predicting drug-like molecules using shape similarity.
method Riemannian geometry applied to molecular surfaces.
result RGMolSA method captures molecular shape effectively.
A new method for virtual drug screening detects top treatments.
problem Understanding model performance in virtual drug screening tasks.
method Regression Enrichment Surfaces (RES) method.
result RES detects more top-performing treatments than existing methods.
ROCS-derived features enhance virtual screening performance.
problem Improving virtual screening accuracy.
method Decomposed ROCS color force field into color components and atom overlaps, creating weighted features.
result Significant improvement in virtual screening performance (ROC AUC scores).
Efficiently allocate budgets for LLM-assisted virtual screening to reduce costs.
problem Reducing the cost of evaluating alternatives in large-scale screening tasks.
method Propose a top-m greedy evaluation mechanism and the EFG-m algorithm for efficient budget allocation. result Prove that EFG-m is both sample-optimal and consistent in large-scale virtual screening. Deep learning predicts protein-small molecule binding.
problem Insufficient benchmark datasets for structure-based virtual screening.
method Learnable atom convolution, non-linear transformation, inner-product for binding potential prediction.
result New benchmark dataset improves testing of structure-based virtual screening methods.
KANEL combines models for early hit enrichment in virtual screening.
problem Assessing model accuracy in chemical bioactivity predictions.
method Ensemble workflow using Kolmogorov-Arnold Networks (KANs) and other models.
result Improves early hit enrichment metrics like PPV@N.
Study improves reliability of neural models for virtual screening.
problem Reliability issues in neural models for molecular property prediction.
method Investigated model architectures, regularization, and loss functions.
result Correct choice of regularization and inference methods improves reliability.
Bayesian learning improves reliability of molecular predictions for hit compound discovery.
problem Improving reliability of machine learning predictions for virtual screening.
method Bayesian learning algorithms applied to graph neural networks.
result Bayesian learning leads to well-calibrated predictions and higher hit compound success.
Novel ligand-based method improves protein representation performance.
problem Improving protein representation for bioinformatics tasks.
method Proposes SMILESVec method to represent ligands and compute protein similarity.
result Ligand-based protein representation performs as well as sequence-based methods.
Deep learning uses ROC cost functions to improve virtual screening accuracy.
problem Challenges in training deep learning models for virtual screening, especially class imbalance and lack of ground truth labels.
method Proposes using ROC cost functions to optimize deep learning models for virtual screening, introduces new training schemes and cost functions.
result Demonstrates improved performance of ROC-based approaches on PubChem datasets.
Improved CPI prediction for unknown compounds using PKM and LIK.
problem Poor performance of CGBVS for new compounds with unknown CPIs.
method Combining PKM with LIK and chemical similarity for link mining.
result Average AUPR value of 0.562, significantly higher than GIP method.
FlowMO uses Gaussian Processes for molecular property prediction with uncertainty.
problem Predicting molecular properties with uncertainty for small datasets.
method Gaussian Processes implemented in FlowMO, built on GPflow and RDKit.
result Comparable predictive performance to deep learning but superior uncertainty calibration.
MOSES benchmarks molecular generation models using a standardized dataset and metrics.
problem Unclear comparison and ranking of molecular generation models.
method Developed MOSES platform with training and testing datasets, metrics.
result Suggested MOSES results as reference for advancements in generative chemistry.
New neural network models predict molecular properties without 3D geometry, speeding up high-throughput screening.
problem Predicting molecular properties for large, complex molecules without computationally expensive 3D geometry.
method Message-passing neural networks trained with and without 3D structural information.
result Message-passing neural networks achieve similar accuracy to state-of-the-art methods without 3D geometry.
Improves molecular activity prediction using graph convolutional neural networks considering graph distances.
problem Predicting molecular activity using graph convolutional neural networks with improved distance representation.
method Proposed three improvements: modified graph distances, distance-dependent weight matrices, and weighted sum conversion.
result The proposed method slightly outperforms the original weave module in compound activity prediction.
PDTS accelerates chemical space exploration using parallel and distributed Thompson sampling.
problem Large chemical space makes brute force searches infeasible; high-throughput screening is needed but current BO methods cannot scale.
method Parallel and distributed Thompson sampling (PDTS) for scalable Bayesian optimization.
result PDTS outperforms other scalable methods in large-scale parallel BO.
CNN scoring function predicts protein-ligand interactions.
problem Scoring protein-ligand interactions for drug discovery.
method Convolutional Neural Networks (CNN) for 3D protein-ligand interactions.
result CNN scoring function outperforms AutoDock Vina in ranking poses.
AI models predict new opioid ligands from molecular dynamics.
problem Lack of crystal structures limits virtual screening of drug candidates.
method Molecular dynamics simulation and machine learning.
result Identified a novel μ opioid chemotype. Deep learning outperforms traditional methods in computational chemistry.
problem Challenges in computational chemistry, such as QSAR, virtual screening, and quantum chemistry.
method Application of deep neural networks to computational chemistry problems.
result Deep neural networks outperform traditional models across various computational chemistry tasks.
New approach safely screens features and samples simultaneously for sparse modeling.
problem Learning sparse models to identify active features and samples.
method Alternating feature and sample screening steps, exploiting synergy between steps.
result Practical advantage in problems with large numbers of features and samples.
RaSE screens variables via random subspaces, identifying joint effects.
problem Missing joint effects of predictors in ultra-high dimensional data.
method Random Subspace Ensemble (RaSE) framework combining subspace evaluation criteria.
result RaSE identifies signals with no marginal effect or high-order interactions.
Extends variable screening for ultrahigh-dimensional models, reducing dimensionality to sample size.
problem Statistical inference challenges in ultrahigh-dimensional linear models.
method Extends correlation-based variable screening to arbitrary linear models and post-screening inference techniques.
result Shows a condition (screening condition) sufficient for successful variable screening in arbitrary linear models.
Novel GNN predicts drug-target interactions using protein-ligand 3D structures.
problem Accurate prediction of drug-target interactions for in silico drug design.
method 3D structure-embedded graph representations and distance-aware graph attention algorithm with gate augmentation.
result Our model outperforms docking and other deep learning methods in virtual screening and pose prediction.
Safe screening reduces the number of triplets in metric learning.
problem Optimizing a metric over many triplets is computationally expensive and impractical.
method Safe triplet screening identifies and removes redundant triplets.
result Safe triplet screening maintains optimality without increasing computational cost.
New Bayesian optimization models for efficient material screening.
problem Efficiently screening materials with expensive and cheap tests.
method Flexible multi-test Bayesian optimization models with complex relationships.
result Demonstrated power on synthetic and real data.
This paper treats the problem of screening for variables with high correlations in high dimensional data in which there can be many fewer samples than variables. We focus on threshold-based correlation screening methods for three related applications: screening for variables with large correlations within a single trea…
New screening rules improve lasso model fitting efficiency.
problem Efficiently solving high-dimensional lasso problems.
method Look-ahead screening rules to discard predictors.
result Look-ahead screening rules outperform existing methods.
A new screening rule 'dynamic Sasvi' improves sparse optimization speed.
problem Sparse optimization problem identification.
method Flexible framework based on Fenchel-Rockafellar duality for norm-regularized least squares.
result Dynamic Sasvi can eliminate more features and increase solver speed.
Study on lightlike submanifolds in metallic semi-Riemannian manifolds.
problem Characterizing and investigating properties of lightlike submanifolds in metallic semi-Riemannian manifolds.
method Introduced and analyzed subclasses of screen transversal lightlike submanifolds and investigated their geometric properties.
result Necessary and sufficient condition for an isotropic screen transversal lightlike submanifold to be totally geodesic.
The paper studies special null hypersurfaces in spacetimes.
problem Characterizing null screen isoparametric hypersurfaces in Lorentzian space forms.
method Developed screen isoparametric hypersurface concept for null hypersurfaces of Robertson-Walker spacetimes, derived Cartan identities, and provided local characterizations.
result Derived Cartan identities for the screen principal curvatures of null screen hypersurfaces in Lorentzian space forms and provided a local characterization.
New AI platform screens portfolios for desirable firms and news.
problem Optimizing portfolio selection with AI.
method Two LLM agents screen for firm fundamentals and news sentiment. Agents deliberate to generate buy/sell signals. High-dimensional estimation determines optimal weights.
result Screened portfolio's Sharpe ratio consistently estimates target, superior to baseline and conventional approaches.
New screening test for LASSO reduces complexity.
problem Efficient screening for LASSO problems.
method Joint screening test for LASSO problem, applied to sphere and dome regions.
result Effective screening of atoms reduces computational complexity.
Recent computational strategies based on screening tests have been proposed to accelerate algorithms addressing penalized sparse regression problems such as the Lasso. Such approaches build upon the idea that it is worth dedicating some small computational effort to locate inactive atoms and remove them from the dictio…
A variable screening procedure via correlation learning was proposed Fan and Lv (2008) to reduce dimensionality in sparse ultra-high dimensional models. Even when the true model is linear, the marginal regression can be highly nonlinear. To address this issue, we further extend the correlation learning to marginal nonp…
A new screening method for high-dimensional data reduces computational cost.
problem Challenges in variable selection for ultrahigh-dimensional linear regression.
method Ordering absolute sample ridge partial correlations to screen variables.
result The method provides sure screening property without strong assumptions.
A new screening rule improves lasso solving speed.
problem Efficiently solving lasso problems with high correlation.
method Uses second-order information from the Hessian to screen predictors.
result Outperforms alternatives on simulated and real data.
Recently, to solve large-scale lasso and group lasso problems, screening rules have been developed, the goal of which is to reduce the problem size by efficiently discarding zero coefficients using simple rules independently of the others. However, screening for overlapping group lasso remains an open challenge because…
Study the geometry of specific submanifolds in Golden Semi-Riemannian manifolds.
problem Investigate the geometry of specific submanifolds in Golden Semi-Riemannian manifolds.
method Investigate the geometry of distributions and induced connections, provide necessary and sufficient conditions for metric connections, and characterize submanifolds.
result Characterization of screen transversal anti-invariant lightlike submanifolds of Golden Semi-Riemannian manifolds.
Safe sample screening improves RSVM performance without sacrificing accuracy.
problem Improving RSVM performance under noisy conditions.
method Proposed two safe sample screening rules based on CCCP framework for RSVM.
result Significant reduction in computational time for RSVM.
In data sets with many more features than observations, independent screening based on all univariate regression models leads to a computationally convenient variable selection method. Recent efforts have shown that in the case of generalized linear models, independent screening may suffice to capture all relevant feat…
Introduces screening rules for non-convex Lasso problems.
problem Efficiently solving non-convex Lasso problems with theoretical guarantees.
method Iterative majorization-minimization strategy with screening rule.
result Significant computational gain compared to classical methods.
Deep learning predicts breast cancer with high accuracy from patient data.
problem Early detection of breast cancer from patient data.
method Feature selection and k-fold Monte Carlo cross-validation using deep learning.
result Deep learning model effectively distinguishes between cancer and healthy patients.