This paper compares feature selection methods for biomarker discovery in toxicant-treated fish.
problem Choosing the most suitable method for biomarker discovery in toxicant exposure studies.
method Three feature selection methods: SAM, mRMR, and GeoDE are compared.
result Different methods perform better in different cases, requiring dataset-specific decisions.
New methods use network data to find genetic indicators for diseases.
problem Finding genetic indicators for diseases from large, complex data.
method Integrating network data to select features from whole-genome data.
result Methods can identify genetic indicators that work together.
AFTNet uses a network-constrained Weibull model for biomarker discovery.
problem Discovering biomarkers from survival data with correlated predictors.
method Survival analysis method based on Weibull AFT model, incorporating network constraints and penalized likelihood for variable selection.
result Theoretical consistency and efficient algorithm for AFTNet estimator validated on synthetic and real data.
New method integrates network knowledge for better clinical risk prediction and biomarker discovery.
problem Improving predictive ability and interpretability of biomarkers using molecular profiling data.
method Introduces a network-regularized sparse Logistic Regression framework with a new penalty term.
result Demonstrates improved performance in simulated and real data compared to existing methods.
The study ranks biomarkers using mutual information, tackling clinical trial challenges.
problem Disentangling predictive and prognostic biomarkers in clinical trials.
method Formalizing biomarker ranking as optimizing mutual information, estimating conditional mutual information terms with efficient approximations and empirical Bayes, introducing a visualisation tool.
result Efficient methods to rank predictive and prognostic biomarkers, visualisation tool for biomarker discovery.
This research uses cooperative game theory to interpret deep learning models for ASD biomarker discovery.
problem Understanding image features used by deep learning models for ASD biomarker discovery.
method Shapley value explanation (SVE) from cooperative game theory applied to deep learning models with graph structure optimization.
result SVE provides more accurate biomarker importance than traditional methods.
ROOFS helps researchers select robust biomarker features from complex data.
problem Challenges in feature selection for biomarker discovery and clinical models.
method ROOFS is a Python package that benchmarks multiple feature selection methods on user data.
result ROOFS identifies a filter method as optimal for identifying predictors of lung cancer resistance.
Deep learning aligns GC-MS peaks for biomarker discovery.
problem Aligning retention times of GC-MS peaks across different samples.
method ChromAlignNet, a deep learning model for peak alignment.
result ChromAlignNet outperforms existing methods on complex data sets.
Feature selection is among the most important components because it not only helps enhance the classification accuracy, but also or even more important provides potential biomarker discovery. However, traditional multivariate methods is likely to obtain unstable and unreliable results in case of an extremely high dimen…
The study discovers digital biomarkers for Parkinson's Disease using optimized transitions and emissions in HSMM.
problem Identifying digital biomarkers for Parkinson's Disease.
method Proposed a Hidden Semi-Markov Model (HSMM) to model Parkinson's Disease patients' step and stride periodic cycles.
result The HSMM allows for more informative characterization of Parkinson's Disease patients/controls by considering the duration spent in each state.
OBF optimally filters features under independent Gaussian models.
problem Biomarker discovery from complex data.
method Optimal Bayesian feature selection under independent Gaussian models.
result OBF is consistent and optimal under mild conditions.
RobKMR improves robustness in multi-omics data analysis for osteoporosis biomarker discovery.
problem Sensitivity to adversarial outliers and lack of comprehensive multi-omics data integration.
method RobKMR, a non-linear M-estimator-based approach using robust kernel centered Gram matrix and robust score test.
result Selected biomarkers (DKK1, MTND5, FASTKD2) significantly bond with four drugs for osteoporosis.
Background: Predictive, stable and interpretable gene signatures are generally seen as an important step towards a better personalized medicine. During the last decade various methods have been proposed for that purpose. However, one important obstacle for making gene signatures a standard tool in clinics is the typica…
New biomarkers for autism detected from R-fMRI data.
problem Challenges in extracting functional biomarkers from multi-site R-fMRI data for complex neuropsychiatric disorders.
method Developed pipelines to extract participant-specific connectomes from functionally-defined brain areas, compared across participants, and predicted neuropsychiatric status.
result 67% prediction accuracy on ABIDE dataset, significantly better than previous results.
A method for biomarker selection using aggregated data under data protection constraints.
problem Manual data exchange and limited data calls hinder joint analyses of clinical biomarkers.
method Distributed multivariate regression modeling for automatic variable selection.
result The heuristic variant reduces data calls from over 10 to 3, making manual data releases feasible.
engGNN combines external and generated graphs to improve disease classification and biomarker discovery.
problem Challenges in integrating omics data due to high dimensionality and small sample sizes.
method Dual-graph framework that integrates external biological networks with data-driven generated graphs.
result engGNN outperforms state-of-the-art methods in disease classification and biomarker discovery.
MKL-based models outperform complex multi-omics integrative approaches.
problem Integrating diverse omics data sources.
method Supervised multiple kernel learning with different kernel fusion strategies.
result MKL-based models outperform more complex architectures.
Stein-Encoder isolates genetic signals in multi-modal biomedical data.
problem Integration of high-dimensional genomic data with clinical data obscures genetic predictive impact.
method White-box supervised framework using Stein's method and residualization.
result Stein-Encoder improves predictive accuracy and reveals specific biological mechanisms.
PKB framework boosts genomic data analysis by integrating pathway knowledge.
problem Boosting discovery power and connecting new findings with biological mechanisms in genomic data.
method Pathway-based Kernel Boosting (PKB) framework integrating clinical and pathway information for prediction of various outcomes.
result PKB substantially outperforms other methods in predicting drug response and cancer survival.
New method detects biomarker-treatment interactions in clinical trials.
problem Detecting interactions between high-dimensional biomarkers and treatments in randomized trials.
method Two-stage penalized regression screening using ridge regression for multivariate screening.
result Ridge regression screening provides greater power than traditional methods in correlated data.
Optimizes biomarker selection for cost-effective treatment rules.
problem Incorporating multiple biomarkers in treatment selection rules can be costly and reduce model performance.
method Developed procedures for estimating linear and nonlinear combinations of biomarkers using 0-norm penalized weighted classification.
result Demonstrated the importance of feature selection and marker cost in treatment selection rules.
A method to explain disease transformation using biomarker covariance matrices.
problem Understanding disease transformation from a healthy baseline.
method Modeling healthy and disease states of biomarker covariance matrices to characterize perturbations.
result Disease perturbs the biomarker covariance structure, allowing for mechanistic explanations and individual patient prognosis.
Neural network models improve ROC curve evaluation of biomarkers, focusing on age's role in physical activity-mortality association.
problem Improving biomarker evaluation using machine learning for complex relationships.
method Proposes neural network-based covariate-adjusted ROC modeling.
result Age has distinct effects on mortality outcomes when physical activity is measured as total activity time.
New DTI model using self-attention molecule representation outperforms state-of-the-art.
problem Predicting drug-target interactions to reduce costs and improve personalized medicine.
method Proposes a new molecule representation using self-attention and a new DTI model.
result Our DTI model outperforms state-of-the-art by up to 4.9% points in precision-recall.
New biomarker predicts MRgFUS treatment outcome without contrast agents.
problem Inaccurate assessment of treated tissue viability after MRgFUS.
method Deep learning on noncontrast multiparametric MRI images, voxel-wise registration.
result Predicted follow-up NPV with DICE coefficient 0.71, outperforming current standard.
RIF prioritizes predictive biomarkers for precision medicine.
problem Lack of tools to select and prioritize predictive biomarkers.
method Random Interaction Forest (RIF) method.
result RIF outperformed conventional methods in various simulation scenarios and clinical trials.
Super-resolution improves MRI resolution and accuracy for biomarker assessment.
problem Inadequate SNR for accurate quantification in high-resolution MRI.
method Utilized deep learning super-resolution to maintain SNR for T2 relaxation time biomarkers while generating high-resolution images.
result Super-resolution successfully maintains high-resolution and accurate biomarkers for MRI.
Graph Neural Network identifies ASD biomarkers from fMRI data.
problem Finding biomarkers for Autism Spectrum Disorder (ASD).
method Graph Neural Network (GNN) for analyzing task-fMRI brain networks, 2-stage pipeline to interpret feature importance.
result GNN achieves high accuracy in identifying ASD biomarkers and reveals their association with social behaviors.
DKT transfers biomarker information between neurodegenerative diseases.
problem Estimating biomarker trajectories in rare neurodegenerative diseases with limited data.
method DKT is a joint-disease generative model that transfers biomarker progressions from common neurodegenerative diseases to rare ones.
result DKT estimates plausible multimodal biomarker trajectories in rare diseases like PCA using only unimodal data.
The Set Covering Machine (SCM) is a greedy learning algorithm that produces sparse classifiers. We extend the SCM for datasets that contain a huge number of features. The whole genetic material of living organisms is an example of such a case, where the number of feature exceeds 10^7. Three human pathogens were used to…
New method identifies predictive biomarkers for subgroup analysis.
problem Identifying predictive biomarkers from large covariates.
method Generalized penalized regression with overlapped group penalties.
result Asymptotically consistent method for sparse, interpretable models.
Method predicts biomarker trajectories with uncertainty bands for Alzheimer's disease.
problem Uncertainty in biomarker predictions poses risks in clinical deployment.
method Conformal prediction for randomly-timed biomarker trajectories.
result Conformal bands achieve desired coverage and are tighter than baseline.
Gaussian OBFS proves strong consistency in feature selection with correlations.
problem Feature selection consistency in the presence of correlations.
method Proves strong consistency of Gaussian OBFS under mild conditions.
result Identifies selected features and rates of convergence for different feature types.
Proposes ConRad model for lung cancer classification using radiomics and interpretable machine learning.
problem Lack of interpretability in deep neural networks for cancer diagnosis.
method Integration of radiomics and DNN-predicted biomarkers in interpretable classifiers (ConRad).
result ConRad models outperform CNNs in five-fold cross-validation.
For mass spectra acquired from cancer patients by MALDI or SELDI techniques, automated discrimination between cancer types or stages has often been implemented by machine learnings. These techniques typically generate "black-box" classifiers, which are difficult to interpret biologically. We develop new and efficient s…
EBM uses high-dimensional imaging biomarkers to improve dementia progression estimation.
problem Current EBMs only use scalar biomarkers, limiting accuracy from cross-sectional data.
method Proposes nDEBM, a novel method using semi-supervised SVM on voxel-wise imaging biomarkers.
result nDEBM outperforms state-of-the-art EBM methods using regional volume biomarkers.
Study identifies biomarkers for lung cancer in female non-smokers.
problem Identifying prognostic biomarkers for stage III NSCLC in non-smoking females.
method Gene expression profiling and XGBoost machine learning algorithm.
result Top biomarkers validated in literature, with AUC score of 0.835.
Novel framework predicts brain biomarker trajectories with superior performance.
problem Challenges in estimating longitudinal brain biomarker trajectories due to variability, inconsistencies, and irregular measurements.
method Personalized deep kernel regression with Adaptive Shrinkage Estimation.
result Superior predictive performance compared to state-of-the-art models.
PR-GNN identifies salient brain regions for ASD biomarkers.
problem Identifying brain regions associated with neurological disorders.
method Pooling Regularized Graph Neural Network (PR-GNN) with novel salient region selection.
result PR-GNN outperforms baseline methods in ASD classification accuracy.
New method uses SHAP for biomarker identification in CATE models.
problem Identifying predictive biomarkers from observational data.
method Surrogate estimation approach using SHAP values for CATE meta-learners.
result SHAP accurately identifies biomarkers in high-dimensional data.
Motivation: Biomarker discovery from high-dimensional data is a crucial problem with enormous applications in biology and medicine. It is also extremely challenging from a statistical viewpoint, but surprisingly few studies have investigated the relative strengths and weaknesses of the plethora of existing feature sele…
The paper develops Gaussian Process models for Alzheimer's disease biomarkers.
problem Designing accurate biomarkers for Alzheimer's disease diagnosis.
method Gaussian Processes, Multiple Kernel Learning, Automatic Relevance Determination.
result Gaussian Process models outperform Support Vector Machines in classification performance.
Study uses LLMs to create personalized treatment plans for rare gynecological tumors.
problem Suboptimal management and poor prognosis due to low incidence and heterogeneity of rare gynecological tumors.
method Developed a digital twin system using LLMs to integrate clinical and biomarker data.
result LLM-enabled digital twins efficiently model individual patient trajectories and identify potential treatment options.
Transformers improve Alzheimer's disease progression prediction by accounting for irregular biomarker histories.
problem Difficult prediction of medium-horizon Alzheimer's disease progression due to tied clinical scores and irregular biomarker observations.
method Developed a residual gap-aware transformer that combines statistical reference with transformer-based residual learning.
result The proposed model reduces mean error and improves prediction-observation correlation compared to baseline models.
Wide and deep neural network predicts Alzheimer's progression from shape and clinical data.
problem Predicting Alzheimer's disease progression from shape and clinical data.
method Fused anatomical shape and tabular clinical data in a neural network, employing survival analysis loss.
result The model outperforms shape and clinical models individually.
Proposes LSTM algorithm for robust Alzheimer's disease progression modeling with missing data.
problem Challenges in modeling disease progression using incomplete longitudinal data.
method Utilizes Long Short-Term Memory (LSTM) networks for Alzheimer's disease progression modeling with a generalized training rule for handling missing data.
result Achieves significantly lower mean absolute error (MAE) than alternatives with p < 0.05.
A new framework models multi-state events and biomarkers.
problem Limited representation of complex multi-state trajectories.
method General multi-state joint modeling framework.
result Accurate parameter recovery and personalized predictions.
This study interprets machine learning models to identify biomarkers for severe COVID-19 infection.
problem The black-box nature of machine learning models makes it difficult for medical researchers to understand and trust their predictions.
method The study uses permutation feature importance, Partial Dependence Plot, Individual Conditional Expectation, Accumulated Local Effects, Local Interpretable Model-agnostic Explanations, and Shapley Additive Explanation to interpret four machine learning models.
result The study identifies NTproBNP, CRP, LDH, LYM, leukocytes, eosinophils, and platelets as biomarkers associated with severe COVID-19 infection.