Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

5101419 · Nov 201919922001200920182026
48 results for Cancer biomarkers

Study identifies biomarkers for lung cancer in female non-smokers.

problem Identifying prognostic biomarkers for stage III NSCLC in non-smoking females.
method Gene expression profiling and XGBoost machine learning algorithm.
result Top biomarkers validated in literature, with AUC score of 0.835.

Robust cancer screening model using pre-trained ensembles for biomarkers.

problem Detecting early-stage cancer, especially in hard-to-diagnose cases like pancreatic cancer.
method Meta-trained Hyperfast model for robust classification, combined with ensembling of XGBoost and LightGBM.
result Achieved highest AUC of 0.9929 and robust performance on imbalanced datasets.

Proposes ConRad model for lung cancer classification using radiomics and interpretable machine learning.

problem Lack of interpretability in deep neural networks for cancer diagnosis.
method Integration of radiomics and DNN-predicted biomarkers in interpretable classifiers (ConRad).
result ConRad models outperform CNNs in five-fold cross-validation.

A method to explain disease transformation using biomarker covariance matrices.

problem Understanding disease transformation from a healthy baseline.
method Modeling healthy and disease states of biomarker covariance matrices to characterize perturbations.
result Disease perturbs the biomarker covariance structure, allowing for mechanistic explanations and individual patient prognosis.

TACOMA improves cancer biomarker validation by incorporating deep features.

problem Improving accuracy and repeatability in TMA image scoring.
method Incorporating deep learning representations learned through unsupervised clustering and recursive space partitioning.
result Reduced error rate by about 6% on breast cancer TMA images.

Bayesian Cox model identifies biomarkers from multi-omics data.

problem Produce interpretable survival prognosis from multi-omics data.
method Penalized semiparametric Bayesian Cox model with graph-structured selection priors.
result Model identifies new biomarkers and improves survival prediction.

TransST improves spatial transcriptomics data analysis by identifying cell clusters and biomarkers.

problem Low resolution and insufficient sequencing depth in spatial transcriptomics data.
method Transfer learning framework to adaptively leverage external cell-labeled information.
result TransST successfully identifies five biologically meaningful cell clusters and separates adipose tissues from connective issues.

In this paper we propose network methodology to infer prognostic cancer biomarkers based on the epigenetic pattern DNA methylation. Epigenetic processes such as DNA methylation reflect environmental risk factors, and are increasingly recognised for their fundamental role in diseases such as cancer. DNA methylation is a…

2015-06-17abs ↗pdf ↗

Network Elastic Net identifies smoking-specific gene expression for lung cancer prognosis.

problem Identifying smoking-specific gene expression biomarkers in lung cancer prognosis.
method Introduces Network Elastic Net, a method that clusters and regresses on graphs based on smoking behavior.
result Shows efficacy of clusters in identifying cancer stages using gene expression and smoking behavior.

Paper proposes inference method for high-dimensional censored quantile regression.

problem Identifying heterogeneous effects of high-dimensional genetic biomarkers on survival outcomes.
method Combines low-dimensional model estimates based on multi-sample splittings and variable selection.
result Proposed estimator is consistent and asymptotically follows a Gaussian process.

For mass spectra acquired from cancer patients by MALDI or SELDI techniques, automated discrimination between cancer types or stages has often been implemented by machine learnings. These techniques typically generate "black-box" classifiers, which are difficult to interpret biologically. We develop new and efficient s…

2014-10-13abs ↗pdf ↗

PKB framework boosts genomic data analysis by integrating pathway knowledge.

problem Boosting discovery power and connecting new findings with biological mechanisms in genomic data.
method Pathway-based Kernel Boosting (PKB) framework integrating clinical and pathway information for prediction of various outcomes.
result PKB substantially outperforms other methods in predicting drug response and cancer survival.

Study uses LLMs to create personalized treatment plans for rare gynecological tumors.

problem Suboptimal management and poor prognosis due to low incidence and heterogeneity of rare gynecological tumors.
method Developed a digital twin system using LLMs to integrate clinical and biomarker data.
result LLM-enabled digital twins efficiently model individual patient trajectories and identify potential treatment options.

Study improves cancer classification using gene selection and projection methods.

problem Overfitting in high-dimensional microarray datasets for cancer classification.
method FSWOR technique, random projection, Kendall test, ensemble classifiers, LDA projection, Naïve Bayes.
result Achieved a test score of 96%, significantly outperforming existing methods.

Omics-GAN uses GANs to generate synthetic multi-omics data for improved disease prediction.

problem Limited sample sizes, noise, and heterogeneity in multi-omics data reduce predictive power.
method Omics-GAN is a GAN-based framework that generates high-quality synthetic multi-omics profiles.
result Synthetic datasets consistently improved prediction accuracy compared to original omics profiles.

Paper proposes clustering model for ICC based on histologic patterns.

problem Challenges in grading rare cancers like ICC due to small sample sizes and difficulty in extracting patterns.
method Unsupervised deep convolutional autoencoder clustering model trained on 246 ICC digitized slides.
result Three clusters significantly associated with recurrence-free survival in Cox-proportional hazard models.

Bayesian neural networks improve cancer dynamics prediction.

problem Predicting cancer dynamics under treatment due to heterogeneity and sparse data.
method Hierarchical Bayesian model using baseline covariates and Bayesian neural networks for nonlinear interactions.
result Bayesian neural networks outperform linear models in predicting cancer dynamics with interactions.

Develops a feature selection method for multi-view data with mixed types.

problem Challenges in feature selection for high-dimensional multi-view data with mixed data types.
method Block Randomized Adaptive Iterative Lasso (B-RAIL) combining randomized Lasso, adaptive weighting, and stability selection.
result Demonstrates effectiveness of B-RAIL in identifying biomarkers and novel candidates for ovarian cancer.

ROOFS helps researchers select robust biomarker features from complex data.

problem Challenges in feature selection for biomarker discovery and clinical models.
method ROOFS is a Python package that benchmarks multiple feature selection methods on user data.
result ROOFS identifies a filter method as optimal for identifying predictors of lung cancer resistance.

Deep object detection improves mitotic nucleus detection in breast cancer biopsies.

problem Challenges in automated mitotic nucleus detection in breast cancer histopathological images.
method Adapted Mask R-CNN for deep object detection, initially selects candidate regions with maximum recall, refines them with multi-object loss function.
result Improved discrimination ability (F-score of 0.86) and significant precision (0.86) for mitotic nuclei compared to two-stage models.

fiBAG integrates multiplatform genomic data to identify disease markers.

problem Understanding complex mechanisms underlying human diseases from multiplatform genomic data.
method fiBAG uses Gaussian process models and Bayes factors to identify functional evidence and guide variable selection.
result fiBAG improves detection of disease-related markers compared to non-integrative methods.

MSB framework improves survival prediction in immunotherapy patients with missing data.

problem High dimensionality and blockwise missingness in multimodal clinical data.
method MSB is a late-fusion framework that independently models modality-specific features before aggregating predictions via cross-validated stacking.
result MSB outperformed baseline algorithms in predicting progression-free survival in lung cancer patients.

EAGLE-Net enhances foundation models by integrating patch-level features for better tissue understanding.

problem Foundation models lack mechanisms for global tissue structure and local context in computational pathology.
method EAGLE-Net combines multi-scale spatial encoding, attention-guided loss functions, and background suppression to aggregate patch-level features into slide-level predictions.
result EAGLE-Net improves classification accuracy and concordance indices across multiple cancer types, producing biologically coherent attention maps.

Modern bio-technologies have produced a vast amount of high-throughput data with the number of predictors far greater than the sample size. In order to identify more novel biomarkers and understand biological mechanisms, it is vital to detect signals weakly associated with outcomes among ultrahigh-dimensional predictor…

2018-05-17abs ↗pdf ↗

The method integrates survival constraints into NMF for identifying survival-associated gene clusters.

problem Understanding and interpreting high-dimensional biological data for disease markers.
method Cox proportional hazards regression integrated with NMF via proportional hazards non-negative matrix factorization.
result The method can uncover survival-associated gene clusters in cancer gene expression data.

We study information theoretic methods for ranking biomarkers. In clinical trials there are two, closely related, types of biomarkers: predictive and prognostic, and disentangling them is a key challenge. Our first step is to phrase biomarker ranking in terms of optimizing an information theoretic quantity. This formal…

2016-12-05abs ↗pdf ↗

New method detects biomarker-treatment interactions in clinical trials.

problem Detecting interactions between high-dimensional biomarkers and treatments in randomized trials.
method Two-stage penalized regression screening using ridge regression for multivariate screening.
result Ridge regression screening provides greater power than traditional methods in correlated data.

Study uses machine learning to identify IBD biomarkers from gut microbiota.

problem Identifying biomarkers for Inflammatory Bowel Disease (IBD) from gut microbiota.
method Ensemble feature selection methods (CMIM, FCBF, mRMR, XGBoost) applied to IBD-associated metagenomics dataset.
result XGBoost minimizes microbiota used for IBD diagnosis, improving classification accuracy.

Optimizes biomarker selection for cost-effective treatment rules.

problem Incorporating multiple biomarkers in treatment selection rules can be costly and reduce model performance.
method Developed procedures for estimating linear and nonlinear combinations of biomarkers using 0-norm penalized weighted classification.
result Demonstrated the importance of feature selection and marker cost in treatment selection rules.

Neural network models improve ROC curve evaluation of biomarkers, focusing on age's role in physical activity-mortality association.

problem Improving biomarker evaluation using machine learning for complex relationships.
method Proposes neural network-based covariate-adjusted ROC modeling.
result Age has distinct effects on mortality outcomes when physical activity is measured as total activity time.

New biomarker predicts MRgFUS treatment outcome without contrast agents.

problem Inaccurate assessment of treated tissue viability after MRgFUS.
method Deep learning on noncontrast multiparametric MRI images, voxel-wise registration.
result Predicted follow-up NPV with DICE coefficient 0.71, outperforming current standard.

Graph Neural Network identifies ASD biomarkers from fMRI data.

problem Finding biomarkers for Autism Spectrum Disorder (ASD).
method Graph Neural Network (GNN) for analyzing task-fMRI brain networks, 2-stage pipeline to interpret feature importance.
result GNN achieves high accuracy in identifying ASD biomarkers and reveals their association with social behaviors.

DKT transfers biomarker information between neurodegenerative diseases.

problem Estimating biomarker trajectories in rare neurodegenerative diseases with limited data.
method DKT is a joint-disease generative model that transfers biomarker progressions from common neurodegenerative diseases to rare ones.
result DKT estimates plausible multimodal biomarker trajectories in rare diseases like PCA using only unimodal data.

New DTI model using self-attention molecule representation outperforms state-of-the-art.

problem Predicting drug-target interactions to reduce costs and improve personalized medicine.
method Proposes a new molecule representation using self-attention and a new DTI model.
result Our DTI model outperforms state-of-the-art by up to 4.9% points in precision-recall.