Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

10.0%20.0%30.0%40.0% · Aug 199419922001200920182026
48 results for Cancer types

Deep learning models improve cancer detection and typing classification from gene expression data.

problem Challenges in establishing specificity for cancer diagnosis using gene expression data.
method Developed deep learning models using mRNA datasets for cancer detection and typing classification.
result Achieved 98% accuracy in cancer detection and 18 out of 32 cancer-typing classifications over 90% accuracy.

We apply our statistically deterministic machine learning/clustering algorithm *K-means (recently developed in https://ssrn.com/abstract=2908286) to 10,656 published exome samples for 32 cancer types. A majority of cancer types exhibit mutation clustering structure. Our results are in-sample stable. They are also out-o…

2017-07-26abs ↗pdf ↗

Machine learning models improve cancer type classification accuracy.

problem Early and accurate cancer diagnosis is challenging due to high costs and biological marker limitations.
method Assessed five machine learning algorithms for 17 cancer types using RNA-seq data.
result Ensemble algorithms achieve 100% accuracy in 14 out of 17 cancer types.

L-Perceptron improves breast cancer diagnosis and survival prediction.

problem Improving early prognosis and survival prediction rates for breast cancer.
method Proposes a novel type of perceptron (L-Perceptron) for better accuracy and sensitivity.
result Achieves 97.42% and 98.73% accuracy and sensitivity in Wisconsin Breast Cancer dataset.

Paper proposes scalable method for analyzing multi-omic data.

problem Integrating high-dimensional multi-omic data for cancer subtyping.
method Mixed graphical model approach using Birth-Death MCMC algorithm.
result Our method outperforms LASSO and standard BDMCMC in computational efficiency and model selection accuracy.

New model identifies cell-specific genes for cancer prognosis.

problem No statistical model to integrate multiscale cancer data.
method Bayesian generalized promotion time cure models (GPTCMs).
result Improves cancer prognosis by identifying cell-specific genes.

Data mining techniques predict breast cancer types with high accuracy.

problem Early detection of breast cancer to reduce mortality rates.
method Twelve classification algorithms applied to the Breast Cancer Wisconsin dataset.
result High accuracy in predicting malignant and benign breast cancer.

BIDIFAC+ factorizes linked matrices for cancer studies.

problem Integrating multiple omics platforms across various cancer types.
method Flexible approach to simultaneous factorization and decomposition of linked matrices using BIDIFAC+.
result Identifies shared and specific modes of variability across multiple omics platforms and cancer types.

MINN-SA enhances cancer detection using TCR sequences with better interpretability.

problem Challenges in detecting cancers using TCR sequences due to one-to-many correspondence.
method Multiple Instance Neural Networks based on Sparse Attention (MINN-SA).
result MINN-SA achieves highest AUC scores on 10 cancer types compared to existing MIL approaches.

System automates identification of cancer drug repurposing from PubMed.

problem Manual extraction of cancer drug repurposing evidence from scientific publications is infeasible.
method NLP pipeline including querying, filtering, entity extraction, classification, and study type classification.
result Automated system extracts cancer drug repurposing evidence from PubMed abstracts.

Machine learning accurately diagnoses cancer from whole genome sequencing data.

problem Accurate cancer diagnosis at all stages.
method Novel MLAC (Machine Learning Against Cancer) method using next-gen RNA sequencing.
result Perfect precision, sensitivity, and specificity achieved for most tumor types.

We present *K-means clustering algorithm and source code by expanding statistical clustering methods applied in https://ssrn.com/abstract=2802753 to quantitative finance. *K-means is statistically deterministic without specifying initial centers, etc. We apply *K-means to extracting cancer signatures from genome data w…

2017-03-02abs ↗pdf ↗

A new method interprets multiple kernel learning for cancer subtypes.

problem Hard to evaluate clustering results in multidimensional diseases like cancer.
method Combines feature clustering with multiple kernel dimensionality reduction.
result Identifies integrative patient subtypes and explains feature importance.

Neural networks improve cancer risk prediction from family history data.

problem Improving cancer risk prediction from family history data using machine learning.
method Developed and trained neural network models on large pedigrees to predict hereditary cancers.
result Neural networks can achieve nearly optimal prediction performance and outperform traditional models in misreported data.

Modeling correlated mutations in cancer for personalized treatment.

problem Identifying mutations for personalized cancer therapy in heterogeneous profiles.
method Proposed correlated zero-inflated negative binomial process with mixed beta-Bernoulli and variational inference.
result Identified biologically relevant correlations between somatic mutations.

This study predicts ovarian cancer from cysts using TVUS and machine learning.

problem Early detection of ovarian cancer from cysts using TVUS screening.
method Employed Random Forest, KNN, and XGBoost machine learning techniques on PLCO dataset.
result Achieved high accuracy, recall, f1 score, and precision in predicting ovarian cancer.

Algorithm segments glandular structures in colon histology images for cancer grading.

problem Manual gland segmentation is time-consuming and risky for patients.
method Local intensity and texture features, Random Forest classifier, multilevel approach.
result Fast, accurate automatic gland segmentation for clinical use.

Deep neural network for cancer classification using autoencoders.

problem Cancer classification using molecular information.
method Using a Denoising Autoencoder (DAE) as weight initialization for a deep neural network, comparing two approaches: fixed weights and fine-tuning. Embedding strategies included encoding layers and complete autoencoder.
result Best F1 score of 98.04% for identifying thyroid cancer samples.

Develops a feature selection method for multi-view data with mixed types.

problem Challenges in feature selection for high-dimensional multi-view data with mixed data types.
method Block Randomized Adaptive Iterative Lasso (B-RAIL) combining randomized Lasso, adaptive weighting, and stability selection.
result Demonstrates effectiveness of B-RAIL in identifying biomarkers and novel candidates for ovarian cancer.

Bayesian model learns cancer subtypes from diverse NGS data.

problem Overdispersed NGS count data and limited samples for specific cancer types.
method Bayesian Multi-Domain Learning (BMDL) model using hierarchical negative binomial factorization.
result BMDL achieves reproducible cancer subtyping without negative transfer effects.

For mass spectra acquired from cancer patients by MALDI or SELDI techniques, automated discrimination between cancer types or stages has often been implemented by machine learnings. These techniques typically generate "black-box" classifiers, which are difficult to interpret biologically. We develop new and efficient s…

2014-10-13abs ↗pdf ↗

AI detects oral pre-cancerous lesions with high accuracy.

problem Manual screening of oral cavity cancer is expensive and lacks specialists.
method Deep convolutional neural networks (DCNNs) using transfer learning.
result DCNN models achieve high accuracy in distinguishing between benign and pre-cancerous tongue lesions.

New dataset and approach improve skin cancer detection accuracy.

problem Lack of patient clinical information in automated skin cancer detection.
method Introduced a new dataset with clinical images and patient information. Combined clinical data with dermoscopy images using deep learning models.
result Combining clinical data improves skin cancer detection accuracy by around 7%.

The paper investigates interpretability techniques for deep learning models in medical data.

problem Understanding the logic behind predictions of black-box models in medical decision-making.
method Applied deep neural networks and random forests to a medical dataset. Used autoencoders and local interpretable models to provide insights.
result Local interpretable models and autoencoders provide meaningful insights into cancer predictions, identifying distinct and non-generalizable features.

Omics-GAN uses GANs to generate synthetic multi-omics data for improved disease prediction.

problem Limited sample sizes, noise, and heterogeneity in multi-omics data reduce predictive power.
method Omics-GAN is a GAN-based framework that generates high-quality synthetic multi-omics profiles.
result Synthetic datasets consistently improved prediction accuracy compared to original omics profiles.

Robust cancer screening model using pre-trained ensembles for biomarkers.

problem Detecting early-stage cancer, especially in hard-to-diagnose cases like pancreatic cancer.
method Meta-trained Hyperfast model for robust classification, combined with ensembling of XGBoost and LightGBM.
result Achieved highest AUC of 0.9929 and robust performance on imbalanced datasets.

Exclusive Lasso improves survival prediction in cancer datasets.

problem Enhanced survival prediction in cancer datasets with high-dimensional genomic and clinical data.
method Proposes Exclusive Lasso regularization for feature selection in Cox regression models for grouped variables.
result Demonstrates improved survival prediction performance using Exclusive Lasso compared to standard Cox regression.

Bayesian neural networks improve cancer dynamics prediction.

problem Predicting cancer dynamics under treatment due to heterogeneity and sparse data.
method Hierarchical Bayesian model using baseline covariates and Bayesian neural networks for nonlinear interactions.
result Bayesian neural networks outperform linear models in predicting cancer dynamics with interactions.

Study compares single vs ensemble feature selection for cancer diagnosis.

problem Identifying relevant variables for cancer diagnosis and prognosis.
method Comparison of single feature selection algorithms and ensemble of diverse algorithms.
result Ensemble approach did not improve predictive performance over individual algorithms.

Robust method estimates self-similarity for mammogram images, improving cancer detection.

problem Statistical assessment of self-similarity in real data with large mean level shifts.
method Theil-type weighted regression for wavelet-based estimation, compared to OLS and AV.
result Robust approach shows nearly 68% accuracy in cancer vs non-cancer classification.

Novel topic modeling approach improves survival prediction for cancer patients.

problem Survival prediction for cancer patients using high-dimensional gene expression data.
method Inspired by topic modeling, a novel methodology using discretized Latent Dirichlet Allocation (dLDA) to derive expressive features from gene expression data.
result Survival estimates are more accurate than standard models, as shown by the Concordance measure.