Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

10.0%20.0%30.0%40.0% · Aug 199419922001200920172026
48 results for cancer types

We apply our statistically deterministic machine learning/clustering algorithm *K-means (recently developed in https://ssrn.com/abstract=2908286) to 10,656 published exome samples for 32 cancer types. A majority of cancer types exhibit mutation clustering structure. Our results are in-sample stable. They are also out-o…

2017-07-26abs ↗pdf ↗

Paper proposes scalable method for analyzing multi-omic data.

problem Integrating high-dimensional multi-omic data for cancer subtyping.
method Mixed graphical model approach using Birth-Death MCMC algorithm.
result Our method outperforms LASSO and standard BDMCMC in computational efficiency and model selection accuracy.

New model identifies cell-specific genes for cancer prognosis.

problem No statistical model to integrate multiscale cancer data.
method Bayesian generalized promotion time cure models (GPTCMs).
result Improves cancer prognosis by identifying cell-specific genes.

BIDIFAC+ factorizes linked matrices for cancer studies.

problem Integrating multiple omics platforms across various cancer types.
method Flexible approach to simultaneous factorization and decomposition of linked matrices using BIDIFAC+.
result Identifies shared and specific modes of variability across multiple omics platforms and cancer types.

MINN-SA enhances cancer detection using TCR sequences with better interpretability.

problem Challenges in detecting cancers using TCR sequences due to one-to-many correspondence.
method Multiple Instance Neural Networks based on Sparse Attention (MINN-SA).
result MINN-SA achieves highest AUC scores on 10 cancer types compared to existing MIL approaches.

System automates identification of cancer drug repurposing from PubMed.

problem Manual extraction of cancer drug repurposing evidence from scientific publications is infeasible.
method NLP pipeline including querying, filtering, entity extraction, classification, and study type classification.
result Automated system extracts cancer drug repurposing evidence from PubMed abstracts.

Machine learning accurately diagnoses cancer from whole genome sequencing data.

problem Accurate cancer diagnosis at all stages.
method Novel MLAC (Machine Learning Against Cancer) method using next-gen RNA sequencing.
result Perfect precision, sensitivity, and specificity achieved for most tumor types.

We present *K-means clustering algorithm and source code by expanding statistical clustering methods applied in https://ssrn.com/abstract=2802753 to quantitative finance. *K-means is statistically deterministic without specifying initial centers, etc. We apply *K-means to extracting cancer signatures from genome data w…

2017-03-02abs ↗pdf ↗

Neural networks improve cancer risk prediction from family history data.

problem Improving cancer risk prediction from family history data using machine learning.
method Developed and trained neural network models on large pedigrees to predict hereditary cancers.
result Neural networks can achieve nearly optimal prediction performance and outperform traditional models in misreported data.

Modeling correlated mutations in cancer for personalized treatment.

problem Identifying mutations for personalized cancer therapy in heterogeneous profiles.
method Proposed correlated zero-inflated negative binomial process with mixed beta-Bernoulli and variational inference.
result Identified biologically relevant correlations between somatic mutations.

This study predicts ovarian cancer from cysts using TVUS and machine learning.

problem Early detection of ovarian cancer from cysts using TVUS screening.
method Employed Random Forest, KNN, and XGBoost machine learning techniques on PLCO dataset.
result Achieved high accuracy, recall, f1 score, and precision in predicting ovarian cancer.

Deep neural network for cancer classification using autoencoders.

problem Cancer classification using molecular information.
method Using a Denoising Autoencoder (DAE) as weight initialization for a deep neural network, comparing two approaches: fixed weights and fine-tuning. Embedding strategies included encoding layers and complete autoencoder.
result Best F1 score of 98.04% for identifying thyroid cancer samples.

For mass spectra acquired from cancer patients by MALDI or SELDI techniques, automated discrimination between cancer types or stages has often been implemented by machine learnings. These techniques typically generate "black-box" classifiers, which are difficult to interpret biologically. We develop new and efficient s…

2014-10-13abs ↗pdf ↗

The paper investigates interpretability techniques for deep learning models in medical data.

problem Understanding the logic behind predictions of black-box models in medical decision-making.
method Applied deep neural networks and random forests to a medical dataset. Used autoencoders and local interpretable models to provide insights.
result Local interpretable models and autoencoders provide meaningful insights into cancer predictions, identifying distinct and non-generalizable features.

Omics-GAN uses GANs to generate synthetic multi-omics data for improved disease prediction.

problem Limited sample sizes, noise, and heterogeneity in multi-omics data reduce predictive power.
method Omics-GAN is a GAN-based framework that generates high-quality synthetic multi-omics profiles.
result Synthetic datasets consistently improved prediction accuracy compared to original omics profiles.

Robust cancer screening model using pre-trained ensembles for biomarkers.

problem Detecting early-stage cancer, especially in hard-to-diagnose cases like pancreatic cancer.
method Meta-trained Hyperfast model for robust classification, combined with ensembling of XGBoost and LightGBM.
result Achieved highest AUC of 0.9929 and robust performance on imbalanced datasets.

Exclusive Lasso improves survival prediction in cancer datasets.

problem Enhanced survival prediction in cancer datasets with high-dimensional genomic and clinical data.
method Proposes Exclusive Lasso regularization for feature selection in Cox regression models for grouped variables.
result Demonstrates improved survival prediction performance using Exclusive Lasso compared to standard Cox regression.

Bayesian neural networks improve cancer dynamics prediction.

problem Predicting cancer dynamics under treatment due to heterogeneity and sparse data.
method Hierarchical Bayesian model using baseline covariates and Bayesian neural networks for nonlinear interactions.
result Bayesian neural networks outperform linear models in predicting cancer dynamics with interactions.

Study compares single vs ensemble feature selection for cancer diagnosis.

problem Identifying relevant variables for cancer diagnosis and prognosis.
method Comparison of single feature selection algorithms and ensemble of diverse algorithms.
result Ensemble approach did not improve predictive performance over individual algorithms.

Robust method estimates self-similarity for mammogram images, improving cancer detection.

problem Statistical assessment of self-similarity in real data with large mean level shifts.
method Theil-type weighted regression for wavelet-based estimation, compared to OLS and AV.
result Robust approach shows nearly 68% accuracy in cancer vs non-cancer classification.

Accurate and robust cell nuclei classification is the cornerstone for a wider range of tasks in digital and Computational Pathology. However, most machine learning systems require extensive labeling from expert pathologists for each individual problem at hand, with no or limited abilities for knowledge transfer between…

2016-06-02abs ↗pdf ↗

EAGLE-Net enhances foundation models by integrating patch-level features for better tissue understanding.

problem Foundation models lack mechanisms for global tissue structure and local context in computational pathology.
method EAGLE-Net combines multi-scale spatial encoding, attention-guided loss functions, and background suppression to aggregate patch-level features into slide-level predictions.
result EAGLE-Net improves classification accuracy and concordance indices across multiple cancer types, producing biologically coherent attention maps.

New framework distinguishes lung cancer subtypes using MALDI mass spectrometry.

problem Distinguishing between adenocarcinoma and squamous cell carcinoma subtypes in lung cancer.
method Supervised topological data analysis on MALDI mass spectrometry imaging data.
result The proposed framework successfully classifies lung cancer subtypes with competitive results.