Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

481115 · Sep 201919922001200920172026
48 results for cancer subtyping

Deep learning model explains breast cancer subtypes using logistic regression.

problem Clarifying the mechanisms of breast cancer subtypes for better treatment.
method Developed a PWL model that generates custom-made logistic regression for each patient.
result The PWL model reveals genes relevant to cell cycle-related pathways.

New framework distinguishes lung cancer subtypes using MALDI mass spectrometry.

problem Distinguishing between adenocarcinoma and squamous cell carcinoma subtypes in lung cancer.
method Supervised topological data analysis on MALDI mass spectrometry imaging data.
result The proposed framework successfully classifies lung cancer subtypes with competitive results.

Paper proposes scalable method for analyzing multi-omic data.

problem Integrating high-dimensional multi-omic data for cancer subtyping.
method Mixed graphical model approach using Birth-Death MCMC algorithm.
result Our method outperforms LASSO and standard BDMCMC in computational efficiency and model selection accuracy.

Bayesian model clusters diverse 'omics data for disease subtyping.

problem Clustering diverse 'omics datasets conflates multiple structures.
method Multi-view Bayesian mixture model with semi-supervised learning.
result Identifies distinct clusters of patients for stratified medicine.

We present a novel method for extracting cancer signatures by applying statistical risk models (http://ssrn.com/abstract=2732453) from quantitative finance to cancer genome data. Using 1389 whole genome sequenced samples from 14 cancers, we identify an "overall" mode of somatic mutational noise. We give a prescription …

2016-04-29abs ↗pdf ↗

Study identifies biomarkers for lung cancer in female non-smokers.

problem Identifying prognostic biomarkers for stage III NSCLC in non-smoking females.
method Gene expression profiling and XGBoost machine learning algorithm.
result Top biomarkers validated in literature, with AUC score of 0.835.

VICatMix clusters categorical biomedical data efficiently and selects relevant variables.

problem Efficient clustering of high-dimensional categorical biomedical data.
method Variational Bayesian finite mixture model with variational inference.
result Improves clustering accuracy and variable selection on noisy, high-dimensional data.

Many researches demonstrated that the DNA methylation, which occurs in the context of a CpG, has strong correlation with diseases, including cancer. There is a strong interest in analyzing the DNA methylation data to find how to distinguish different subtypes of the tumor. However, the conventional statistical methods …

2018-08-02abs ↗pdf ↗

Multi-view data, that is matched sets of measurements on the same subjects, have become increasingly common with advances in multi-omics technology. Often, it is of interest to find associations between the views that are related to the intrinsic class memberships. Existing association methods cannot directly incorpora…

2018-11-20abs ↗pdf ↗

OPAL optimizes labeling strategy for precise inference from uncertain models.

problem Inference from uncertain machine learning models is brittle.
method OPAL learns a smooth policy to adaptively label data points based on model uncertainty.
result OPAL yields estimators with the lowest variance and achieves nominal coverage in finite samples.

The task of clustering a set of objects based on multiple sources of data arises in several modern applications. We propose an integrative statistical model that permits a separate clustering of the objects for each data source. These separate clusterings adhere loosely to an overall consensus clustering, and hence the…

2013-02-28abs ↗pdf ↗

Many real-world tasks such as classification of digital histopathology images and 3D object detection involve learning from a set of instances. In these cases, only a group of instances or a set, collectively, contains meaningful information and therefore only the sets have labels, and not individual data instances. In…

2019-11-18abs ↗pdf ↗

Integrative analysis of disparate data blocks measured on a common set of experimental subjects is a major challenge in modern data analysis. This data structure naturally motivates the simultaneous exploration of the joint and individual variation within each data block resulting in new insights. For instance, there i…

2017-04-07abs ↗pdf ↗

Study identifies five AD subtypes using graph diffusion and similarity learning.

problem Identifying homogeneous AD subtypes to improve diagnosis and treatment.
method Unsupervised clustering with graph diffusion and similarity learning.
result Five distinct AD subtypes identified with significant differences in biomarkers and clinical features.

Efficient algorithm for Bayesian networks reduces marginal probability distribution computation.

problem Exact computation of marginal probability distribution is NP-hard for categorical variables in Bayesian networks.
method Divide-and-conquer approach exploiting graphical properties of Bayesian networks.
result Novel algorithm outperforms state-of-the-art methods in classification and cancer subtype identification.

The linking genotype to phenotype is the fundamental aim of modern genetics. We focus on study of links between gene expression data and phenotype data through integrative analysis. We propose three approaches. 1) The inherent complexity of phenotypes makes high-throughput phenotype profiling a very difficult and labor…

2015-06-29abs ↗pdf ↗

While developing their software, professional object-oriented (OO) software developers keep in their minds an image of the subtyping relation between types in their software. The goal of this paper is to present an observation about the graph of the subtyping relation in Java, namely the observation that, after the add…

2014-11-19abs ↗pdf ↗

Unsupervised method selects genes for tumor subtype discovery.

problem High-dimensional tumor gene expression data with noisy variables and heterogeneity.
method Autoencoders for latent space learning, Multiple Kernel Learning for feature selection, clustering.
result Lower redundancy and better clustering performance compared to benchmarks.

We consider high-dimensional regression over subgroups of observations. Our work is motivated by biomedical problems, where disease subtypes, for example, may differ with respect to underlying regression models, but sample sizes at the subgroup-level may be limited. We focus on the case in which subgroup-specific model…

2016-11-03abs ↗pdf ↗

We introduce a general framework for estimation of inverse covariance, or precision, matrices from heterogeneous populations. The proposed framework uses a Laplacian shrinkage penalty to encourage similarity among estimates from disparate, but related, subpopulations, while allowing for differences among matrices. We p…

2016-01-02abs ↗pdf ↗

Smile-GANs clusters brain MRI scans to reveal disease subtypes and progression.

problem Understanding disease heterogeneity in brain MRI scans.
method Generative Adversarial Networks (GANs) for semi-supervised clustering.
result Discovered four subtypes of Alzheimer's and prodromal phases, with two progressive pathways.