Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

3.6%7.1%10.7%14.3% · Oct 199219922001200920172026
48 results for subtype identification

Bayesian model clusters diverse 'omics data for disease subtyping.

problem Clustering diverse 'omics datasets conflates multiple structures.
method Multi-view Bayesian mixture model with semi-supervised learning.
result Identifies distinct clusters of patients for stratified medicine.

Study identifies five AD subtypes using graph diffusion and similarity learning.

problem Identifying homogeneous AD subtypes to improve diagnosis and treatment.
method Unsupervised clustering with graph diffusion and similarity learning.
result Five distinct AD subtypes identified with significant differences in biomarkers and clinical features.

While developing their software, professional object-oriented (OO) software developers keep in their minds an image of the subtyping relation between types in their software. The goal of this paper is to present an observation about the graph of the subtyping relation in Java, namely the observation that, after the add…

2014-11-19abs ↗pdf ↗

Unsupervised method selects genes for tumor subtype discovery.

problem High-dimensional tumor gene expression data with noisy variables and heterogeneity.
method Autoencoders for latent space learning, Multiple Kernel Learning for feature selection, clustering.
result Lower redundancy and better clustering performance compared to benchmarks.

Computational approaches to transcription factor binding site identification have been actively researched for the past decade. Negative examples have long been utilized in de novo motif discovery and have been shown useful in transcription factor binding site search as well. However, understanding of the roles of nega…

2011-04-07abs ↗pdf ↗

The task of clustering a set of objects based on multiple sources of data arises in several modern applications. We propose an integrative statistical model that permits a separate clustering of the objects for each data source. These separate clusterings adhere loosely to an overall consensus clustering, and hence the…

2013-02-28abs ↗pdf ↗

New framework distinguishes lung cancer subtypes using MALDI mass spectrometry.

problem Distinguishing between adenocarcinoma and squamous cell carcinoma subtypes in lung cancer.
method Supervised topological data analysis on MALDI mass spectrometry imaging data.
result The proposed framework successfully classifies lung cancer subtypes with competitive results.

Smile-GANs clusters brain MRI scans to reveal disease subtypes and progression.

problem Understanding disease heterogeneity in brain MRI scans.
method Generative Adversarial Networks (GANs) for semi-supervised clustering.
result Discovered four subtypes of Alzheimer's and prodromal phases, with two progressive pathways.

Cluster analysis methods are used to identify homogeneous subgroups in a data set. In biomedical applications, one frequently applies cluster analysis in order to identify biologically interesting subgroups. In particular, one may wish to identify subgroups that are associated with a particular outcome of interest. Con…

2013-04-13abs ↗pdf ↗

Model learns to select relevant clinical variables for disease subtype prediction from small data.

problem Few-shot disease subtype prediction from small genomic data.
method Meta learning Prototypical Network with feature selection and sample reweighting.
result Superior performance in predicting disease subtypes and identifying genes.

Study identifies biomarkers for lung cancer in female non-smokers.

problem Identifying prognostic biomarkers for stage III NSCLC in non-smoking females.
method Gene expression profiling and XGBoost machine learning algorithm.
result Top biomarkers validated in literature, with AUC score of 0.835.

Deep learning has demonstrated success in health risk prediction especially for patients with chronic and progressing conditions. Most existing works focus on learning disease Network (StageNet) model to extract disease stage information from patient data and integrate it into risk prediction. StageNet is enabled by (1…

2020-01-24abs ↗pdf ↗

UCSL combines clustering with supervised learning to discover interpretable subtypes.

problem Discovering interpretable subtypes in datasets relevant to supervised tasks.
method UCSL (Unsupervised Clustering driven by Supervised Learning) framework integrating clustering and supervised learning.
result UCSL achieves +1.9 points in balanced accuracy for psychiatric diseases clustering.

More than two thirds of mental health problems have their onset during childhood or adolescence. Identifying children at risk for mental illness later in life and predicting the type of illness is not easy. We set out to develop a platform to define subtypes of childhood social-emotional development using longitudinal,…

2016-12-04abs ↗pdf ↗

Study examines XAI methods for ECG analysis to improve model transparency.

problem Lack of transparency in deep learning models for ECG analysis.
method Investigates post-hoc XAI methods for local and global perspectives, establishes sanity checks, and demonstrates knowledge discovery.
result Quantitative evidence supports expert rules for sensible attribution methods and demonstrates XAI's utility for knowledge discovery.

VICatMix clusters categorical biomedical data efficiently and selects relevant variables.

problem Efficient clustering of high-dimensional categorical biomedical data.
method Variational Bayesian finite mixture model with variational inference.
result Improves clustering accuracy and variable selection on noisy, high-dimensional data.

The study finds obstructions for certain Weyl curvature tensors on manifolds.

problem Can manifolds admit metrics with purely electric or magnetic Weyl tensors?
method Analyzes algebraic curvature tensors and their Pontryagin classes on scalar product spaces.
result Obstructions to the existence of metrics with PE or PM Weyl tensors in top-degree cohomology.

Multi-view data, that is matched sets of measurements on the same subjects, have become increasingly common with advances in multi-omics technology. Often, it is of interest to find associations between the views that are related to the intrinsic class memberships. Existing association methods cannot directly incorpora…

2018-11-20abs ↗pdf ↗

OPAL optimizes labeling strategy for precise inference from uncertain models.

problem Inference from uncertain machine learning models is brittle.
method OPAL learns a smooth policy to adaptively label data points based on model uncertainty.
result OPAL yields estimators with the lowest variance and achieves nominal coverage in finite samples.

In clinical practice and biomedical research, measurements are often collected sparsely and irregularly in time while the data acquisition is expensive and inconvenient. Examples include measurements of spine bone mineral density, cancer growth through mammography or biopsy, a progression of defective vision, or assess…

2018-09-24abs ↗pdf ↗

The study examines Cox models for lifetime loan default risk, addressing biased estimates by incorporating recurrent events.

problem Ignoring recurrent default events in Cox models leads to biased and inaccurate PD estimates.
method Investigates and compares different Cox models (Andersen-Gill and Prentice-Williams-Peterson) for lifetime loan default risk.
result The Andersen-Gill model underperforms compared to the Prentice-Williams-Person model and the time to first default model.

Paper proposes scalable method for analyzing multi-omic data.

problem Integrating high-dimensional multi-omic data for cancer subtyping.
method Mixed graphical model approach using Birth-Death MCMC algorithm.
result Our method outperforms LASSO and standard BDMCMC in computational efficiency and model selection accuracy.