Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

5.9%11.8%17.8%23.7% · Jun 202019922001200920182026
48 results for phenotyping tasks

Multitask learning improves phenotyping in EHR data, but its benefits vary by phenotype complexity.

problem Improving phenotyping accuracy in EHR data using multitask learning.
method Investigated multitask learning for phenotyping rare and common phenotypes in EHR data using neural nets and logistic regression.
result Multitask learning with neural nets consistently outperforms single-task neural nets for rare phenotypes but underperforms for common phenotypes.

Deep learning outperforms classical methods in patient phenotyping.

problem Classifying patients into medical conditions using clinical notes.
method Comparison of CNNs, n-gram models, and cTAKES-based approaches on 10 phenotyping tasks.
result CNNs achieve an average F1-score of 76, significantly outperforming other methods.

We develop a model to cluster time-series data with interval censoring, improving disease phenotyping.

problem Noise and interval censoring hinder clustering in disease phenotyping.
method Deep generative, continuous-time model that clusters time-series data while correcting for censorship.
result Our model corrects for interval censoring and recovers known clinical subtypes.

Model identifies key problems in HIV patients' records.

problem Complex and time-consuming task of identifying patient problems from electronic health records.
method Unsupervised phenotyping approach that jointly learns phenotypes from structured and unstructured data.
result Learned phenotypes and their relatedness are clinically valid and surpass existing methods.

SWoTTeD discovers hidden temporal patterns in EHR data.

problem Complex temporal patterns in EHR data.
method Sliding Window for Temporal Tensor Decomposition (SWoTTeD) with constraints and regularizations.
result SWoTTeD achieves at least as accurate reconstruction as state-of-the-art models and extracts meaningful temporal phenotypes.

The linking genotype to phenotype is the fundamental aim of modern genetics. We focus on study of links between gene expression data and phenotype data through integrative analysis. We propose three approaches. 1) The inherent complexity of phenotypes makes high-throughput phenotype profiling a very difficult and labor…

2015-06-29abs ↗pdf ↗

Paper develops federated tensor factorization for phenotyping without sharing patient data.

problem Deriving phenotypes across multiple hospitals without patient-level data sharing.
method Secure data harmonization and federated computation using ADMM.
result Method generates phenotypes similar to centralized training while respecting privacy.

Bayesian model enhances phenotype discovery in asthma EHRs.

problem Lack of interpretability in unsupervised learning phenotyping of EHR data.
method Operationalized a Bayesian latent class framework with clinical knowledge priors.
result Identified an asthma sub-phenotype with elevated eosinophil levels and allergy markers.

Method predicts brain regions for neuroimaging phenotypes.

problem Predicting phenotypes from brain networks without static community structure.
method Supervised community detection using block-structured regularization and ADMM optimization.
result The method identifies task-specific brain regions that improve phenotype prediction.

New model predicts traits from gene expression, accounting for heterogeneity and gene networks.

problem Predicting phenotypes from gene expression data, considering heterogeneity and gene networks.
method Developed a novel model that considers heterogeneity and gene regulatory networks.
result Model performs well on prediction and provides clusters and gene regulatory networks.

auton-survival simplifies survival analysis for healthcare data.

problem Handling censored time-to-event data in healthcare.
method Open-source package for survival regression, adjustment, counterfactual estimation, phenotyping, and treatment effects.
result Demonstrates auton-survival's ability to support complex health and epidemiological questions.

This paper reviews methods for discovering patient subgroups from EHR data.

problem Discovering subgroups of patients and co-occurring medical conditions from EHR data.
method Low-rank data approximation methods like matrix and tensor decompositions.
result These methods provide transparent and interpretable insights into patient phenotypes.

New method learns parameter groups and structures in multi-response models.

problem Discovering unknown grouping structures in multi-response models.
method Proposes two convex regularization formulations and optimization approaches.
result Validated on simulations and real datasets, providing more accurate parameter estimation.

Develops methods for GWAS of high dimensional phenotypes using summary statistics.

problem Lack of methods to model pleiotropy in multi-phenotype GWAS.
method Bayesian inference model using summary statistics, fast computation, and biologically informed priors.
result Demonstrates utility in metabolite GWAS with interpretable pathway-level inference.

ODBAE detects complex phenotypes in biological data.

problem Challenges in identifying complex phenotypes from high-dimensional biological data.
method ODBAE (Outlier Detection using Balanced Autoencoders) identifies influential and high leverage points in latent relationships among multiple physiological parameters.
result ODBAE reveals novel metabolism-related genes and uncovers coordinated abnormalities across metabolic indicators.

Systematic review of electronic health record phenotyping approaches.

problem Detecting patient cohorts using electronic health records.
method Comprehensive literature review of preprocessing and modeling approaches.
result Natural language processing shows promise for electronic phenotyping.

TASTE combines static and temporal data for phenotyping EHRs.

problem Phenotyping EHRs with both static and temporal data.
method Jointly models static and temporal tensors using PARAFAC2 and non-negative matrix factorization, alternatingly solving sub-problems.
result TASTE outperforms existing methods in speed and clinical meaningfulness of phenotypes.

Study identifies three sub-phenotypes of AKI with different severity.

problem Tackles the heterogeneity of AKI to improve targeted interventions.
method Used a memory network-based deep learning approach on EHR data.
result Identified three distinct sub-phenotypes of AKI with varying severity.

Study uses LCA to identify ARDS sub-phenotypes improving predictive models.

problem Complex and heterogeneous nature of ARDS makes early recognition difficult.
method Applied latent class analysis to identify sub-groups, then built predictive models.
result Significantly improved prediction performance for two sub-phenotypes of ARDS.

Scientists interact with deep learning models to avoid misleading results.

problem Deep neural networks can misinterpret data and achieve high performance by exploiting confounding factors.
method Introduce explanatory interactive learning (XIL) where scientists revise models based on explanations.
result XIL helps prevent misleading results and encourages model trust.

WEST uses EHRs and expert cases to improve rare disease phenotyping.

problem Limited labeled data for rare diseases.
method Weakly supervised transformer model trained on probabilistic silver-standard labels.
result WEST outperforms existing methods in phenotype classification and subphenotyping.

Study develops electronic phenotypes of ICU patient acuity.

problem Limited time for patient acuity assessments and imprecise clinical trajectory prediction.
method Developed electronic phenotypes using automated variable retrieval in electronic health records.
result Identified three phenotypes: persistently stable, persistently unstable, and transitioning from unstable to stable.

Binary PheNorm extends phenotype labeling for EHRs using binary silver labels.

problem Lack of gold-standard phenotype labels in EHR studies.
method Proposes Binary PheNorm, an extension that uses binary silver labels directly in phenotype scoring.
result Binary PheNorm achieved strong discrimination using binary labels alone and improved performance when combined with count labels.

New method phenotypes sleep apnea patients using time series analysis.

problem Traditional diagnosis of sleep apnea is insufficient for capturing its multi-faceted outcomes.
method Fuzzy clustering in time and frequency domains, and persistent homology for topological analysis.
result Phenotyping patients improves understanding of sleep apnea.

Machine learning predicts plant phenotypes from soil microbiome data.

problem Predicting plant phenotypes from soil microbiome data.
method Two models (random forest and Bayesian neural network) were used to predict plant phenotypes from soil properties and microbial population density.
result Human decisions and normalization strategies significantly impact model performance.

Transfer learning improves clinical time series prediction with limited data.

problem Training deep RNNs for clinical tasks requires large labeled data and tuning.
method Transfer learning from pre-trained RNNs on multiple tasks to new tasks.
result Features from pre-trained RNNs improve model performance and robustness.

Unsupervised learning uncovers hidden patterns in health data.

problem Limited scalability and accuracy of supervised learning in identifying complex clinical patterns.
method Derives from Lasko et al. method, implemented in Apache Spark and Python, generalized for MIMIC-III lab data.
result Unsupervised learning finds patterns in EHRs that reveal subtypes of diseases.

Tree-based regularization improves latent variable inference from related datasets.

problem Inferring latent variables from multiple related datasets in causal systems.
method Tree-Based Regularization (TBR) for sparse changes across environments.
result TBR identifies true latent variables up to simple transformations under sparse changes.

New algorithm improves plant breeding by clustering soybean genotypes more accurately and efficiently.

problem Low accuracy and high computational complexity in clustering plant genotypes.
method Spectral Clustering with Pivotal Sampling for phenotypic data.
result Our algorithm achieves substantially more accuracy than existing methods.