Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

2457 · Nov 201819922001200920182026
48 results for Behavioral-clinical phenotyping

Unsupervised clustering identifies meal patterns in T2DM self-monitoring data.

problem Identifying individual-level behavioral-clinical phenotypes in T2DM self-monitoring data.
method Hierarchical clustering of blood glucose and macronutrient consumption.
result All 9 gold standard patterns were re-discovered using HC, and most clusters were rated positively by CDEs.

Multitask learning improves phenotyping in EHR data, but its benefits vary by phenotype complexity.

problem Improving phenotyping accuracy in EHR data using multitask learning.
method Investigated multitask learning for phenotyping rare and common phenotypes in EHR data using neural nets and logistic regression.
result Multitask learning with neural nets consistently outperforms single-task neural nets for rare phenotypes but underperforms for common phenotypes.

The linking genotype to phenotype is the fundamental aim of modern genetics. We focus on study of links between gene expression data and phenotype data through integrative analysis. We propose three approaches. 1) The inherent complexity of phenotypes makes high-throughput phenotype profiling a very difficult and labor…

2015-06-29abs ↗pdf ↗

Deep learning outperforms classical methods in patient phenotyping.

problem Classifying patients into medical conditions using clinical notes.
method Comparison of CNNs, n-gram models, and cTAKES-based approaches on 10 phenotyping tasks.
result CNNs achieve an average F1-score of 76, significantly outperforming other methods.

Paper develops federated tensor factorization for phenotyping without sharing patient data.

problem Deriving phenotypes across multiple hospitals without patient-level data sharing.
method Secure data harmonization and federated computation using ADMM.
result Method generates phenotypes similar to centralized training while respecting privacy.

Bayesian model enhances phenotype discovery in asthma EHRs.

problem Lack of interpretability in unsupervised learning phenotyping of EHR data.
method Operationalized a Bayesian latent class framework with clinical knowledge priors.
result Identified an asthma sub-phenotype with elevated eosinophil levels and allergy markers.

Model identifies key problems in HIV patients' records.

problem Complex and time-consuming task of identifying patient problems from electronic health records.
method Unsupervised phenotyping approach that jointly learns phenotypes from structured and unstructured data.
result Learned phenotypes and their relatedness are clinically valid and surpass existing methods.

This paper reviews methods for discovering patient subgroups from EHR data.

problem Discovering subgroups of patients and co-occurring medical conditions from EHR data.
method Low-rank data approximation methods like matrix and tensor decompositions.
result These methods provide transparent and interpretable insights into patient phenotypes.

Develops methods for GWAS of high dimensional phenotypes using summary statistics.

problem Lack of methods to model pleiotropy in multi-phenotype GWAS.
method Bayesian inference model using summary statistics, fast computation, and biologically informed priors.
result Demonstrates utility in metabolite GWAS with interpretable pathway-level inference.

ODBAE detects complex phenotypes in biological data.

problem Challenges in identifying complex phenotypes from high-dimensional biological data.
method ODBAE (Outlier Detection using Balanced Autoencoders) identifies influential and high leverage points in latent relationships among multiple physiological parameters.
result ODBAE reveals novel metabolism-related genes and uncovers coordinated abnormalities across metabolic indicators.

Systematic review of electronic health record phenotyping approaches.

problem Detecting patient cohorts using electronic health records.
method Comprehensive literature review of preprocessing and modeling approaches.
result Natural language processing shows promise for electronic phenotyping.

TASTE combines static and temporal data for phenotyping EHRs.

problem Phenotyping EHRs with both static and temporal data.
method Jointly models static and temporal tensors using PARAFAC2 and non-negative matrix factorization, alternatingly solving sub-problems.
result TASTE outperforms existing methods in speed and clinical meaningfulness of phenotypes.

Study identifies three sub-phenotypes of AKI with different severity.

problem Tackles the heterogeneity of AKI to improve targeted interventions.
method Used a memory network-based deep learning approach on EHR data.
result Identified three distinct sub-phenotypes of AKI with varying severity.

Study uses LCA to identify ARDS sub-phenotypes improving predictive models.

problem Complex and heterogeneous nature of ARDS makes early recognition difficult.
method Applied latent class analysis to identify sub-groups, then built predictive models.
result Significantly improved prediction performance for two sub-phenotypes of ARDS.

SWoTTeD discovers hidden temporal patterns in EHR data.

problem Complex temporal patterns in EHR data.
method Sliding Window for Temporal Tensor Decomposition (SWoTTeD) with constraints and regularizations.
result SWoTTeD achieves at least as accurate reconstruction as state-of-the-art models and extracts meaningful temporal phenotypes.

WEST uses EHRs and expert cases to improve rare disease phenotyping.

problem Limited labeled data for rare diseases.
method Weakly supervised transformer model trained on probabilistic silver-standard labels.
result WEST outperforms existing methods in phenotype classification and subphenotyping.

We develop a model to cluster time-series data with interval censoring, improving disease phenotyping.

problem Noise and interval censoring hinder clustering in disease phenotyping.
method Deep generative, continuous-time model that clusters time-series data while correcting for censorship.
result Our model corrects for interval censoring and recovers known clinical subtypes.

Study develops electronic phenotypes of ICU patient acuity.

problem Limited time for patient acuity assessments and imprecise clinical trajectory prediction.
method Developed electronic phenotypes using automated variable retrieval in electronic health records.
result Identified three phenotypes: persistently stable, persistently unstable, and transitioning from unstable to stable.

Binary PheNorm extends phenotype labeling for EHRs using binary silver labels.

problem Lack of gold-standard phenotype labels in EHR studies.
method Proposes Binary PheNorm, an extension that uses binary silver labels directly in phenotype scoring.
result Binary PheNorm achieved strong discrimination using binary labels alone and improved performance when combined with count labels.

New method phenotypes sleep apnea patients using time series analysis.

problem Traditional diagnosis of sleep apnea is insufficient for capturing its multi-faceted outcomes.
method Fuzzy clustering in time and frequency domains, and persistent homology for topological analysis.
result Phenotyping patients improves understanding of sleep apnea.

Machine learning predicts plant phenotypes from soil microbiome data.

problem Predicting plant phenotypes from soil microbiome data.
method Two models (random forest and Bayesian neural network) were used to predict plant phenotypes from soil properties and microbial population density.
result Human decisions and normalization strategies significantly impact model performance.

New algorithm identifies multiple medical conditions from EHRs.

problem Automatically phenotyping multiple medical conditions from clinical notes.
method Constrained Non-Negative Matrix Factorization (NMF) with domain-specific constraints.
result Learned phenotypes are clinically interpretable and predictive of mortality.

Unsupervised learning uncovers hidden patterns in health data.

problem Limited scalability and accuracy of supervised learning in identifying complex clinical patterns.
method Derives from Lasko et al. method, implemented in Apache Spark and Python, generalized for MIMIC-III lab data.
result Unsupervised learning finds patterns in EHRs that reveal subtypes of diseases.

New algorithm improves plant breeding by clustering soybean genotypes more accurately and efficiently.

problem Low accuracy and high computational complexity in clustering plant genotypes.
method Spectral Clustering with Pivotal Sampling for phenotypic data.
result Our algorithm achieves substantially more accuracy than existing methods.

We propose a non-parametric regression methodology, Random Forests on Distance Matrices (RFDM), for detecting genetic variants associated to quantitative phenotypes representing the human brain's structure or function, and obtained using neuroimaging techniques. RFDM, which is an extension of decision forests, requires…

2013-09-24abs ↗pdf ↗

Paper models Alzheimer's disease using genotypic, phenotypic, and cognitive data.

problem Early detection and risk factor identification for Alzheimer's disease.
method Probabilistic generative subspace learning from multi-view medical data.
result Proposes a method to model Alzheimer's disease that combines genotypic, phenotypic, and cognitive data.

Linear Mixed Models (LMMs) are important tools in statistical genetics. When used for feature selection, they allow to find a sparse set of genetic traits that best predict a continuous phenotype of interest, while simultaneously correcting for various confounding factors such as age, ethnicity and population structure…

2015-07-16abs ↗pdf ↗

New model predicts traits from gene expression, accounting for heterogeneity and gene networks.

problem Predicting phenotypes from gene expression data, considering heterogeneity and gene networks.
method Developed a novel model that considers heterogeneity and gene regulatory networks.
result Model performs well on prediction and provides clusters and gene regulatory networks.

sGLMM corrects genetic associations with complex relatedness and confounding.

problem Correcting spurious associations in complex genetic data with population stratification and relatedness.
method Sparse graph-structured linear mixed model (sGLMM) that incorporates relatedness information and confounding correction.
result sGLMM outperforms existing approaches in modeling correlation from population structure and shared signals.

Paper models treatment effects by clustering patients with distinct survival characteristics.

problem Estimating treatment efficacy in clinical settings with censored outcomes.
method Latent variable approach to model heterogeneous treatment effects.
result The latent structure can mediate base survival rates and reveal actionable phenotypes.

New method identifies key genes affecting phenotypes in biological systems.

problem Identifying genes that drive specific phenotypes in complex biological systems.
method Data-driven observability decomposition using Koopman operators.
result Koopman operator representation identifies genes that drive phenotypes.