Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

13.0%26.0%39.1%52.1% · Jun 202019922001200920182026
48 results for clinical data analysis

The paper explores using historical data to improve clinical trial analysis by optimizing covariate weights.

problem Limited covariates in small clinical trials reduce the effectiveness of analysis.
method Leverage historical data to pre-specify covariate weights as a composite covariate.
result A composite covariate improves the cost/benefit ratio and reduces overfitting in small clinical trials.

Wide and deep neural network predicts Alzheimer's progression from shape and clinical data.

problem Predicting Alzheimer's disease progression from shape and clinical data.
method Fused anatomical shape and tabular clinical data in a neural network, employing survival analysis loss.
result The model outperforms shape and clinical models individually.

Paper proposes multimodal contrastive learning for EHR data.

problem Separate treatment of structured and unstructured EHR data.
method Proposes a multimodal feature embedding generative model and a multimodal contrastive loss.
result Multimodal learning yields better feature representation than single-modality learning.

Study investigates XAI methods in clinical gait analysis.

problem Limited understanding of machine learning models in healthcare.
method XAI methods, specifically Layer-wise Relevance Propagation (LRP), to explain ML predictions.
result Explanations from LRP show promising statistical and clinical relevance.

Analysis of flow cytometry data is an essential tool for clinical diagnosis of hematological and immunological conditions. Current clinical workflows rely on a manual process called gating to classify cells into their canonical types. This dependence on human annotation limits the rate, reproducibility, and complexity …

2017-11-21abs ↗pdf ↗

Machine learning predicts sperm motility from videos and participant data.

problem Manual evaluation of semen samples is time-consuming and unreliable.
method Used machine learning techniques including simple linear regression and convolutional neural networks.
result Sperm motility prediction using deep learning is rapid and consistent.

Health care is one of the most exciting frontiers in data mining and machine learning. Successful adoption of electronic health records (EHRs) created an explosion in digital clinical data available for analysis, but progress in machine learning for healthcare research has been difficult to measure because of the absen…

2017-03-22abs ↗pdf ↗

Paper presents an efficient method for selecting machine learning algorithms and hyper-parameters.

problem Efficient selection of machine learning algorithms and hyper-parameters is challenging for large datasets.
method Progressive sampling-based Bayesian optimization
result Significantly reduces search time, classification error rate, and error rate variability.

Study improves healthcare time series imputation by considering structured missingness.

problem Structured missingness in clinical data impacts time series imputation models.
method Analysis of different masking strategies on imputation methods using PhysioNet Challenge 2012 dataset.
result Masking choices significantly affect imputation accuracy and clinical prediction.

Qwant Research improves clinical case matching and information retrieval.

problem Matching and retrieving relevant clinical cases and discussions.
method Approach based on language models and preprocessings, information extraction system using neural networks and linguistic analysis.
result Very encouraging results in information extraction accuracy.

Develops a TL framework for estimating RMST difference in clinical trials.

problem Estimating RMST difference in clinical trials with time-to-event outcomes.
method Targeted learning (TL) framework using pseudo-observations and copy reference (CR) approach for sensitivity analysis.
result Demonstrated the effectiveness of the TL framework using real data.

This research uses machine learning to identify Alzheimer's disease subtypes and predict progression.

problem Heterogeneity in Alzheimer's disease clinical manifestations and progression rate limit personalized care and treatment planning.
method Unsupervised and supervised machine learning approaches applied to ADNI data.
result Identification of patient subtypes and prediction of disease progression zones.

HPPCA improves imputation of longitudinal data with missing values.

problem Handling incomplete, high-dimensional longitudinal data with nested sources of variation and temporal dependency.
method Hierarchical probabilistic principal component analysis (HPPCA) with a two-level latent factor model and Gaussian process.
result HPPCA outperforms standard PPCA and multivariate functional PCA in imputation accuracy, even under heavy missingness and model misspecification.

In this work we explored building automatic speech recognition models for transcribing doctor patient conversation. We collected a large scale dataset of clinical conversations (14,00014,000 hr), designed the task to represent the real word scenario, and explored several alignment approaches to iteratively improve data qua…

2017-11-20abs ↗pdf ↗

TrialGraph uses graph machine learning to improve clinical trial design and predict side effects.

problem Complexity and cost in clinical trials hinder drug development.
method Curated clinical trial data set converted to graph-structured formats, applied graph machine learning algorithms.
result MetaPath2Vec algorithm performed exceptionally well, improving prediction accuracy.

New approach for random forests protects privacy in collaborative prediction.

problem Privacy-preserving machine learning for ensemble methods in collaborative analysis.
method Each entity learns a model locally, and predictions are computed using all locally trained models without revealing extra information.
result High efficiency and potential accuracy benefit demonstrated on real-world datasets, including EHR data.

AdaptiveNet tackles disease progression prediction in rheumatoid arthritis using deep neural networks.

problem Predicting disease progression in rheumatoid arthritis using clinical data.
method AdaptiveNet, a novel recurrent neural network architecture, that handles multiple lists of different events and missing data.
result AdaptiveNet outperforms classical baselines in disease progression prediction.

FedRD improves risk difference estimation in federated learning for clinical outcomes.

problem Privacy-preserving model co-training in medical research is hindered by server-dependent architectures and focus on relative effect measures.
method FedRD is a server-independent, communication-efficient framework for federated risk difference estimation in distributed survival data.
result FedRD provides valid confidence intervals and hypothesis testing, and is asymptotically equivalent to pooled individual-level analysis.

Pairwise ranking aligns subjective clinical evaluations with objective indicators.

problem Aligning subjective clinical evaluations with objective indicators for improved diagnosis.
method Pairwise ranking methods to align subjective evaluations with objective indicators.
result The resulting score improves classification accuracy and provides a nuanced severity assessment.

Transfer learning improves clinical time series prediction with deep RNNs.

problem Training deep neural networks for clinical time series analysis requires large labeled data and expertise.
method Investigated transfer learning scenarios for deep RNNs: domain-adaptation and task-adaptation.
result Pre-trained deep models allow robust, efficient, and data-efficient clinical time series prediction.

CRBM generates digital twins for MS patients, aiding in disease progression analysis.

problem Characterizing and analyzing disease progression in MS patients.
method Unsupervised machine learning with Conditional Restricted Boltzmann Machines (CRBMs).
result Generated digital twins are statistically indistinguishable from actual subjects.

PKB framework boosts genomic data analysis by integrating pathway knowledge.

problem Boosting discovery power and connecting new findings with biological mechanisms in genomic data.
method Pathway-based Kernel Boosting (PKB) framework integrating clinical and pathway information for prediction of various outcomes.
result PKB substantially outperforms other methods in predicting drug response and cancer survival.

Improved model predicts ICU readmission and mortality with interpretable results.

problem Lack of clinically interpretable predictions from deep learning models on clinical notes.
method Augmented a convolutional model with an attention mechanism for clinical note prediction.
result Attention mechanism improves prediction performance while providing interpretable results.

New method targets relative risk heterogeneity in clinical trials.

problem Identifying treatment effects across subgroups with absolute risk differences.
method Modified causal forests using a novel node-splitting procedure based on relative risk.
result Relative risk causal forests can capture heterogeneity not detected by absolute risk methods.

The study explores using unlabeled data to improve survival time predictions.

problem Challenges in clinical follow-up studies due to drop-out and data collection issues.
method Investigates three approaches to incorporate unlabeled data in survival analysis.
result All approaches improve predictive performance over independent test data.

A new method uses deep Gaussian processes to handle missing values in irregularly sampled healthcare data.

problem Missing values and irregular sampling in healthcare data.
method Deep Gaussian process emulation with stochastic imputation.
result The method outperforms conventional imputation methods in clinical datasets.

C3T-Budget optimizes drug efficacy in dose-finding trials with budget and safety constraints.

problem Heterogeneous patient populations and budget constraints make dose-finding clinical trials challenging.
method Contextual constrained clinical trial algorithm that maximizes drug efficacy while learning subgroup responses.
result Demonstrates efficient budget usage and balanced learning-treatment trade-off in simulated trials.

KAPLAN-HR models survival data without manual interactions, outperforming existing methods.

problem Survival analysis challenges with complex covariates and time-varying effects.
method Kolmogorov-Arnold Networks (KAN) for nonparametric hazard estimation.
result KAPLAN-HR matches or exceeds existing methods in clinical survival data.

Framework improves clinical timeline reconstruction from text and tables.

problem Temporal precision and event timing in clinical narratives and EHRs.
method Retrieval-augmented multimodal alignment framework.
result Consistently improves absolute timestamp accuracy and temporal concordance.

Study compares LSTM, Transformer, and Mamba for bladder cancer recurrence analysis.

problem Complex time-dependent data in bladder cancer recurrence analysis.
method Evaluation of LSTM, Transformer, and Mamba models using Cox proportional hazards model.
result LSTM-Cox model outperforms Transformer-Cox and Mamba-Cox models in prediction accuracy.