A tutorial on various methods for clustering longitudinal data.
problem Identifying groups with different trends in longitudinal data.
method Group-based trajectory modeling, growth mixture modeling, longitudinal k-means.
result Strengths, limitations, and model extensions of the methods are discussed.
New framework for forecasting psychological processes from ILD.
problem Forecasting psychological processes at the individual level from ILD.
method A novel modeling framework addressing challenges in ILD.
result Improved forecasting of psychological processes at the individual level.
We present a generative approach to classify scarcely observed longitudinal patient trajectories. The available time series are represented as tensors and factorized using generative deep recurrent neural networks. The learned factors represent the patient data in a compact way and can then be used in a downstream clas…
TPSQRs model longitudinal event data, detecting ADRs from EHRs.
problem Detecting adverse drug reactions from longitudinal event data.
method Learned by estimating a collection of interrelated PSQRs, using Poisson pseudo-likelihood for estimation.
result TPSQRs effectively and efficiently recover ADR signals from EHRs.
Proposes Deep LTMLE for estimating dynamic treatment effects in longitudinal studies.
problem Estimating counterfactual mean outcomes under dynamic treatment policies in longitudinal settings.
method Uses a transformer architecture with temporal-difference learning for initial estimation, followed by TMLE correction and statistical inference.
result Demonstrates superior performance in complex, long-term scenarios compared to existing methods.
Method uses ANN to estimate incentive salience from large behavioral data.
problem Estimating incentive salience in naturalistic settings.
method Artificial Neural Networks (ANNs) for latent state approximation.
result ANNs produce better representations for predicting future behaviour.
Novel method combines neural network features with survival models for ICU infections.
problem Improving predictive models of ICU infections while maintaining interpretability.
method Semi-parametric approach combining low-resolution and high-resolution data.
result Improved predictive power with interpretability maintained.
The episodic, irregular and asynchronous nature of medical data render them difficult substrates for standard machine learning algorithms. We would like to abstract away this difficulty for the class of time-stamped categorical variables (or events) by modeling them as a renewal process and inferring a probability dens…
LPCI provides valid prediction intervals for longitudinal data.
problem Current conformal prediction methods for time series data lack cross-sectional coverage when applied to longitudinal datasets.
method Modeling residual data as a quantile fixed-effects regression problem, constructing prediction intervals with a trained quantile regressor.
result LPCI achieves valid cross-sectional coverage and outperforms existing benchmarks in terms of longitudinal coverage rates.
This paper introduces a novel framework for modeling temporal events with complex longitudinal dependency that are generated by dependent sources. This framework takes advantage of multidimensional point processes for modeling time of events. The intensity function of the proposed process is a mixture of intensities, a…
GANs improve longitudinal data imputation but face challenges in missing data and class imbalance.
problem Missing data and class imbalance in longitudinal data.
method GANs applied to longitudinal data imputation (LDI).
result GANs show potential but need more versatile approaches.
New method models longitudinal data using variational inference and normalizing flows.
problem Handling high-dimensional longitudinal data with time dependency.
method Variational inference with normalizing flows for latent variables.
result The method achieves better likelihood estimates and more reliable missing data imputation.
Deep neural networks are a family of computational models that have led to a dramatical improvement of the state of the art in several domains such as image, voice or text analysis. These methods provide a framework to model complex, non-linear interactions in large datasets, and are naturally suited to the analysis of…
latrend simplifies longitudinal clustering for numeric measurements.
problem Clustering of longitudinal data to identify common trends over time.
method Unified framework for applying various clustering methods.
result Facilitates comparison and rapid prototyping of new methods.
HL-VAE extends VAE for heterogeneous temporal and longitudinal data.
problem Handling heterogeneous data in temporal and longitudinal datasets.
method Proposes HL-VAE, an extension of existing VAEs for temporal and longitudinal data, incorporating likelihood models for various data types.
result HL-VAE achieves competitive performance in missing value imputation and predictive accuracy.
Automates kernel discovery for longitudinal data analysis.
problem Handling irregularly sampled, sparse longitudinal data with multilevel correlation.
method Combines deep neural networks and non-parametric kernel methods to discover complex multilevel correlation structure.
result Significantly outperforms state-of-the-art methods on benchmark data sets.
Develops new algorithms for QRF to handle mixed-frequency and longitudinal data.
problem Handling mixed-frequency and longitudinal data in quantile regression.
method Mixed-Frequency Quantile Regression Forest (MIDAS-QRF) and Finite Mixture Quantile Regression Forest (FM-QRF).
result Valid and flexible models for complex empirical settings in financial risk management and climate-change impact evaluation.
A scalable model for high-dimensional longitudinal data.
problem Modeling high-dimensional, non-linear, time-varying longitudinal data.
method LMM-VAE, combining linear mixed models and amortized variational inference.
result Competitive performance across simulated and real-world datasets.
New method calibrates asynchronous, error-prone covariates for longitudinal data.
problem Estimation biases and slow convergence in analyzing time-varying covariates with measurement error.
method Functional calibration approach based on functional principal component analysis.
result Asymptotically unbiased and consistent estimators for time-invariant coefficients; optimal convergence rate for time-varying coefficients.
Finite mixture models have become a popular tool for clustering. Amongst other uses, they have been applied for clustering longitudinal data and clustering high-dimensional data. In the latter case, a latent Gaussian mixture model is sometimes used. Although there has been much work on clustering using latent variables…
CausalLongPFN predicts counterfactual outcomes from time-series data.
problem Predicting future outcomes under varying treatments in time-series data with confounding and heterogeneity.
method Prior-fitted network pretrained on synthetic episodes of temporal structural causal models.
result CausalLongPFN outperforms domain-trained models on factual and counterfactual prediction tasks.
Develops algorithm to find subgroups with different treatment effects in HIV patients.
problem Estimating treatment effects in EHR data with challenges like time-varying confounding.
method SDLD algorithm combining generalized interaction tree and longitudinal targeted maximum likelihood estimation.
result Identifies subgroups of HIV patients at higher risk of weight gain with dolutegravir-containing ARTs.
This paper reviews random forest methods for analyzing longitudinal data in precision medicine.
problem Analyzing longitudinal data for precision medicine.
method Extensions of random forest for longitudinal data analysis.
result Categorization of random forest methods for different data structures and repeated measurements.
We study regularized estimation in high-dimensional longitudinal classification problems, using the lasso and fused lasso regularizers. The constructed coefficient estimates are piecewise constant across the time dimension in the longitudinal problem, with adaptively selected change points (break points). We present an…
VTD uses deep embeddings to estimate treatment effects from longitudinal data without unconfoundedness assumption.
problem Challenges in estimating individualized treatment effects from longitudinal observational data due to confounding bias.
method Leverages deep variational embeddings and observed proxies to learn hidden confounders.
result Effective in estimating treatment effects when hidden confounding is the leading bias.
While studying response trajectory, often the population of interest may be diverse enough to exist distinct subgroups within it and the longitudinal change in response may not be uniform in these subgroups. That is, the timeslope and/or influence of covariates in longitudinal profile may vary among these different sub…
TransformerLSR models longitudinal, recurrent, and survival data jointly.
problem Joint modeling of longitudinal measurements, recurrent events, and survival data with dependencies.
method Transformer-based deep learning framework integrating deep temporal point processes and latent structure representation.
result TransformerLSR effectively models all three components simultaneously, demonstrating necessity and effectiveness through simulations and real-world data.
CDM models counterfactual outcomes in longitudinal data with improved accuracy.
problem Predicting counterfactual outcomes in longitudinal data with complex time-dependent confounding.
method Causal Diffusion Model (CDM) using denoising diffusion architecture with relational self-attention.
result CDM outperforms state-of-the-art methods in generating full probabilistic distributions of counterfactual outcomes.
TraCeR uses transformers to analyze survival data with longitudinal covariates.
problem Handling longitudinal covariates and assessing model calibration in survival analysis.
method Transformer-based survival analysis framework with factorized self-attention architecture.
result TraCeR achieves significant performance improvements over state-of-the-art methods.
DEBIAS learns causal effects from psychiatric longitudinal data by optimizing outcome weights.
problem Causal inference challenges in psychiatric longitudinal data due to symptom heterogeneity and latent confounding.
method DEBIAS algorithm that optimizes outcome weights to maximize durable treatment effects and minimize confounding.
result DEBIAS consistently outperforms state-of-the-art methods in recovering causal effects for clinically interpretable composite outcomes.
Model uses smartphone data to assess MS trajectories.
problem Personalized longitudinal MS assessment.
method Imputation, generalized estimation equation, ensemble learning, fine-tuning.
result Promising model for predicting MS over time.
Study finds little progress in medical machine learning benchmarks over 3 years.
problem Lack of meaningful progress in medical machine learning benchmarks for structured healthcare data.
method Comprehensive review and meta-analysis of benchmarks in medical machine learning for structured data.
result Deep recurrent models perform only better than logistic regression on certain clinical prediction tasks.
Proposes L-VAE for longitudinal data analysis.
problem Analyse high-dimensional longitudinal data with missing values.
method Uses a multi-output additive Gaussian process (GP) prior to extend VAE's capability.
result Achieves highly accurate predictive performance.
We consider the problem of learning predictive models from longitudinal data, consisting of irregularly repeated, sparse observations from a set of individuals over time. Such data often exhibit {\em longitudinal correlation} (LC) (correlations among observations for each individual over time), {\em cluster correlation…
Pipeline integrates cross-sectional and longitudinal multi-omics data for IBD research.
problem Integrating diverse data types from the same individuals for disease understanding.
method Statistical and deep learning methods for variable selection, feature extraction, and joint integration.
result Identified microbial pathways, metabolites, and genes discriminating IBD status.
Bayesian Causal Forests model assesses part-time work's impact on student growth.
problem Estimating causal effects of part-time work on student growth in mathematics achievement.
method Longitudinal Bayesian Causal Forests model combining non-parametric and difference-in-differences methods.
result Negative impact of part-time work for most students, potential benefits for those with low school belonging, widening achievement gap identified.
New method for analyzing multiple longitudinal data processes.
problem Exploring associations between multiple random processes observed jointly.
method Functional Generalized Canonical Correlation Analysis (FGCCA) based on multiblock Regularized Generalized Canonical Correlation Analysis (RGCCA).
result FGCCA framework is robust to sparsely and irregularly observed data.
FREEtree improves tree-based methods for correlated longitudinal data.
problem Poor performance of Random Forests in high dimensional longitudinal data with correlated features.
method FREEtree uses a piecewise random effects model and clustering with WGCNA to select features and maintain interpretability.
result FREEtree outperforms other tree-based methods in prediction and feature selection accuracy.
A Longitudinal Attribute-Conditioned Neural Network (LANTERN) framework for modeling health-state transition probabilities in irregular longitudinal data.
problem Estimating long-term care transition probabilities in irregular longitudinal health data.
method A neural network that learns from individual health history, incorporates time elapsed, and conditions on demographic and socioeconomic attributes.
result Improves severe disability discrimination and maintains strong calibration.
HPPCA improves imputation of longitudinal data with missing values.
problem Handling incomplete, high-dimensional longitudinal data with nested sources of variation and temporal dependency.
method Hierarchical probabilistic principal component analysis (HPPCA) with a two-level latent factor model and Gaussian process.
result HPPCA outperforms standard PPCA and multivariate functional PCA in imputation accuracy, even under heavy missingness and model misspecification.
Develops privacy-preserving methods for longitudinal linear regression.
problem Protecting individual information in longitudinal data with privacy-preserving statistics.
method Proposes a user-level private regression estimator and a privatized covariance estimator for longitudinal linear regression under user-level differential privacy.
result Establishes theoretical guarantees for practical user-level differential privacy estimation and inference in longitudinal linear regression.
Traditional Recurrent Neural Networks assume vectorized data as inputs. However many data from modern science and technology come in certain structures such as tensorial time series data. To apply the recurrent neural networks for this type of data, a vectorisation process is necessary, while such a vectorisation leads…
Proposes a method for valid inference in GPLSIMs with longitudinal data.
problem Challenges in longitudinal data inference due to within-subject correlation and unstable variance estimation.
method Profile estimating-equation approach using spline approximation and block empirical likelihood.
result Block empirical likelihood ratio statistic with Wilks-type chi-square limit for joint inference.
Develops a hierarchical model for analyzing longitudinal data on manifolds.
problem Correlation between intra-subject measurements in nonlinear manifolds.
method Geodesic hierarchical models extended to Bézier spline trends.
result Validated on osteoarthritis data, improving disease progression classification.
Analyzing electronic health records (EHR) poses significant challenges because often few samples are available describing a patient's health and, when available, their information content is highly diverse. The problem we consider is how to integrate sparsely sampled longitudinal data, missing measurements informative …
MMM model clusters mixed-type longitudinal data efficiently.
problem Challenges in clustering multivariate longitudinal mixed-type data.
method MMM model reorganizes data into a three-way structure, using a mixture of matrix-variate normal distributions.
result MMM model handles various data types (continuous, ordinal, binary, nominal, count) and temporal dependence.
Unified method for discovering biclusters and triclusters in longitudinal data.
problem High-dimensional, sparsely sampled, irregularly observed longitudinal data.
method Tri-SfSVD, a unified sparse functional Singular Value Decomposition framework.
result Identified localized structures at the subject, subject-feature, and subject-feature-time levels.
Joint models for longitudinal and time-to-event data are commonly used in longitudinal studies to forecast disease trajectories over time. While there are many advantages to joint modeling, the standard forms suffer from limitations that arise from a fixed model specification, and computational difficulties when applie…