Cohort effects are important factors in determining the evolution of human mortality for certain countries. Extensions of dynamic mortality models with cohort features have been proposed in the literature to account for these factors under the generalised linear modelling framework. In this paper we approach the proble…
The Lee Carter modelling framework is widely used because of its simplicity and robustness despite its inability to model specific cohort effects. A large number of extensions have been proposed that model cohort effects but there is no consensus. It is difficult to simultaneously account for cohort effects and age-adj…
Proposes a new model for mortality forecasting considering age groups and cohort effects.
problem Longevity risk due to ageing population.
method Mixed-effects time-series approach with age groups dependency and random cohort effects.
result Remarkable improvements in forecast accuracy compared to the CBD model.
Study identifies 1,012 persistent wallet cohorts on Solana pump.fun, showing coordinated buying behavior.
problem Understanding coordinated buying behavior on Solana pump.fun.
method Two-stage detection pipeline: first-buyer-window extraction followed by persistent-cohort surfacing via graph co-occurrence.
result 1,012 persistent wallet cohorts identified, showing systematic co-buying across multiple launches.
A new method to explain black box models using Shapley values.
problem Quantifying the impact of individual input variables in black box functions.
method Cohort Shapley measure based on cooperative game theory, using similarity cohorts.
result Introduces a new squared cohort Shapley value for variable importance.
Randomized Controlled Trials (RCTs) are the gold standard for comparing the effectiveness of a new treatment to the current one (the control). Most RCTs allocate the patients to the treatment group and the control group by uniform randomization. We show that this procedure can be highly sub-optimal (in terms of learnin…
Deep neural networks improve sleep stage classification across diverse datasets.
problem Manual sleep scoring is subjective and lacks reliability; automatic systems generalize poorly.
method Developed a deep neural network using 15,684 polysomnography studies from five cohorts.
result Classification accuracy improved with more training data and multiple data sources.
In systems biomedicine, an experimenter encounters different potential sources of variation in data such as individual samples, multiple experimental conditions, and multi-variable network-level responses. In multiparametric cytometry, which is often used for analyzing patient samples, such issues are critical. While c…
Develops a new method to model overlapping asymmetric datasets effectively.
problem Handling overlapping asymmetric datasets in data science.
method Twice penalized P-Spline approximation method.
result Improves model fit by over 65% in a real-life dataset.
Research aims to ensure fair classification across explicit and implicit sensitive features.
problem Ensuring fairness in machine learning models when sensitive features are not explicitly provided.
method Defined explicit and implicit cohorts, used clustering of embeddings, modified loss function.
result Improved classification parity across explicit and implicit sensitive features.
COHORTNEY groups web users based on activity patterns.
problem Lack of academic discussion on cohort analysis for user behavior.
method Unsupervised non-parametric machine learning approach.
result COHORTNEY outperforms traditional methods in cohort analysis.
Improves treatment effect estimation by reducing sample size needed.
problem Estimating causal treatment effects from observational data requires many covariates, increasing sample size.
method Proposes a nonconvex joint sparsity regularization objective function to recover a sparse subset of covariates.
result Improves sample complexity to scale with the size of the sparse subset and log of the total covariates.
ODVICE augments EHR cohorts using ontology to improve analysis robustness.
problem Limited records in cohorts for rare diseases hamper robust analysis.
method Ontology-driven Monte-Carlo graph spanning algorithm for data augmentation.
result ODVICE augmented cohorts show ~30% improvement in AUC over non-augmented datasets.
Develops a GP framework for age and year-specific mortality surfaces.
problem Learning the covariance structure of age and year-specific mortality surfaces.
method Genetic programming algorithm to search for the most expressive GP kernel.
result Reveals the presence/absence of cohort effects in different populations.
Study compares neural and statistical models for Parkinson's disease progression from voice data.
problem Difficult statistical analysis of longitudinal voice biomarkers due to subject correlation, small cohorts, and varied disease trajectories.
method Evaluated Neural Mixed Effects (NME), Generalized Neural Network Mixed Models (GNMMs), and semi-parametric Generalized Additive Mixed Models (GAMMs).
result GAMMs achieve stronger predictive performance and retain interpretable smooth effects and subject-level structure.
The paper develops stochastic models for mortality rates using infinite dimensional processes.
problem Uncertainty in demographic projections of future mortality rates.
method Forward mortality models driven by Wiener process and Poisson random measure.
result Consistency conditions for forward mortality improvements and mortality rates.
CAT framework improves AI medical screening fairness and reliability.
problem Imbalanced data, varying performance across cohorts, and patient-level inconsistencies in traditional metrics.
method CAT framework introduces patient-level assessment, entropy-based distribution weighting, and cohort-weighted sensitivity and specificity.
result Enhanced predictive reliability, fairness, and interpretability of AI-driven medical screening models.
Cohort analysis speeds up Bitcoin blockchain data queries.
problem Efficiently querying Bitcoin blockchain data for economic insights.
method Cohort analysis applied to Bitcoin transaction data.
result Creation of datasets and visualizations for key Bitcoin transaction indicators.
Mean-field approximations simplify insurance liability calculations.
problem High-dimensional system of equations makes insurance liability calculation infeasible.
method Use mean-field model to replace high-dimensional system with a low-dimensional non-linear system.
result Insurance liability converges to mean-field approximation as cohort size increases.
Causal analysis reveals regional discrepancies in TOPCAT trial results.
problem Inconclusive results in TOPCAT trial for heart failure treatment.
method Causal discovery methods with domain knowledge integration.
result Significant causal effects shown for some subgroups globally.
Typical cohorts in brain imaging studies are not large enough for systematic testing of all the information contained in the images. To build testable working hypotheses, investigators thus rely on analysis of previous work, sometimes formalized in a so-called meta-analysis. In brain imaging, this approach underlies th…
System exposes study population descriptions in clinical guidelines.
problem Challenges in understanding applicability of clinical guidelines.
method Developed an ontology-enabled prototype system using SIO.
result Allows medical practitioners to better understand study populations.
Brain imaging analysis on clinically acquired computed tomography (CT) is essential for the diagnosis, risk prediction of progression, and treatment of the structural phenotypes of traumatic brain injury (TBI). However, in real clinical imaging scenarios, entire body CT images (e.g., neck, abdomen, chest, pelvis) are t…
A new method for variable importance measures without impossible data.
problem Using impossible data for variable importance measures in black box models.
method Cohort Shapley, a method grounded in economic game theory using only observed data.
result Cohort Shapley provides a more trustworthy explanation of black box models' decisions.
Enhancing spectral embedding for low-dimensional embeddings in rare disease cohorts
problem Representing clinical concepts and patients in electronic health records
method Spectral-based unsupervised learning with flexible knowledge transfer
result Outperforms competing approaches in challenging scenarios
Advances in molecular "omics'" technologies have motivated new methodology for the integration of multiple sources of high-content biomedical data. However, most statistical methods for integrating multiple data matrices only consider data shared vertically (one cohort on multiple platforms) or horizontally (different …
Contextualized ML learns context-dependent effects using deep learning.
problem Learning heterogeneous and context-dependent effects in data.
method Applying deep learning to the meta-relationship between contextual information and context-specific parametric models.
result Unified framework for cluster analysis and cohort modeling.
Transformers simplify modeling of small longitudinal cohort data by reducing parameters and incorporating attention mechanisms.
problem Challenges in modeling longitudinal cohort data due to complex temporal dependencies and large dataset requirements.
method Simplified transformer architecture with attention mechanism, autoregressive model, and kernel-based temporal decay.
result The approach recovers contextual dependencies even with small datasets, identifying temporal patterns in stress and mental health.
Approaches for testing sets of variants, such as a set of rare or common variants within a gene or pathway, for association with complex traits are important. In particular, set tests allow for aggregation of weak signal within a set, can capture interplay among variants, and reduce the burden of multiple hypothesis te…
Meta Fusion integrates various multimodal data fusion strategies into a unified framework.
problem Improving predictive power of machine learning methods across diverse applications.
method Meta Fusion constructs a cohort of models based on latent representations across modalities, sharing soft information to boost performance.
result Meta Fusion consistently outperforms conventional fusion strategies in simulation and real-world applications.
Real world observational data, together with causal inference, allow the estimation of causal effects when randomized controlled trials are not available. To be accepted into practice, such predictive models must be validated for the dataset at hand, and thus require a comprehensive evaluation toolkit, as introduced he…
Compact formulas for evaluating insurance policies' risks.
problem Quantifying demographic risk in insurance portfolios.
method Cohort-based approach with market-consistent valuation.
result Formal closed formula for idiosyncratic risk (accidental mortality).
Deep learning predicts response to HER2-targeted breast cancer therapy.
problem Predicting response to HER2-targeted neoadjuvant chemotherapy.
method Developed and validated a deep learning approach using pre-treatment dynamic breast MRI.
result Deep learning model achieved strong performance in predicting pathological complete response.
Mobile technologies offer opportunities for higher resolution monitoring of health conditions. This opportunity seems of particular promise in psychiatry where diagnoses often rely on retrospective and subjective recall of mood states. However, getting actionable information from these rather complex time series is cha…
The paper presents a systematic review of state-of-the-art approaches to identify patient cohorts using electronic health records. It gives a comprehensive overview of the most commonly de-tected phenotypes and its underlying data sets. Special attention is given to preprocessing of in-put data and the different modeli…
Optimizes pension mix of PAYGO, EET, and individual savings.
problem Balancing PAYGO, EET, and individual savings in funded pension schemes.
method Solves a Nash equilibrium between pension participants and government, considering age-dependent preferences and optimal asset allocation.
result Identifies critical ages and optimal contribution rates for maximizing overall utility.
A new estimator reduces bias and improves efficiency for staggered adoption studies.
problem Bias in difference-in-differences estimates for staggered adoption studies.
method Fused Extended Two-Way Fixed Effects (FETWFE) estimator with automatic parameter selection.
result FETWFE identifies correct restrictions with probability tending to one, improving efficiency.
Deep transfer learning improves automatic sleep staging accuracy.
problem Small cohort data variability and inefficiency in sleep studies.
method Deep transfer learning approach using a large dataset to a small cohort.
result Significant performance improvement on automatic sleep staging.
A framework detects nonlinear and interaction effects in epidemiological data with uncertainty quantification.
problem Lack of reliable inference for ML-discovered nonlinearities and interactions in epidemiological data.
method Combines Bayesian sparse regression, tree ensembles, and Shapley values.
result Valid uncertainty quantification for feature effects at the individual level.
Paper develops NN models for diabetes screening using NHANES data.
problem Developing accurate predictive models for diabetes in diverse populations.
method Proposes a neural network framework with survey weights, uncertainty quantification.
result Robust risk score models for diabetes in US population.
Study investigates deep learning model's reliability in clinical MRI data.
problem Tackles reliability of DL models in clinical out-of-distribution MRI data.
method Investigated performance of DL model trained on diverse datasets compared to clinical data.
result Model performs better in similar protocols but worse in clinical data with different tissue contrasts.
Federated survival analysis outperforms local and centralized training, with RSF offering the best balance of discrimination, calibration, and robustness.
problem Survival analysis models require large, diverse cohorts but are limited by privacy regulations and lack of centralized data.
method Federated learning (FL) is used to train shared models without exchanging raw data.
result FL consistently outperforms local training and approaches, and occasionally exceeds centralized performance.
The paper compares two methods for handling missing data in causal discovery.
problem Handling missing data in causal discovery algorithms.
method Test-wise deletion and multiple imputation.
result Multiple imputation is more challenging for causal discovery than for estimation.
There is growing interest in the design of pension annuities that insure against idiosyncratic longevity risk while pooling and sharing systematic risk. This is partially motivated by the desire to reduce capital and reserve requirements while retaining the value of mortality credits; see for example Piggott, Valdez an…
How do macro-financial shocks affect investor behavior and market dynamics? Recent evidence on experience effects suggests a long-lasting influence of personally experienced outcomes on investor beliefs and investment, but also significant differences across older and younger generations. We formalize experience-based …
At this moment, databanks worldwide contain brain images of previously unimaginable numbers. Combined with developments in data science, these massive data provide the potential to better understand the genetic underpinnings of brain diseases. However, different datasets, which are stored at different institutions, can…
The widespread availability of electronic health records (EHRs) promises to usher in the era of personalized medicine. However, the problem of extracting useful clinical representations from longitudinal EHR data remains challenging. In this paper, we explore deep neural network models with learned medical feature embe…
Proposes counterfactual explanations for deep two-sample tests on high-dimensional data.
problem Limited interpretability of deep two-sample tests on high-dimensional data.
method Combines diffusion autoencoder and pretrained deep two-sample test model to generate counterfactuals.
result Counterfactual transformations increase p-values, indicating closer distribution similarity.