Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,878 papers · 148 categories

Trend · papers per month

2.4%4.7%7.1%9.5% · Jan 201919922001200920172026
48 results for small cohort

Transformers simplify modeling of small longitudinal cohort data by reducing parameters and incorporating attention mechanisms.

problem Challenges in modeling longitudinal cohort data due to complex temporal dependencies and large dataset requirements.
method Simplified transformer architecture with attention mechanism, autoregressive model, and kernel-based temporal decay.
result The approach recovers contextual dependencies even with small datasets, identifying temporal patterns in stress and mental health.

ODVICE augments EHR cohorts using ontology to improve analysis robustness.

problem Limited records in cohorts for rare diseases hamper robust analysis.
method Ontology-driven Monte-Carlo graph spanning algorithm for data augmentation.
result ODVICE augmented cohorts show ~30% improvement in AUC over non-augmented datasets.

Deep neural networks improve sleep stage classification across diverse datasets.

problem Manual sleep scoring is subjective and lacks reliability; automatic systems generalize poorly.
method Developed a deep neural network using 15,684 polysomnography studies from five cohorts.
result Classification accuracy improved with more training data and multiple data sources.

Cohort effects are important factors in determining the evolution of human mortality for certain countries. Extensions of dynamic mortality models with cohort features have been proposed in the literature to account for these factors under the generalised linear modelling framework. In this paper we approach the proble…

2017-03-24abs ↗pdf ↗

Study identifies 1,012 persistent wallet cohorts on Solana pump.fun, showing coordinated buying behavior.

problem Understanding coordinated buying behavior on Solana pump.fun.
method Two-stage detection pipeline: first-buyer-window extraction followed by persistent-cohort surfacing via graph co-occurrence.
result 1,012 persistent wallet cohorts identified, showing systematic co-buying across multiple launches.

Study compares neural and statistical models for Parkinson's disease progression from voice data.

problem Difficult statistical analysis of longitudinal voice biomarkers due to subject correlation, small cohorts, and varied disease trajectories.
method Evaluated Neural Mixed Effects (NME), Generalized Neural Network Mixed Models (GNMMs), and semi-parametric Generalized Additive Mixed Models (GAMMs).
result GAMMs achieve stronger predictive performance and retain interpretable smooth effects and subject-level structure.

Research aims to ensure fair classification across explicit and implicit sensitive features.

problem Ensuring fairness in machine learning models when sensitive features are not explicitly provided.
method Defined explicit and implicit cohorts, used clustering of embeddings, modified loss function.
result Improved classification parity across explicit and implicit sensitive features.

Proposes a new model for mortality forecasting considering age groups and cohort effects.

problem Longevity risk due to ageing population.
method Mixed-effects time-series approach with age groups dependency and random cohort effects.
result Remarkable improvements in forecast accuracy compared to the CBD model.

CAT framework improves AI medical screening fairness and reliability.

problem Imbalanced data, varying performance across cohorts, and patient-level inconsistencies in traditional metrics.
method CAT framework introduces patient-level assessment, entropy-based distribution weighting, and cohort-weighted sensitivity and specificity.
result Enhanced predictive reliability, fairness, and interpretability of AI-driven medical screening models.

Enhancing spectral embedding for low-dimensional embeddings in rare disease cohorts

problem Representing clinical concepts and patients in electronic health records
method Spectral-based unsupervised learning with flexible knowledge transfer
result Outperforms competing approaches in challenging scenarios

Advances in molecular "omics'" technologies have motivated new methodology for the integration of multiple sources of high-content biomedical data. However, most statistical methods for integrating multiple data matrices only consider data shared vertically (one cohort on multiple platforms) or horizontally (different …

2019-06-09abs ↗pdf ↗

In statistical connectomics, the quantitative study of brain networks, estimating the mean of a population of graphs based on a sample is a core problem. Often, this problem is especially difficult because the sample or cohort size is relatively small, sometimes even a single subject. While using the element-wise sampl…

2016-09-06abs ↗pdf ↗

Deep generative model for healthcare data identifies coherent substructures and mutational clusters.

problem Analytical challenges in healthcare data, including sparsity, missingness, and small sample sizes.
method Proposes a deep generative Bayesian model with collapsed Gibbs sampling for multinomial count data.
result Identifies coherent substructures and biologically meaningful mutational clusters in cancer data.

Paper develops NN models for diabetes screening using NHANES data.

problem Developing accurate predictive models for diabetes in diverse populations.
method Proposes a neural network framework with survey weights, uncertainty quantification.
result Robust risk score models for diabetes in US population.

Study investigates deep learning model's reliability in clinical MRI data.

problem Tackles reliability of DL models in clinical out-of-distribution MRI data.
method Investigated performance of DL model trained on diverse datasets compared to clinical data.
result Model performs better in similar protocols but worse in clinical data with different tissue contrasts.

Clinical notes are a rich source of information about patient state. However, using them to predict clinical events with machine learning models is challenging. They are very high dimensional, sparse and have complex structure. Furthermore, training data is often scarce because it is expensive to obtain reliable labels…

2017-05-19abs ↗pdf ↗

The paper develops stochastic models for mortality rates using infinite dimensional processes.

problem Uncertainty in demographic projections of future mortality rates.
method Forward mortality models driven by Wiener process and Poisson random measure.
result Consistency conditions for forward mortality improvements and mortality rates.

Study finds that only a fraction of data is needed for accurate patient-level prediction models.

problem Developing predictive models for patient-level outcomes using large observational data.
method Empirical assessment of sample size effects on model performance and complexity using learning curves.
result A median reduction of 9.5% to 78.5% in the number of observations and 8.6% to 68.3% in the number of predictors can be achieved with adequate sample size.

A new model-free variable importance method (IGCS) is introduced for high-dimensional data.

problem Model-free variable importance for high-dimensional data, especially when prediction functions are proprietary or expensive.
method Integrated Gradient (IG) version of Cohort Shapley (CS) method with O(nd)\mathcal{O}(nd) cost.
result IGCS closely matches Cohort Shapley (CS) in performance, especially for binary predictors.

Quantum neural networks improve causal inference in biomedical studies, especially for small samples.

problem Addressing selection bias in comparing surgical techniques using observational data.
method Developed QNN-based propensity score models focusing on four key covariates (Age, Sex, Stage, BMI). Employed a linear ZFeatureMap for data encoding, SummedPaulis for predictions, and CMA-ES for optimization. Integrated noise modeling to enhance predictive stability.
result QNNs, particularly with noise-aware strategies, outperformed classical models in small samples, achieving AUC up to 0.750 for n=100.

Bayesian model predicts mental health symptoms from IAT data, improving accuracy over D-score.

problem Limited predictive performance of D-score method for mental health assessment.
method Sparse hierarchical Bayesian model leveraging multi-modal data.
result AUCs of 0.73 (E-IAT) and 0.76 (PSY-IAT) in best modality configurations, significant after FDR correction.