New method interprets deep embeddings for diabetes patient clustering.
problem Interpreting deep embeddings for disease progression.
method Patient clustering approach using deep embeddings.
result Clinically meaningful insights into diabetes progression patterns.
Study develops a dynamic risk model for COVID-19 mortality using UK Biobank data.
problem Developing tools to monitor high-risk patients during the COVID-19 pandemic.
method Data-driven random forest classification model using baseline characteristics and symptoms.
result Model predicts COVID-19 mortality with excellent performance (AUC: 0.91), identifying novel predictors.
Stochastic encoding improves gender classification of brain networks from UK Biobank data.
problem Complexity and bias in interpreting deep learning models of brain connectivity.
method Stochastic encoding in ensemble of CNNs, multivariate balancing algorithm.
result AUROC of 0.8459, with resting-state data more accurate than task data.
New methods improve genetic studies of complex diseases.
problem Improving genetic studies of complex diseases using high-dimensional clinical data.
method Evaluation of unsupervised disentangled representation learning methods (autoencoders, VAE, beta-VAE, FactorVAE) for genetic association studies.
result FactorVAEs and beta-VAEs outperform standard VAEs and non-variational autoencoders in genetic studies of asthma and COPD.
Maintaining good cardiac function for as long as possible is a major concern for healthcare systems worldwide and there is much interest in learning more about the impact of different risk factors on cardiac health. The aim of this study is to analyze the impact of systolic blood pressure (SBP) on cardiac function whil…
We propose a method based on deep learning to perform cardiac segmentation on short axis MRI image stacks iteratively from the top slice (around the base) to the bottom slice (around the apex). At each iteration, a novel variant of U-net is applied to propagate the segmentation of a slice to the adjacent slice below it…
Method tackles missing covariates in large-scale datasets.
problem Cross-population missing data problem in large-scale datasets.
method Augmented transfer regression learning method combining importance-weighted estimating equations and imputation terms.
result Estimator is n1/2-consistent and asymptotically normal, attaining semiparametric efficiency bound under correct specification. Framework for imputing missing heart data to simulate brain-heart interactions.
problem Lack of multi-modal patient data representing heart and brain processes.
method Probabilistic framework for joint cardiac data imputation and mechanistic model personalization.
result Accurate imputation of missing cardiac features in incomplete datasets.
The exploitation of large-scale population data has the potential to improve healthcare by discovering and understanding patterns and trends within this data. To enable high throughput analysis of cardiac imaging data automatically, a pipeline should comprise quality monitoring of the input images, segmentation of the …
Study compares estimators for causal mediation analysis with multiple mediators.
problem Estimating causal effects through multiple mediators in observational studies.
method Parametric and non-parametric estimators, including multiply robust and double machine learning approaches.
result Advanced estimators perform well across various settings and real data.
Bayesian meta-learning improves health prediction models across similar diseases.
problem Inter- and intra-task variability in healthcare predictions due to disease heterogeneity and patient differences.
method Bayesian meta-learning approach that models task similarity to mitigate negative transfer and improve generalizability.
result Significant generalizability improvements in stroke prediction tasks using electronic health record data.
Digital risk scores predict depression and anxiety over 10 years.
problem Identifying individuals at risk of depression and anxiety.
method Developed a 10-year predictive algorithm using UKB cohort, selecting predictors via Cox proportional hazards model and DeepSurv.
result Highly discriminating models for depression and anxiety were developed.
Modeling how individuals evolve over time is a fundamental problem in the natural and social sciences. However, existing datasets are often cross-sectional with each individual observed only once, making it impossible to apply traditional time-series methods. Motivated by the study of human aging, we present an interpr…
At this moment, databanks worldwide contain brain images of previously unimaginable numbers. Combined with developments in data science, these massive data provide the potential to better understand the genetic underpinnings of brain diseases. However, different datasets, which are stored at different institutions, can…
Ensemble models provide more accurate feature importance estimates than single models.
problem Inaccurate variable importance estimates due to model instability and stochasticity.
method Theoretical analysis and validation on benchmarks and real data.
result Ensembling at the model level reduces excess risk and provides more accurate variable-importance estimates.
Paper develops conformalized survival analysis method for better prediction.
problem Survival analysis models often misspecify and require strong assumptions.
method Uses conformal prediction to wrap around any survival prediction algorithm.
result Lower predictive bounds provide guaranteed coverage without strong assumptions.
Identification of disease subtypes and corresponding biomarkers can substantially improve clinical diagnosis and treatment selection. Discovering these subtypes in noisy, high dimensional biomedical data is often impossible for humans and challenging for machines. We introduce a new approach to facilitate the discovery…
CNNs achieve remarkable performance by leveraging deep, over-parametrized architectures, trained on large datasets. However, they have limited generalization ability to data outside the training domain, and a lack of robustness to noise and adversarial attacks. By building better inductive biases, we can improve robust…
We construct genomic predictors for heritable and extremely complex human quantitative traits (height, heel bone density, and educational attainment) using modern methods in high dimensional statistics (i.e., machine learning). Replication tests show that these predictors capture, respectively, ∼40, 20, and 9 perc…
Bayesian hypergraph inference models disease pathways from EHR data.
problem Modeling rare diseases influenced by shared risk factors.
method Bayesian hypergraph inference framework reframing multi-disease modeling.
result Interpretable disease pathways and well-calibrated uncertainty quantification.
New method for mixed data types in graphical models.
problem Challenges in analyzing data with mixed variable types.
method Latent Gaussian copula models with leveraged polychoric and polyserial correlations.
result Flexible and scalable methodology for mixed data types.
ICAM creates interpretable feature attribution maps for brain images.
problem Challenges in predicting class relevance from brain images due to heterogeneity and background variation.
method A VAE-GAN framework for disentangling class relevance from background features.
result FA maps generated by ICAM outperform baseline methods and support phenotype variation exploration.
New method improves statistical inference using machine learning-imputed data.
problem Improving statistical inference with imputed data from machine learning.
method Two-phase sampling approach for Z-estimation with ML-imputed outcomes.
result Guaranteed efficiency matching or exceeding classical inference, regardless of prediction quality.
In this paper, we investigate the capability of the universal Kriging (UK) model for single-objective global optimization applied within an efficient global optimization (EGO) framework. We implemented this combined UK-EGO framework and studied four variants of the UK methods, that is, a UK with a first-order polynomia…
A transfer learning method builds high-dimensional models using disparate datasets.
problem Building comprehensive prediction models with small sample sizes and limited features.
method Transfer learning approach using external data to build a reduced model and apply calibration equations.
result Proposes a penalized generalized method of moment framework for inference and one-step estimation.
Unified CCA methods for large-scale data with fast SGD algorithms.
problem Computational infeasibility of classical CCA methods for large-scale data.
method Unconstrained objective, stochastic gradient descent (SGD) algorithms.
result Significantly faster convergence and higher correlations than previous methods.
UK hosts 62.89% of all HYIPs, many registered as 'limited company'.
problem Understanding the prevalence and characteristics of HYIPs in the UK.
method Examined HYIPs' registration in UK, analyzed social media and payment processors, used Cox proportional regression analysis.
result HYIPs with valid UK addresses tend to have longer lifespans.
Unified normative modeling for neuroimaging phenotypes using denoising diffusion models.
problem Discarding multivariate dependence in neuroimaging pipelines.
method Denoising diffusion probabilistic models (DDPMs) with FiLM and SAINT backbones.
result Unified multivariate normative modeling with better calibration and dependence preservation.
Deep CITs test conditional independence in images, improving brain MRI scan analysis.
problem Testing conditional independence in complex, high-dimensional variables like images.
method Combines embedding maps and nonparametric CITs for feature representations.
result Valid DNCITs for brain MRI scans and behavioral traits, confirming null results.
This paper studies business cycle patterns in UK sectoral output. It analyzes the distinction between white noise processes and their non-white noise counterparts in the frequency domain and further examines the associated features and patterns for the process where white noise conditions are violated. The characterist…
Develops a Bayesian method for causal inference with partly censored time-to-event data.
problem Estimating causal effects with unobserved confounders and measurement errors in partly censored time-to-event data.
method Semiparametric Bayesian instrumental variable analysis using a two-stage Dirichlet process mixture model.
result The proposed method outperforms competing methods in simulations and real-world data analysis.
Locational Marginal Pricing aims to free UK power markets.
problem Unfree and regulated power markets.
method Implementing Locational Marginal Pricing.
result Increased economic freedom, reduced prices, decreased losses, incentivized investment.
Federated learning improves bioinformatics by sharing data legally.
problem Lack of access to diverse data in bioinformatics.
method Combines data from multiple institutions legally.
result Federated learning accelerates clinical discovery and robust exploration.
UK universities pension scheme valuation study shows high dependence on gilt yields.
problem High dependence of UK universities pension scheme on UK government bond yields.
method Analysis of USS valuations from 2014 to 2023, examination of self-sufficiency conditions, and evaluation of metrics.
result Second self-sufficiency condition amplifies gilt yield dependence, leading to inflated liabilities and excessive prudence.
Paper develops SKPD framework for signal region detection in image regression.
problem Limited research on image region detection in high-resolution image regression.
method Sparse Kronecker Product Decomposition (SKPD) framework for matrices and tensors.
result Computed solutions converge to truth with guaranteed consistency.
Study improves prediction of UK road accidents' severity using AI.
problem Improving prediction of UK road traffic accident severity.
method Combination of machine learning, econometric, and statistical methods on historical data.
result XGBoost model with RMSE of 0.176 and MAE of 0.087 outperforms naive forecasting.
We detail distributed algorithms for scalable, secure multiparty linear regression and feature selection at essentially the same speed as plaintext regression. While the core geometric ideas are simple, the recognition of their broad utility when combined is novel. Our scheme opens the door to efficient and secure geno…
Gradient-flow optimization is reinterpreted as a statistical inference problem.
problem Optimizing training duration and assessing model performance in deep learning.
method Develops a statistical framework for gradient-flow training, treating it as a random-effects model.
result Establishes asymptotic optimality for prediction and reduces reliance on validation splits.
SLOE speeds up logistic regression in high dimensions with accurate signal strength estimation.
problem Poor performance of logistic regression in high-dimensional settings.
method SLOE reparameterizes the signal strength for faster and more accurate estimation.
result SLOE provides a fast and accurate method for dimensionality correction in logistic regression.
This paper proposes non-stationary factor models for financial stress in the UK.
problem Managing financial vulnerabilities in the UK's complex financial system.
method Creation of non-stationary factor models to capture financial stress.
result Non-stationary factor models can better capture financial stress, especially tail events.
Study evaluates UK CDC schemes, finding intergenerational cross-subsidies in flat-accrual schemes and dynamic-accrual schemes can reduce but not eliminate them.
problem Intergenerational cross-subsidies in UK CDC schemes, particularly in flat-accrual schemes.
method Comparison of flat-accrual and dynamic-accrual CDC schemes, analysis of performance and level of cross-subsidies.
result Dynamic-accrual schemes can reduce but not eliminate intergenerational cross-subsidies, while flat-accrual schemes often have significant cross-subsidies.
A clinician desires to use a risk-stratification method that achieves confident risk-stratification - the risk estimates of the different patients reflect the true risks with a high probability. This allows him/her to use these risks to make accurate predictions about prognosis and decisions about screening, treatments…
In an analysis of the US, the UK, and the German stock market we find a change in the behavior based on the stock's beta values. Before 2006 risky trades were concentrated on stocks in the IT and technology sector. Afterwards risky trading takes place for stocks from the financial sector. We show that an agent-based mo…
UK's rapid vaccine rollout linked to reduced COVID-19 mortality.
problem Assessing the impact of accelerated vaccine rollout on public health outcomes.
method Flexible probabilistic models combining interrupted time series analysis and synthetic control methods with multi-output Gaussian processes.
result Substantial reduction in COVID-19 mortality with little effect on transmission rates.
Key to the imposition of appropriate minimum capital requirements on a daily basis requires accurate volatility estimation. Here, measures are presented based on discrete estimation of aggregated high frequency UK futures realisations underpinned by a continuous time framework. Squared and absolute returns are incorpor…
Measurement noise limits the advantage of nonlinear models over linear models in biomedical prediction
problem Nonlinear models vs. linear models in biomedical prediction
method Measurement reliability
result Measurement noise blurs the population-optimal predictor
Research quantifies financial exclusion risks in UK, focusing on cash infrastructure and socio-economic factors.
problem Localised financial exclusion in the UK as cash infrastructure declines.
method Developed a composite indicator using various input variables.
result Financial exclusion is more prevalent in deprived communities and affluent areas.
Study identifies two borrowing patterns in UK payday loan users.
problem Financial vulnerability of payday loan users.
method Two-state hidden Markov model (HMM) using Open Banking data.
result 36.4% of borrowers experience high-intensity exposure for 12 weeks or more.