Method fuses low and high-resolution data for better health estimates.
problem Improving high-resolution health estimates from mixed data sources.
method Fusion of unbiased low-resolution and potentially biased high-resolution data, learning a distribution consistent with sampling bias.
result Significant reduction in bias in high-resolution estimates.
Deep learning improves CVD risk prediction from health records.
problem Predicting cardiovascular disease risk from administrative health data.
method Combined survival analysis and deep learning models.
result Deep learning models outperform traditional Cox models in accuracy and explained time-to-event occurrence.
Accurate prediction of suicide risk in mental health patients remains an open problem. Existing methods including clinician judgments have acceptable sensitivity, but yield many false positives. Exploiting administrative data has a great potential, but the data has high dimensionality and redundancies in the recording …
New model maps malaria prevalence across Kenya's changing administrative boundaries.
problem Mapping disease prevalence with changing administrative boundaries.
method Combines deep learning and MCMC with aggVAE for disease mapping.
result Solves the change-of-support problem in disease surveillance.
The past decade has seen an explosion in the amount of digital information stored in electronic health records (EHR). While primarily designed for archiving patient clinical information and administrative healthcare tasks, many researchers have found secondary use of these records for various clinical informatics tasks…
Risk prediction is central to both clinical medicine and public health. While many machine learning models have been developed to predict mortality, they are rarely applied in the clinical literature, where classification tasks typically rely on logistic regression. One reason for this is that existing machine learning…
Study uses data to analyze COPD patients' impact on hospital systems.
problem Understanding and quantifying resource requirements for COPD patients.
method Combines segmentation, queuing theory, and data recovery techniques.
result Finding useful operational results from incomplete administrative data.
Germany's tax admin costs likely exceed 20% of total revenue, requiring system improvement.
problem High tax administrative costs in Germany and other jurisdictions.
method Statistical data, surveys, and a novel approach to measure total administrative cost as a percentage of total tax revenue.
result Germany's 2021 tax administrative costs likely exceeded 20% of total tax revenue.
The paper learns personalized treatment rules from observational data.
problem Developing effective treatment policies for individual patients.
method Contextual bandit approach to minimize expected risk of treatment policies.
result The proposed method outperforms physicians and baseline approaches in IV and VP administration.
The availability of a large amount of electronic health records (EHR) provides huge opportunities to improve health care service by mining these data. One important application is clinical endpoint prediction, which aims to predict whether a disease, a symptom or an abnormal lab test will happen in the future according…
A2A metric evaluates bias correction methods, reducing ATE estimation errors.
problem Selection biases in non-randomized studies of medical treatments.
method Propensity score matching (PSM) with novel metric A2A.
result Reduces ATE estimation errors by up to 90% across synthetic and real-world datasets.
Machine learning methods have gained a great deal of popularity in recent years among public administration scholars and practitioners. These techniques open the door to the analysis of text, image and other types of data that allow us to test foundational theories of public administration and to develop new theories. …
This paper addresses issues with the Brier score in administrative censoring scenarios.
problem Problems with the Brier score in administrative censoring scenarios.
method Proposes an alternative Brier score for administratively censored data.
result The administrative Brier score is valid even when censoring times can be identified from covariates.
Research funding agencies routinely use a proportion of their total revenues to support internal administration and marketing costs. The ratio of administration to total costs, referred to as the administration ratio, is highly variable and within any single fund depends on many factors including the number and average…
The occurrence of drug-drug-interactions (DDI) from multiple drug dispensations is a serious problem, both for individuals and health-care systems, since patients with complications due to DDI are likely to reenter the system at a costlier level. We present a large-scale longitudinal study (18 months) of the DDI phenom…
New model outperforms traditional disease models in forecasting COVID-19.
problem Forecasting COVID-19 spread with high accuracy and reliability.
method Developed a novel neural forecasting model called ACTS using inter-series attention.
result ACTS outperforms leading forecasters in multiple metrics.
The paper analyzes fairness of compensation-based risk-sharing schemes for fund payouts.
problem Fair allocation of payouts in an endowment contingency fund.
method Analyzes two types of administrators and general non-negative loss distributions.
result General conditions for actuarial fairness are provided.
New method simplifies data analysis.
problem Complex data analysis challenges.
method Innovative algorithm for data simplification.
result Significant reduction in analysis time.
From medical charts to national census, healthcare has traditionally operated under a paper-based paradigm. However, the past decade has marked a long and arduous transformation bringing healthcare into the digital age. Ranging from electronic health records, to digitized imaging and laboratory reports, to public healt…
New DP mechanism SWAG-PPM improves privacy in deep learning models.
problem Differential privacy struggles with real-world distributions, especially imbalanced data.
method SWAG-PPM uses a pseudo posterior distribution to downweight high-risk records.
result SWAG-PPM outperforms DP-SGD with similar privacy budget and modest utility degradation.
Study improves risk evaluation timing with right-censored reporting delays.
problem Improving risk evaluation under short observation windows due to administrative censoring.
method Jointly models parametric hazards for event and reporting processes, uses Monte Carlo expectation-maximization algorithm, and proposes transfer-learning procedure.
result Improves accuracy of timely risk evaluation under administrative censoring.
It is not clear how to target patients who are most likely to benefit from digital care management programs ex-ante, a shortcoming of current risk score based approaches. This study focuses on defining impactability by identifying those patients most likely to benefit from technology enabled care management, delivered …
Study on tax administration issues and their impact on Georgia's budget revenues.
problem Problems in revenue administration and tax rates in Georgia.
method Analyzed foreign experience and proposed a progressive tax system.
result A progressive tax system would benefit Georgia's business and economy.
Mathematical models help keep vaccine prices low.
problem Pricing COVID-19 vaccines to ensure affordability and profitability.
method Optimization and game theory approaches modeling a duopoly market.
result Government can negotiate low prices while manufacturers earn profits.
Study estimates personalized effects of maternal PM2.5 exposure on birth weight.
problem Identify critical windows and heterogeneity in maternal PM2.5 exposure effects on birth weight.
method Heterogeneous Distributed Lag Models and Bayesian Additive Regression Trees.
result Evidence of heterogeneity in PM2.5-birth weight relationship, with some dyads showing 3x larger decrease.
PHASE predicts surgical complications from physiological signals.
problem Predicting adverse surgical outcomes from physiological signals.
method Self-supervised transfer learning for physiological signals.
result PHASE outperforms other approaches in predicting five surgical complications.
Study shows how missing data from certain groups can unfairly bias risk models.
problem Data missingness without indicators of missingness can unfairly bias risk models.
method Developed an analytically tractable model of differential feature under-reporting and proposed new methods to mitigate bias.
result Under-reporting typically leads to increasing disparities in risk models.
In this paper, we consider the problem of predicting demographics of geographic units given geotagged Tweets that are composed within these units. Traditional survey methods that offer demographics estimates are usually limited in terms of geographic resolution, geographic boundaries, and time intervals. Thus, it would…
Study uses machine learning to predict future health from various health data types.
problem Predicting future health using diverse health data types.
method Applied machine learning (neural networks and XGBoost) to longitudinal data from 6830 individuals.
result Health-related measures were the strongest predictors of future health status, while genetic data performed poorly.
This article was withdrawn by the arXiv.org administrators since it plagiarizes math.AT/0401211.
This article was withdrawn by the arXiv.org administrators since it plagiarizes math.GT/0011056.
Although aviation accidents are rare, safety incidents occur more frequently and require a careful analysis to detect and mitigate risks in a timely manner. Analyzing safety incidents using operational data and producing event-based explanations is invaluable to airline companies as well as to governing organizations s…
Paper proposes a method to estimate confidence bands for survival random forests.
problem No statistically valid and computationally feasible approach for estimating confidence bands for survival random forests.
method Extending recent developments in infinite-order incomplete U-statistics, the paper proposes an unbiased confidence band estimation.
result The proposed method accurately estimates the confidence band and achieves desired coverage rate.
Study shows Lula's Zero Hunger program reduced income inequality in Brazil.
problem Income inequality in Brazil during Lula's administration.
method Breakpoint regression analysis using detailed descriptive statistics.
result The Zero Hunger program substantially reduced income inequality and provided income security for the poor.
HealthSyn generates synthetic user behavior data for health interventions.
problem Lack of representative data for testing AI health interventions.
method Uses Markov processes to simulate diverse user actions, generating logs for ML algorithms.
result Synthetic data can be used to develop, test, and evaluate ML algorithms and RL-based interventions.
This paper has been withdrawn by arXiv administrators because of disputed claims of authorship among former collaborators
DeepCoDA provides personalized interpretability for complex health data.
problem Interpreting complex health data, especially compositional data, is challenging.
method DeepCoDA framework for high-dimensional compositional data, personalized interpretability through patient-specific weights.
result DeepCoDA maintains state-of-the-art performance and provides coherent, personalized interpretations.
Policy shifts between Trump and Biden impact ESG investments, creating volatility.
problem Dramatic policy shifts between Trump and Biden administrations affect ESG investments.
method Analyzes contrasting policies of Trump and Biden administrations and their impacts on ESG investments.
result Policy changes significantly influence ESG investments, leading to volatility and portfolio reassessment.
Repository tackles fake health news in cancer research.
problem Spread of fake health news over the internet.
method Developed comprehensive FakeHealth repository with rich features and detailed explanations.
result Repository helps in understanding and validating health fake news datasets.
This version withdrawn by arXiv administrators because the submitter did not have the right to agree to our license at the time of submission.
Automatically assesses the quality of online health articles.
problem Lack of automated tools to evaluate the quality of online health information.
method Data mining approach using 10 quality criteria and feature selection.
result Classifier achieved 84%-90% accuracy on 10 criteria.
We present the Network-based Biased Tree Ensembles (NetBiTE) method for drug sensitivity prediction and drug sensitivity biomarker identification in cancer using a combination of prior knowledge and gene expression data. Our devised method consists of a biased tree ensemble that is built according to a probabilistic bi…
Equity-Directed Bootstrapping improves model performance across groups in imbalanced datasets.
problem Improving model performance across different groups in imbalanced datasets.
method Equity-Directed Bootstrapping to balance training data with respect to both labels and group identity.
result The equity-directed bootstrap brings test set sensitivities and specificities closer to satisfying the equal odds criterion.
Bayesian model predicts patient survival from sparse EHR data.
problem Analyzing EHR data with few samples and diverse information.
method Nonparametric probabilistic model using Bayesian trees.
result Improved survival trajectory predictions on patient data.
Semiparametric STAR model improves mental health data analysis.
problem Overdispersed, zero-inflated, bounded count data in self-reported mental health surveys.
method STAR transformation and rounding of latent Gaussian model, nonparametric transformation estimation, EM algorithm for maximum likelihood.
result Substantial improvements in goodness-of-fit compared to existing models.
New algorithm forecasts health indicators for better equipment lifespan prediction.
problem Improving equipment lifespan prediction through health indicator forecasting.
method Generative + scenario matching approach using Gaussian Process.
result Superior performance compared to existing methods.
This study is motivated by the magnitude of the problem of Louisiana high school dropout and its negative impacts on individual and public well-being. Our goal is to predict students who are at risk of high school dropout, by examining Louisiana administrative dataset. Due to the imbalanced nature of the dataset, imbal…
Proposes a model to handle mobile health data with irregular measurements.
problem Handling heterogeneous, multi-resolution data in mobile health.
method Individualized dynamic latent factor model for irregular multi-resolution time series data.
result Superior performance compared to existing methods in simulation and smartwatch data applications.