This paper provides a guide to using machine learning in public administration.
problem Lack of clarity in proper use and potential pitfalls of machine learning methods.
method Provides a foundational view of machine learning and demonstrates its use in public administration research.
result Machine learning techniques can enrich public administration research and practice.
Germany's tax admin costs likely exceed 20% of total revenue, requiring system improvement.
problem High tax administrative costs in Germany and other jurisdictions.
method Statistical data, surveys, and a novel approach to measure total administrative cost as a percentage of total tax revenue.
result Germany's 2021 tax administrative costs likely exceeded 20% of total tax revenue.
Research funding agencies routinely use a proportion of their total revenues to support internal administration and marketing costs. The ratio of administration to total costs, referred to as the administration ratio, is highly variable and within any single fund depends on many factors including the number and average…
Policy shifts between Trump and Biden impact ESG investments, creating volatility.
problem Dramatic policy shifts between Trump and Biden administrations affect ESG investments.
method Analyzes contrasting policies of Trump and Biden administrations and their impacts on ESG investments.
result Policy changes significantly influence ESG investments, leading to volatility and portfolio reassessment.
New model maps malaria prevalence across Kenya's changing administrative boundaries.
problem Mapping disease prevalence with changing administrative boundaries.
method Combines deep learning and MCMC with aggVAE for disease mapping.
result Solves the change-of-support problem in disease surveillance.
We demonstrate by mathematical analysis and systematic computer simulations that redistribution can lead to sustainable growth in a society. The human capital dynamics of each agent is described by a stochastic multiplicative process which, in the long run, leads to the destruction of individual human capital and the e…
Study predicts high school dropout risk in Louisiana using imbalanced learning techniques.
problem Predicting high school dropout risk in Louisiana.
method Applied imbalanced learning techniques including resampling, case weighting, and cost-sensitive learning.
result Imbalanced learning techniques improve recall but decrease precision.
This article presents results from the first statistically significant study of cost escalation in transportation infrastructure projects. Based on a sample of 258 transportation infrastructure projects worth US$90 billion and representing different project types, geographical regions, and historical periods, it is fou…
Method fuses low and high-resolution data for better health estimates.
problem Improving high-resolution health estimates from mixed data sources.
method Fusion of unbiased low-resolution and potentially biased high-resolution data, learning a distribution consistent with sampling bias.
result Significant reduction in bias in high-resolution estimates.
Mathematical models help keep vaccine prices low.
problem Pricing COVID-19 vaccines to ensure affordability and profitability.
method Optimization and game theory approaches modeling a duopoly market.
result Government can negotiate low prices while manufacturers earn profits.
The paper analyzes fairness of compensation-based risk-sharing schemes for fund payouts.
problem Fair allocation of payouts in an endowment contingency fund.
method Analyzes two types of administrators and general non-negative loss distributions.
result General conditions for actuarial fairness are provided.
Study shows how missing data from certain groups can unfairly bias risk models.
problem Data missingness without indicators of missingness can unfairly bias risk models.
method Developed an analytically tractable model of differential feature under-reporting and proposed new methods to mitigate bias.
result Under-reporting typically leads to increasing disparities in risk models.
Network science reveals corruption risk in EU procurement markets.
problem Identifying corruption risk in EU procurement markets.
method Analyzing a large dataset of public procurement contracts using network science.
result Corruption risk is clustered and varies by country, not just by market core or periphery.
This paper addresses issues with the Brier score in administrative censoring scenarios.
problem Problems with the Brier score in administrative censoring scenarios.
method Proposes an alternative Brier score for administratively censored data.
result The administrative Brier score is valid even when censoring times can be identified from covariates.
Study improves risk evaluation timing with right-censored reporting delays.
problem Improving risk evaluation under short observation windows due to administrative censoring.
method Jointly models parametric hazards for event and reporting processes, uses Monte Carlo expectation-maximization algorithm, and proposes transfer-learning procedure.
result Improves accuracy of timely risk evaluation under administrative censoring.
Bayesian models forecast COVID-19 hospitalizations at single sites.
problem Forecasting daily COVID-19 hospitalizations at a single hospital.
method Hierarchical Bayesian models with generalized Poisson likelihood and autoregressive/Gaussian process latent processes.
result Demonstrated superior performance compared to baselines in public datasets.
A2A metric evaluates bias correction methods, reducing ATE estimation errors.
problem Selection biases in non-randomized studies of medical treatments.
method Propensity score matching (PSM) with novel metric A2A.
result Reduces ATE estimation errors by up to 90% across synthetic and real-world datasets.
New model outperforms traditional disease models in forecasting COVID-19.
problem Forecasting COVID-19 spread with high accuracy and reliability.
method Developed a novel neural forecasting model called ACTS using inter-series attention.
result ACTS outperforms leading forecasters in multiple metrics.
Study on tax administration issues and their impact on Georgia's budget revenues.
problem Problems in revenue administration and tax rates in Georgia.
method Analyzed foreign experience and proposed a progressive tax system.
result A progressive tax system would benefit Georgia's business and economy.
This paper explains tax policy for crypto assets in a rapidly evolving tech landscape.
problem Rapid technological changes in crypto assets create regulatory and tax policy blind spots.
method Explains principles of crypto assets, their technology, and tax issues.
result Tax policies are lagging behind innovation in blockchain and crypto.
This study simulates the evolution of artificial economies in order to understand the tax relevance of administrative boundaries in the quality of life of its citizens. The modeling involves the construction of a computational algorithm, which includes citizens, bounded into families; firms and governments; all of them…
New method simplifies data analysis.
problem Complex data analysis challenges.
method Innovative algorithm for data simplification.
result Significant reduction in analysis time.
New DP mechanism SWAG-PPM improves privacy in deep learning models.
problem Differential privacy struggles with real-world distributions, especially imbalanced data.
method SWAG-PPM uses a pseudo posterior distribution to downweight high-risk records.
result SWAG-PPM outperforms DP-SGD with similar privacy budget and modest utility degradation.
This article was withdrawn by the arXiv.org administrators since it plagiarizes math.AT/0401211.
This article was withdrawn by the arXiv.org administrators since it plagiarizes math.GT/0011056.
Deep learning improves CVD risk prediction from health records.
problem Predicting cardiovascular disease risk from administrative health data.
method Combined survival analysis and deep learning models.
result Deep learning models outperform traditional Cox models in accuracy and explained time-to-event occurrence.
Risk prediction is central to both clinical medicine and public health. While many machine learning models have been developed to predict mortality, they are rarely applied in the clinical literature, where classification tasks typically rely on logistic regression. One reason for this is that existing machine learning…
This paper has been withdrawn by arXiv administrators because of disputed claims of authorship among former collaborators
The appeal of metric evaluation of research impact has attracted considerable interest in recent times. Although the public at large and administrative bodies are much interested in the idea, scientists and other researchers are much more cautious, insisting that metrics are but an auxiliary instrument to the qualitati…
Paper proposes a recursive PLS model for optimal response to security threats.
problem Optimal response to security threats after violations have occurred.
method Recursive Partial Least Squares (PLS) model with factorial analysis of security events.
result The model optimally estimates security administrators' responses to threats.
The occurrence of drug-drug-interactions (DDI) from multiple drug dispensations is a serious problem, both for individuals and health-care systems, since patients with complications due to DDI are likely to reenter the system at a costlier level. We present a large-scale longitudinal study (18 months) of the DDI phenom…
Dataset of Italian municipalities' income taxes from 2007-2011.
problem Understanding the economic structure of Italian municipalities.
method Annual aggregated income taxes of all Italian municipalities, clustered by regions and provinces.
result Data useful for economic comparisons and understanding municipal structures.
System automates identification of cancer drug repurposing from PubMed.
problem Manual extraction of cancer drug repurposing evidence from scientific publications is infeasible.
method NLP pipeline including querying, filtering, entity extraction, classification, and study type classification.
result Automated system extracts cancer drug repurposing evidence from PubMed abstracts.
New method uses geometric mean to avoid non-collapsibility in case-control studies.
problem Non-collapsibility of odds ratio under outcome-dependent sampling.
method Proposes geometric mean aggregation to avoid non-collapsibility and provides estimation and inference methods.
result Geometric odds ratio is collapsible under outcome-dependent sampling.
Predicts stock price changes based on clinical trial announcements.
problem Forecasting the impact of clinical trial results on pharma stock prices.
method BERT for sentiment analysis, Temporal Fusion Transformer for forecasting, graph convolution network for event relationships, gradient boosting for price change prediction.
result Identifies two crucial factors: drug portfolio size and network effect of related events.
Optimizes quantum circuits using evolutionary strategies.
problem Optimizing quantum circuits for efficiency.
method Uses evolution strategies to optimize circuits.
result Improves quantum circuit performance.
Accurate prediction of suicide risk in mental health patients remains an open problem. Existing methods including clinician judgments have acceptable sensitivity, but yield many false positives. Exploiting administrative data has a great potential, but the data has high dimensionality and redundancies in the recording …
Paper uses CT-IV to estimate causal effects in non-randomized settings.
problem Estimating causal effects in non-randomized observational studies.
method Modified Causal Tree (CT-IV) algorithm combining CART and IV framework.
result Demonstrates efficiency in handling heterogeneity of causal effects.
Study shows Lula's Zero Hunger program reduced income inequality in Brazil.
problem Income inequality in Brazil during Lula's administration.
method Breakpoint regression analysis using detailed descriptive statistics.
result The Zero Hunger program substantially reduced income inequality and provided income security for the poor.
Modeling student course choices using latent variables.
problem Understanding student enrollment patterns in large universities.
method Probabilistic approach based on multilabel classification and mixture models.
result Demonstrated the model's ability to infer student interests guiding enrollment decisions.
Equity-Directed Bootstrapping improves model performance across groups in imbalanced datasets.
problem Improving model performance across different groups in imbalanced datasets.
method Equity-Directed Bootstrapping to balance training data with respect to both labels and group identity.
result The equity-directed bootstrap brings test set sensitivities and specificities closer to satisfying the equal odds criterion.
NetBiTE predicts drug sensitivity and identifies biomarkers in cancer.
problem Predicting drug sensitivity and identifying biomarkers in cancer.
method NetBiTE combines prior knowledge and gene expression data using a biased tree ensemble approach.
result NetBiTE outperforms RF in predicting IC50 drug sensitivity for drugs targeting membrane receptor pathways.
Framework for fast CAT calibration and administration using AutoML and IRT.
problem Calibrating and administering large-scale CAT tests with limited data.
method AutoIRT (AutoML + IRT) for calibration, BanditCAT for administration.
result Framework successfully launched new item types on DET practice test.
Study uses data to analyze COPD patients' impact on hospital systems.
problem Understanding and quantifying resource requirements for COPD patients.
method Combines segmentation, queuing theory, and data recovery techniques.
result Finding useful operational results from incomplete administrative data.
Paper proposes a method to estimate confidence bands for survival random forests.
problem No statistically valid and computationally feasible approach for estimating confidence bands for survival random forests.
method Extending recent developments in infinite-order incomplete U-statistics, the paper proposes an unbiased confidence band estimation.
result The proposed method accurately estimates the confidence band and achieves desired coverage rate.
Fine-grained event tagging system for SEC 8-K filings improves precision to 96%.
problem Coarse SEC item codes mislabel routine and significant events.
method Two-stage system tagging 8-K disclosures against a 119-event taxonomy.
result LLM judge finds precision rises to 96% with quality scores.
Neural networks outperform logistic regression for predicting HF readmission.
problem Predicting 30-day all-cause readmission in heart failure patients.
method Used a large administrative claims dataset to compare neural network models (RNNCRF) with logistic regression models (LASSO) for predicting readmission.
result RNNCRF model achieved best performance with 0.642 AUC, while logistic regression with LASSO had equal performance.
From medical charts to national census, healthcare has traditionally operated under a paper-based paradigm. However, the past decade has marked a long and arduous transformation bringing healthcare into the digital age. Ranging from electronic health records, to digitized imaging and laboratory reports, to public healt…