Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

11223344 · Jul 202019922001200920182026
48 results for public administration

This paper provides a guide to using machine learning in public administration.

problem Lack of clarity in proper use and potential pitfalls of machine learning methods.
method Provides a foundational view of machine learning and demonstrates its use in public administration research.
result Machine learning techniques can enrich public administration research and practice.

Germany's tax admin costs likely exceed 20% of total revenue, requiring system improvement.

problem High tax administrative costs in Germany and other jurisdictions.
method Statistical data, surveys, and a novel approach to measure total administrative cost as a percentage of total tax revenue.
result Germany's 2021 tax administrative costs likely exceeded 20% of total tax revenue.

Policy shifts between Trump and Biden impact ESG investments, creating volatility.

problem Dramatic policy shifts between Trump and Biden administrations affect ESG investments.
method Analyzes contrasting policies of Trump and Biden administrations and their impacts on ESG investments.
result Policy changes significantly influence ESG investments, leading to volatility and portfolio reassessment.

Study predicts high school dropout risk in Louisiana using imbalanced learning techniques.

problem Predicting high school dropout risk in Louisiana.
method Applied imbalanced learning techniques including resampling, case weighting, and cost-sensitive learning.
result Imbalanced learning techniques improve recall but decrease precision.

This article presents results from the first statistically significant study of cost escalation in transportation infrastructure projects. Based on a sample of 258 transportation infrastructure projects worth US$90 billion and representing different project types, geographical regions, and historical periods, it is fou…

2013-03-06abs ↗pdf ↗

Study shows how missing data from certain groups can unfairly bias risk models.

problem Data missingness without indicators of missingness can unfairly bias risk models.
method Developed an analytically tractable model of differential feature under-reporting and proposed new methods to mitigate bias.
result Under-reporting typically leads to increasing disparities in risk models.

This paper addresses issues with the Brier score in administrative censoring scenarios.

problem Problems with the Brier score in administrative censoring scenarios.
method Proposes an alternative Brier score for administratively censored data.
result The administrative Brier score is valid even when censoring times can be identified from covariates.

Study improves risk evaluation timing with right-censored reporting delays.

problem Improving risk evaluation under short observation windows due to administrative censoring.
method Jointly models parametric hazards for event and reporting processes, uses Monte Carlo expectation-maximization algorithm, and proposes transfer-learning procedure.
result Improves accuracy of timely risk evaluation under administrative censoring.

Bayesian models forecast COVID-19 hospitalizations at single sites.

problem Forecasting daily COVID-19 hospitalizations at a single hospital.
method Hierarchical Bayesian models with generalized Poisson likelihood and autoregressive/Gaussian process latent processes.
result Demonstrated superior performance compared to baselines in public datasets.

A2A metric evaluates bias correction methods, reducing ATE estimation errors.

problem Selection biases in non-randomized studies of medical treatments.
method Propensity score matching (PSM) with novel metric A2A.
result Reduces ATE estimation errors by up to 90% across synthetic and real-world datasets.

This study simulates the evolution of artificial economies in order to understand the tax relevance of administrative boundaries in the quality of life of its citizens. The modeling involves the construction of a computational algorithm, which includes citizens, bounded into families; firms and governments; all of them…

2015-10-16abs ↗pdf ↗

New DP mechanism SWAG-PPM improves privacy in deep learning models.

problem Differential privacy struggles with real-world distributions, especially imbalanced data.
method SWAG-PPM uses a pseudo posterior distribution to downweight high-risk records.
result SWAG-PPM outperforms DP-SGD with similar privacy budget and modest utility degradation.

Deep learning improves CVD risk prediction from health records.

problem Predicting cardiovascular disease risk from administrative health data.
method Combined survival analysis and deep learning models.
result Deep learning models outperform traditional Cox models in accuracy and explained time-to-event occurrence.

Paper proposes a recursive PLS model for optimal response to security threats.

problem Optimal response to security threats after violations have occurred.
method Recursive Partial Least Squares (PLS) model with factorial analysis of security events.
result The model optimally estimates security administrators' responses to threats.

Dataset of Italian municipalities' income taxes from 2007-2011.

problem Understanding the economic structure of Italian municipalities.
method Annual aggregated income taxes of all Italian municipalities, clustered by regions and provinces.
result Data useful for economic comparisons and understanding municipal structures.

System automates identification of cancer drug repurposing from PubMed.

problem Manual extraction of cancer drug repurposing evidence from scientific publications is infeasible.
method NLP pipeline including querying, filtering, entity extraction, classification, and study type classification.
result Automated system extracts cancer drug repurposing evidence from PubMed abstracts.

New method uses geometric mean to avoid non-collapsibility in case-control studies.

problem Non-collapsibility of odds ratio under outcome-dependent sampling.
method Proposes geometric mean aggregation to avoid non-collapsibility and provides estimation and inference methods.
result Geometric odds ratio is collapsible under outcome-dependent sampling.

Predicts stock price changes based on clinical trial announcements.

problem Forecasting the impact of clinical trial results on pharma stock prices.
method BERT for sentiment analysis, Temporal Fusion Transformer for forecasting, graph convolution network for event relationships, gradient boosting for price change prediction.
result Identifies two crucial factors: drug portfolio size and network effect of related events.

Accurate prediction of suicide risk in mental health patients remains an open problem. Existing methods including clinician judgments have acceptable sensitivity, but yield many false positives. Exploiting administrative data has a great potential, but the data has high dimensionality and redundancies in the recording …

2016-05-03abs ↗pdf ↗

Paper uses CT-IV to estimate causal effects in non-randomized settings.

problem Estimating causal effects in non-randomized observational studies.
method Modified Causal Tree (CT-IV) algorithm combining CART and IV framework.
result Demonstrates efficiency in handling heterogeneity of causal effects.

Study shows Lula's Zero Hunger program reduced income inequality in Brazil.

problem Income inequality in Brazil during Lula's administration.
method Breakpoint regression analysis using detailed descriptive statistics.
result The Zero Hunger program substantially reduced income inequality and provided income security for the poor.

Equity-Directed Bootstrapping improves model performance across groups in imbalanced datasets.

problem Improving model performance across different groups in imbalanced datasets.
method Equity-Directed Bootstrapping to balance training data with respect to both labels and group identity.
result The equity-directed bootstrap brings test set sensitivities and specificities closer to satisfying the equal odds criterion.

NetBiTE predicts drug sensitivity and identifies biomarkers in cancer.

problem Predicting drug sensitivity and identifying biomarkers in cancer.
method NetBiTE combines prior knowledge and gene expression data using a biased tree ensemble approach.
result NetBiTE outperforms RF in predicting IC50 drug sensitivity for drugs targeting membrane receptor pathways.

Study uses data to analyze COPD patients' impact on hospital systems.

problem Understanding and quantifying resource requirements for COPD patients.
method Combines segmentation, queuing theory, and data recovery techniques.
result Finding useful operational results from incomplete administrative data.

Paper proposes a method to estimate confidence bands for survival random forests.

problem No statistically valid and computationally feasible approach for estimating confidence bands for survival random forests.
method Extending recent developments in infinite-order incomplete U-statistics, the paper proposes an unbiased confidence band estimation.
result The proposed method accurately estimates the confidence band and achieves desired coverage rate.

Neural networks outperform logistic regression for predicting HF readmission.

problem Predicting 30-day all-cause readmission in heart failure patients.
method Used a large administrative claims dataset to compare neural network models (RNNCRF) with logistic regression models (LASSO) for predicting readmission.
result RNNCRF model achieved best performance with 0.642 AUC, while logistic regression with LASSO had equal performance.