Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

2805608391,119 · Jun 202019922001200920182026
48 results for administrative healthcare data

Study predicts which patients will benefit from digital health interventions.

problem Unclear targeting of patients for digital care management programs.
method Analyzed claims data, combined with sociodemographic and app-generated data. Created two models: cost prediction and impactability classification.
result Random forest model accurately categorized patients as impactable or not, achieving 71.9% accuracy.

Machine learning in healthcare faces challenges due to complex data attributes.

problem Complex data attributes hinder accurate insights from machine learning models.
method Discusses preprocessing, model building, and interpretation challenges.
result Understanding data attributes is crucial for successful machine learning in healthcare.

LMM predicts healthcare costs and risks with improved accuracy.

problem Wasteful healthcare spending and inefficiencies in risk prediction.
method Generative pre-trained transformer trained on patient event sequences.
result Improves cost prediction by 14.1% and chronic conditions prediction by 1.9%.

Germany's tax admin costs likely exceed 20% of total revenue, requiring system improvement.

problem High tax administrative costs in Germany and other jurisdictions.
method Statistical data, surveys, and a novel approach to measure total administrative cost as a percentage of total tax revenue.
result Germany's 2021 tax administrative costs likely exceeded 20% of total tax revenue.

Neural networks outperform logistic regression for predicting HF readmission.

problem Predicting 30-day all-cause readmission in heart failure patients.
method Used a large administrative claims dataset to compare neural network models (RNNCRF) with logistic regression models (LASSO) for predicting readmission.
result RNNCRF model achieved best performance with 0.642 AUC, while logistic regression with LASSO had equal performance.

This paper provides a guide to using machine learning in public administration.

problem Lack of clarity in proper use and potential pitfalls of machine learning methods.
method Provides a foundational view of machine learning and demonstrates its use in public administration research.
result Machine learning techniques can enrich public administration research and practice.

Study analyzes costs of managing research funds, developing a model for optimal administration.

problem High variability in administration costs among research funding agencies.
method Identified standard agency activities, developed a model estimating optimum portfolio success rate and administration ratio.
result Model estimates optimum portfolio success rate and administration ratio based on input variables.

Randomized machine learning methods improve suicide risk prediction from administrative data.

problem Accurately predicting suicide risk in mental health patients.
method Three randomized machine learning techniques: random forests, gradient boosting machines, and deep neural nets with dropout.
result Randomized methods outperform traditional approaches in predicting suicide risk with robustness against data redundancies.

This paper addresses issues with the Brier score in administrative censoring scenarios.

problem Problems with the Brier score in administrative censoring scenarios.
method Proposes an alternative Brier score for administratively censored data.
result The administrative Brier score is valid even when censoring times can be identified from covariates.

System automates identification of cancer drug repurposing from PubMed.

problem Manual extraction of cancer drug repurposing evidence from scientific publications is infeasible.
method NLP pipeline including querying, filtering, entity extraction, classification, and study type classification.
result Automated system extracts cancer drug repurposing evidence from PubMed abstracts.

Dataset of Italian municipalities' income taxes from 2007-2011.

problem Understanding the economic structure of Italian municipalities.
method Annual aggregated income taxes of all Italian municipalities, clustered by regions and provinces.
result Data useful for economic comparisons and understanding municipal structures.

Study improves risk evaluation timing with right-censored reporting delays.

problem Improving risk evaluation under short observation windows due to administrative censoring.
method Jointly models parametric hazards for event and reporting processes, uses Monte Carlo expectation-maximization algorithm, and proposes transfer-learning procedure.
result Improves accuracy of timely risk evaluation under administrative censoring.

Deep learning improves CVD risk prediction from health records.

problem Predicting cardiovascular disease risk from administrative health data.
method Combined survival analysis and deep learning models.
result Deep learning models outperform traditional Cox models in accuracy and explained time-to-event occurrence.

MPVAA learns holistic patient representations from mixed healthcare data.

problem Learning personalized patient representations from heterogeneous healthcare data.
method Mixed Pooling Multi-View Attention Autoencoder (MPVAA) that integrates non-linear relationships among multiple data modalities.
result MPVAA generates more effective patient representations than state-of-the-art methods.

Study uses data to analyze COPD patients' impact on hospital systems.

problem Understanding and quantifying resource requirements for COPD patients.
method Combines segmentation, queuing theory, and data recovery techniques.
result Finding useful operational results from incomplete administrative data.

Broad learning integrates diverse healthcare data for diagnostics and precision medicine.

problem Integrating various types of healthcare data for better diagnostics and personalized medicine.
method Fusing multi-view data including scalar, tensor, graph, and sequence data for knowledge discovery and machine learning tasks.
result Accurate user profiles and brain connectivity patterns can be created for improved diagnostics and personalized medicine.

Paper proposes a method to estimate confidence bands for survival random forests.

problem No statistically valid and computationally feasible approach for estimating confidence bands for survival random forests.
method Extending recent developments in infinite-order incomplete U-statistics, the paper proposes an unbiased confidence band estimation.
result The proposed method accurately estimates the confidence band and achieves desired coverage rate.

VHGM-MAE generates synthetic humans from healthcare data.

problem Handling high-dimensional, sparse healthcare data with missing values.
method Masked autoencoder (MAE) tailored for healthcare data, addressing heterogeneity, missingness, and high-dimensionality.
result VHGM-MAE outperforms existing methods in missing value imputation and synthetic data generation.

Study shows Lula's Zero Hunger program reduced income inequality in Brazil.

problem Income inequality in Brazil during Lula's administration.
method Breakpoint regression analysis using detailed descriptive statistics.
result The Zero Hunger program substantially reduced income inequality and provided income security for the poor.

Policy shifts between Trump and Biden impact ESG investments, creating volatility.

problem Dramatic policy shifts between Trump and Biden administrations affect ESG investments.
method Analyzes contrasting policies of Trump and Biden administrations and their impacts on ESG investments.
result Policy changes significantly influence ESG investments, leading to volatility and portfolio reassessment.

Deep generative model for healthcare data identifies coherent substructures and mutational clusters.

problem Analytical challenges in healthcare data, including sparsity, missingness, and small sample sizes.
method Proposes a deep generative Bayesian model with collapsed Gibbs sampling for multinomial count data.
result Identifies coherent substructures and biologically meaningful mutational clusters in cancer data.

Paper proposes a recursive PLS model for optimal response to security threats.

problem Optimal response to security threats after violations have occurred.
method Recursive Partial Least Squares (PLS) model with factorial analysis of security events.
result The model optimally estimates security administrators' responses to threats.

Study improves healthcare time series imputation by considering structured missingness.

problem Structured missingness in clinical data impacts time series imputation models.
method Analysis of different masking strategies on imputation methods using PhysioNet Challenge 2012 dataset.
result Masking choices significantly affect imputation accuracy and clinical prediction.

Isthmus platform simplifies ML/AI integration in healthcare.

problem Challenges in deploying ML models in healthcare due to data quality, regulatory, and security issues.
method Turnkey, cloud-based platform addressing data quality, clinical relevance, and regulatory compliance.
result Reduces time to market for operationalizing ML/AI in healthcare.

Equity-Directed Bootstrapping improves model performance across groups in imbalanced datasets.

problem Improving model performance across different groups in imbalanced datasets.
method Equity-Directed Bootstrapping to balance training data with respect to both labels and group identity.
result The equity-directed bootstrap brings test set sensitivities and specificities closer to satisfying the equal odds criterion.

Research predicts healthcare index movements using historical OHLC data.

problem Predicting the directional movement of healthcare indices based on historical data.
method Supervised classification task with a one-step-ahead rolling window, using a diverse feature set including OHLC ratios.
result Robust predictive performance with accuracy exceeding 0.8 and Matthews correlation coefficients above 0.6, highlighting the importance of nowcasting features.

A new method uses Hamiltonian Monte Carlo for imputation and augmentation of healthcare data.

problem Missing values in clinical studies lead to biased results and loss of statistical power.
method Folded Hamiltonian Monte Carlo (F-HMC) with Bayesian inference to handle high-dimensional, small sample size datasets.
result The method enriches the quality of data in precision, accuracy, recall, F1 score, and propensity metric.