Study shows racial bias in health data, which can be reduced with simple techniques.
problem Racial bias in health indicators measured by the Medical Expenditure Panel Survey (MEPS).
method Used publicly available and nationally representative MEPS data to show bias in predictive models for care management.
result Racial bias can be significantly reduced using simple mitigation techniques.
Study decomposes racial healthcare disparities via shifts in mediator distributions.
problem Racial disparities in healthcare expenditures and their underlying drivers.
method Framework decomposing disparities into mediator distribution shifts and residual components, using MEPS data.
result Substantial disparities persist even when mediators are equalized, suggesting unmeasured or structural factors.
Preserving the privacy of individuals by protecting their sensitive attributes is an important consideration during microdata release. However, it is equally important to preserve the quality or utility of the data for at least some targeted workloads. We propose a novel framework for privacy preservation based on the …
We empirically investigate distributions of individual consumption expenditure f or four commodity categories conditional on fixed income levels. The data stems from the Family Expenditure Survey carried out annually in the United Kingdom. W e use graphical techniques to test for normality and lognormality of these dis…
Early diagnosis is important for type 2 diabetes (T2D) to improve patient prognosis, prevent complications and reduce long-term treatment costs. We present a novel risk profiling approach based exclusively on health expenditure data that is available to Belgian mutual health insurers. We used expenditure data related t…
Simulation framework assesses ROI of chronic disease adherence and policy timing.
problem Uncertainty in ROI of adherence-enhancing interventions under heterogeneous patient behavior and socioeconomic variation.
method Simulation-based framework integrating disease progression, time-varying adherence, and policy timing.
result Early and adaptive interventions yield highest ROI, exceeding 20% under certain conditions.
Surveying machine learning methods for economic forecasting.
problem Improving accuracy of economic forecasts using machine learning.
method Nowcasting, textual data, panel and tensor data, high-dimensional Granger causality tests, time series cross-validation, classification with economic losses.
result Recent advances in machine learning methods enhance economic forecasting accuracy.
This work presents an empirical study of the evolution of the consumer expenditure distribution in India during 1982-2007. We have used the National Sample Survey Organization data and analysed the expenditure distribution for the urban and rural sectors. It is found that this distribution is a mixture of two distribut…
A statistical description and model of individual healthcare expenditures in the US has been developed for measuring value in healthcare. We find evidence that healthcare expenditures are quantifiable as an infusion-diffusion process, which can be thought of intuitively as a steady change in the intensity of treatment …
ChatGPT scores corporate investment plans, predicting future spending and returns.
problem Measuring and predicting corporate investment plans.
method Created a firm-level ChatGPT investment score based on conference calls.
result The investment score predicts future capital expenditures and returns.
High quality risk adjustment in health insurance markets weakens insurer incentives to engage in inefficient behavior to attract lower-cost enrollees. We propose a novel methodology based on Markov Chain Monte Carlo methods to improve risk adjustment by clustering diagnostic codes into risk groups optimal for health ex…
Survey of deep learning methods for medical anomaly detection.
problem Medical anomaly detection using machine learning.
method Thorough review of deep learning techniques across various medical domains.
result Comparison and contrast of deep learning models and their limitations.
LHIEM model predicts health, income, and employment over years.
problem Lack of path dependency in health policy simulations.
method Discrete-time microsimulation with Markov chain modules.
result Validates health care financing proposal through detailed modeling.
We analyze three sets of income data: the US Panel Study of Income Dynamics PSID), the British Household Panel Survey (BHPS), and the German Socio-Economic Panel (GSOEP). It is shown that the empirical income distribution is consistent with a two-parameter lognormal function for the low-middle income group (97%-99% of …
Proposes a robust EM algorithm for analyzing incomplete panel count data.
problem Missing reports in panel count data.
method Functional EM algorithm for non-parametric counting process mean function estimation.
result Robust to misspecification of Poisson process assumption and missing completely at random.
Most papers which explored so far macroeconomic variables took into account income and wealth. Equally important as the previous macroeconomic variables is the expenditure or consumption, which shows the amount of goods and services that a person or a household purchased. Using statistical distributions from Physics, s…
The study introduces backward baselines to distinguish past prediction from future prediction in machine learning models.
problem Differentiating between past and future prediction in machine learning models.
method Theoretical, empirical, and normative arguments support a family of simple and efficient statistical tests called backward baselines.
result The study provides a meaningful backward baseline for auditing black-box prediction systems.
Survey on factor models and their applications in econometrics.
problem Estimating low-rank structures in high-dimensional models.
method Low-rank recovery techniques for factor model estimation.
result New insights into factor model applications in econometrics.
Paper proposes a human-algorithm approach to reduce medical device recall risk and workload.
problem High recall rate and regulatory workload in FDA's 510(k) pathway.
method Developed machine learning models to estimate recall risk and proposed a data-driven clearance policy.
result Conservative evaluation of policy shows a 32.9% improvement in recall rate and 40.5% reduction in workload.
Digital personas improve survey results for stable attributes but fail for subjective responses.
problem When can digital personas reliably approximate human survey findings?
method Using LISS panel, constructed personas from background variables and survey histories, tested against held-out post-cutoff answers.
result Digital personas improve alignment with human response distributions for stable attributes but fail for subjective responses.
Survey on understanding neural networks for medical applications.
problem Black-box nature of deep neural networks hinders their use in critical applications.
method Comprehensive review of interpretability studies in neural networks.
result Interpretability research is crucial for the acceptance of neural networks in medical diagnosis.
Quantifying the improvement in human living standard, as well as the city growth in developing countries, is a challenging problem due to the lack of reliable economic data. Therefore, there is a fundamental need for alternate, largely unsupervised, computational methods that can estimate the economic conditions in the…
Super learner with Huber loss improves cost prediction and causal effect estimation in healthcare expenditure data.
problem Challenges in modeling healthcare expenditure distributions with standard super learning methods.
method Proposes a super learner using Huber loss, a robust loss function that down-weights outliers.
result Demonstrates appreciable finite-sample gains in cost prediction and causal effect estimation.
Paper develops NN models for diabetes screening using NHANES data.
problem Developing accurate predictive models for diabetes in diverse populations.
method Proposes a neural network framework with survey weights, uncertainty quantification.
result Robust risk score models for diabetes in US population.
Financial planners helped preserve and increase household net financial assets during the Great Recession.
problem Impact of financial planners on household net financial assets during the Great Recession.
method Utilized 2007-2009 Survey of Consumer Finances (SCF) panel dataset, analyzed 3,862 respondents.
result Starting to use a financial planner during the Great Recession had a positive impact on preserving and increasing household net financial assets.
The paper analyzes the pricing of a new compute futures asset.
problem Uncertainty in AI adoption and pricing of compute capital.
method An asset-pricing framework for compute futures, including synthetic futures pricing.
result Preliminary evidence suggests a positive compute risk premium.
The paper improves machine learning for heavy-tailed panel data.
problem Improving estimates for financial and economic data with fat tails.
method Sparse-group LASSO regularization and Fuk-Nagaev concentration inequality.
result Oracle inequalities for panel data estimators.
Optimizes profit in targeted marketing across multiple markets with varying marketing expenditures.
problem Maximizing profit in a sequential marketing strategy with multiple markets and varying marketing costs.
method Near-optimal algorithms in an adversarial bandit setting, proving regret bounds for different demand curve types.
result Proved near-optimal regret bounds for the profit-maximization problem in targeted marketing.
New method tests Granger non-causality in panel data with cross-sectional dependencies.
problem Testing Granger non-causality in panel data with cross-sectional dependencies.
method Proposes a new approach to aggregate p-values from panel members to test Granger non-causality, showing lower FDR.
result Our approach discovers true causal relations in panel data, unlike state-of-the-art methods.
Improving the precision of heart diseases detection has been investigated by many researchers in the literature. Such improvement induced by the overwhelming health care expenditures and erroneous diagnosis. As a result, various methodologies have been proposed to analyze the disease factors aiming to decrease the phys…
In the present paper, we identify several distributions from Physics and study their applicability to phenomena such as distribution of income, wealth, and expenditure. Firstly, we apply logistic distribution to these data and we find that it fits very well the annual data for the entire income interval including for u…
Machine learning predicts exercise load from heart rate data post-exercise.
problem Monitoring energy expenditure in real life.
method Machine learning methods (linear regression, etc.) applied to heart rate data.
result Random forest and k-nearest neighbors classifiers predict load levels accurately.
Paper develops a new estimator for panel data with endogenous treatments, improving causal inference.
problem Challenges in causal inference for static panel data with endogenous treatments and confounding variables.
method Develops Double Machine Learning (DML) estimator for static panel models with endogenous treatments (panel IV DML). Introduces weak-identification diagnostics.
result Panel IV DML estimator improves estimation accuracy and delivers more reliable inference under weak identification.
New method for estimating heterogeneous treatment effects in panel data.
problem Estimating heterogeneous treatment effects in non-stationary, temporally dependent panel data.
method Proposes H1SL and H2SL, synthetic learners for panel data, based on existing non-panel data estimators.
result Established convergence rates for proposed estimators and demonstrated superior performance.
We propose a robust implementation of the Nerlove--Arrow model using a Bayesian structural time series model to explain the relationship between advertising expenditures of a country-wide fast-food franchise network with its weekly sales. Thanks to the flexibility and modularity of the model, it is well suited to gener…
What has happened in machine learning lately, and what does it mean for the future of medical image analysis? Machine learning has witnessed a tremendous amount of attention over the last few years. The current boom started around 2009 when so-called deep artificial neural networks began outperforming other established…
Polynomial distribution can be applied to dynamical systems in certain situations. Macroeconomic systems characterized by economic variables such as income and wealth can be modelled similarly using polynomials. We extend our previous work to data regarding income from a more diversified pool of countries, which contai…
This paper addresses privacy in federated learning for medical imaging by estimating model uncertainty.
problem Privacy concerns in federated learning for medical imaging.
method Federated Learning (FL) for collaborative model training while preserving patient data privacy.
result Accurate uncertainty estimation in federated learning for medical imaging.
Deep learning models accurately recognize and estimate physical activity types and energy expenditure from wrist accelerometer data.
problem Rigorous evaluation of wrist-worn accelerometers for assessing physical activity across the lifespan.
method Built deep learning networks to extract spatial and temporal representations from time-series data, recognizing physical activity types and estimating energy expenditure.
result Deep learning models achieved high performance: F1 scores of 0.82, 0.81, and 95 for sedentary, locomotor, and lifestyle activities, respectively; root mean square error of 1.1 for EE estimation.
Study assesses health plan risk measures for Solvency Capital Requirement.
problem Assessing risk measures for health plans to meet Solvency Capital Requirement.
method Three-part regression model with three GLMs for claim counts, episode allocation, and severity.
result Reduction in regression models compared to traditional methods.
We analyze expenditure patterns of discretionary funds by Brazilian congress members. This analysis is based on a large dataset containing over 7 million expenses made publicly available by the Brazilian government. This dataset has, up to now, remained widely untouched by machine learning methods. Our main contribut…
Motivated by applications in architecture and design, we present a novel method for increasing the developability of a B-spline surface. We use the property that the Gauss image of a developable surface is 1-dimensional and can be locally well approximated by circles. This is cast into an algorithm for thinning the Gau…
A new method for online prediction uncertainty quantification in non-exchangeable panel data.
problem Challenges in quantifying predictive uncertainty for non-exchangeable panel data.
method Online conformal prediction framework for non-exchangeable panel data, using similarity weights and adaptive miscoverage levels.
result Improves coverage on worst-covered target units through adaptive interval-width allocation.
Estimates mean and covariance for large, unbalanced stock returns panels.
problem Estimating mean and covariance in large, unbalanced panel data.
method Nonparametric, kernel-based joint estimator for conditional mean and covariance matrices.
result The idiosyncratic risk explains more than 75% of cross-sectional variance.
In this paper, the development of a probabilistic network for the diagnosis of acute cardiopulmonary diseases is presented. This paper is a draft version of the article published after peer review in 2018 (https://doi.org/10.1002/bimj.201600206). A panel of expert physicians collaborated to specify the qualitative part…
Proposes CoDEAL for estimating heterogeneous treatment effects in panel data models.
problem Estimating heterogeneous treatment effects in causal panel data models with covariate effects.
method Covariate-Adjusted Deep Causal Learning (CoDEAL) integrating neural networks and autoencoders.
result Establishes theoretical guarantees and demonstrates compelling performance in simulations and real data.
Deep neural network predicts diabetic readmission with high accuracy.
problem Predicting 30-day readmission for diabetic patients.
method Categorical embeddings and deep neural network.
result 95.2% accuracy and 97.4% AUROC on diabetic readmission data.
Paper uses machine learning for nowcasting corporate earnings from mixed-frequency data.
problem Predicting corporate earnings for a large cross-section of firms with different frequency data.
method Structured machine learning regressions with sparse-group LASSO regularization for panel data.
result Machine learning models outperform traditional methods in nowcasting corporate earnings.