An Explainable Machine Learning Framework for Predicting Healthcare Utilization and Quantifying Economic Burden in US Health Systems

Authors

Samiha Binte Abdullah

Master of Science in Business Analytics Ambassador Crawford College of Business and Entrepreneurship Kent State University, Kent, Ohio, USA (USA)

Anurodh Singh

Master of Science in Business Analytics Ambassador Crawford College of Business and Entrepreneurship Kent State University, Kent, Ohio, USA (USA)

Gaurav Kudeshia

Master of Science in Business Analytics Ambassador Crawford College of Business and Entrepreneurship Kent State University, Kent, Ohio, USA (USA)

Article Information

DOI: 10.51244/IJRSI.2026.1306000345

Subject Category: Education

Volume/Issue: 13/6 | Page No: 4643-4655

Publication Timeline

Submitted: 2026-06-17

Accepted: 2026-06-22

Published: 2026-07-10

Abstract

Background: Rising healthcare expenditure in the United States represents one of the most critical challenges facing health systems today. Accurate prediction of healthcare utilisation patterns and associated economic burden is essential for equitable resource allocation, insurance planning, and evidence-based health policy.
Objective: This study presents an explainable machine learning framework integrating a two-stage economic modelling architecture with Shapley Additive Explanations (SHAP) to simultaneously predict healthcare utilisation across three care settings and quantify individual and system-level economic burden.
Methods: We utilise the Medical Expenditure Panel Survey (MEPS) as the primary dataset, supplemented by HCUP for external validation. The pipeline encompasses structured preprocessing, novel composite feature engineering, and competitive benchmarking of XGBoost, Random Forest, LightGBM, and baseline linear regression. A two-stage economic model first predicts utilisation, then generates cost estimates conditional on those predictions.
Results: XGBoost achieved superior performance (RMSE = 1.24 ED visits, R2 = 0.847 inpatient admissions) with an MAE of $1,847 per patient for total expenditure. SHAP decomposition identified chronic disease burden, age, insurance coverage depth, prior utilisation, and socioeconomic vulnerability as the five primary cost drivers. System-level forecasts matched CMS figures within 3.2%.
Conclusions: The framework advances the state of the art by unifying utilisation prediction, econometric modelling, and explainable AI in a single reproducible pipeline. Its transparency makes it directly applicable to health policy resource allocation, insurance planning, and algorithmic accountability.

Keywords

Machine learning; Healthcare utilisation; Economic burden; SHAP explainability; XGBoost; Medical expenditure; Predictive modelling; Health equity; United States health systems

Downloads

References

1. Agency for Healthcare Research and Quality. (2023). Medical Expenditure Panel Survey (MEPS) HC-224: 2021 Full Year Consolidated Data File. US Department of Health and Human Services. [Google Scholar] [Crossref]

2. Centers for Medicare & Medicaid Services. (2023). National Health Expenditure Data. US Department of Health and Human Services. [Google Scholar] [Crossref]

3. Chen, Y., Zhang, L., & Wang, H. (2022). SHAP-based interpretable machine learning for insurance premium prediction. Insurance: Mathematics and Economics, 103, 45–62. [Google Scholar] [Crossref]

4. Dieleman, J. L., Squires, E., Bui, A. L., Campbell, M., Chapin, A., Hamavid, H., & Murray, C. J. L. (2017). Factors associated with increases in US health care spending, 1996–2013. JAMA, 318(17), 1668–1678. [Google Scholar] [Crossref]

5. Futoma, J., Morris, J., & Lucas, J. (2015). A comparison of models for predicting early hospital readmissions. Journal of Biomedical Informatics, 56, 229–238. [Google Scholar] [Crossref]

6. Goto, T., Hirayama, A., Faridi, M. K., Camargo, C. A., & Hasegawa, K. (2019). Machine learning for prediction of emergency department return visits. Academic Emergency Medicine, 26(5), 546–555. [Google Scholar] [Crossref]

7. Lundberg, S. M., & Lee, S. I. (2017). A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems, 30, 4765–4774. [Google Scholar] [Crossref]

8. Lundberg, S. M., Erion, G., Chen, H., DeGrave, A., Prutkin, J. M., Nair, B., & Lee, S. I. (2020). From local explanations to global understanding with explainable AI for trees. Nature Machine Intelligence, 2(1), 56–67. [Google Scholar] [Crossref]

9. Manning, W. G., & Mullahy, J. (2001). Estimating log models: To transform or not to transform? Journal of Health Economics, 20(4), 461–494. [Google Scholar] [Crossref]

10. Obermeyer, Z., & Emanuel, E. J. (2016). Predicting the future Big data, machine learning, and clinical medicine. New England Journal of Medicine, 375(13), 1216–1219. [Google Scholar] [Crossref]

11. Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453. [Google Scholar] [Crossref]

12. Rajpurkar, P., Irvin, J., Ball, R. L., Zhu, K., Yang, B., Mehta, H., & Ng, A. Y. (2017). CheXNet: Radiologist-level pneumonia detection on chest X-rays with deep learning. arXiv preprint arXiv:1711.05225. [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles