An Explainable Machine Learning Framework for Predicting Healthcare Utilization and Quantifying Economic Burden in US Health Systems
Authors
Master of Science in Business Analytics Ambassador Crawford College of Business and Entrepreneurship Kent State University, Kent, Ohio, USA (USA)
Master of Science in Business Analytics Ambassador Crawford College of Business and Entrepreneurship Kent State University, Kent, Ohio, USA (USA)
Master of Science in Business Analytics Ambassador Crawford College of Business and Entrepreneurship Kent State University, Kent, Ohio, USA (USA)
Article Information
DOI: 10.51244/IJRSI.2026.1306000345
Subject Category: Education
Volume/Issue: 13/6 | Page No: 4643-4655
Publication Timeline
Submitted: 2026-06-17
Accepted: 2026-06-22
Published: 2026-07-10
Abstract
Background: Rising healthcare expenditure in the United States represents one of the most critical challenges facing health systems today. Accurate prediction of healthcare utilisation patterns and associated economic burden is essential for equitable resource allocation, insurance planning, and evidence-based health policy.
Objective: This study presents an explainable machine learning framework integrating a two-stage economic modelling architecture with Shapley Additive Explanations (SHAP) to simultaneously predict healthcare utilisation across three care settings and quantify individual and system-level economic burden.
Methods: We utilise the Medical Expenditure Panel Survey (MEPS) as the primary dataset, supplemented by HCUP for external validation. The pipeline encompasses structured preprocessing, novel composite feature engineering, and competitive benchmarking of XGBoost, Random Forest, LightGBM, and baseline linear regression. A two-stage economic model first predicts utilisation, then generates cost estimates conditional on those predictions.
Results: XGBoost achieved superior performance (RMSE = 1.24 ED visits, R2 = 0.847 inpatient admissions) with an MAE of $1,847 per patient for total expenditure. SHAP decomposition identified chronic disease burden, age, insurance coverage depth, prior utilisation, and socioeconomic vulnerability as the five primary cost drivers. System-level forecasts matched CMS figures within 3.2%.
Conclusions: The framework advances the state of the art by unifying utilisation prediction, econometric modelling, and explainable AI in a single reproducible pipeline. Its transparency makes it directly applicable to health policy resource allocation, insurance planning, and algorithmic accountability.
Keywords
Machine learning; Healthcare utilisation; Economic burden; SHAP explainability; XGBoost; Medical expenditure; Predictive modelling; Health equity; United States health systems
Downloads
References
1. Agency for Healthcare Research and Quality. (2023). Medical Expenditure Panel Survey (MEPS) HC-224: 2021 Full Year Consolidated Data File. US Department of Health and Human Services. [Google Scholar] [Crossref]
2. Centers for Medicare & Medicaid Services. (2023). National Health Expenditure Data. US Department of Health and Human Services. [Google Scholar] [Crossref]
3. Chen, Y., Zhang, L., & Wang, H. (2022). SHAP-based interpretable machine learning for insurance premium prediction. Insurance: Mathematics and Economics, 103, 45–62. [Google Scholar] [Crossref]
4. Dieleman, J. L., Squires, E., Bui, A. L., Campbell, M., Chapin, A., Hamavid, H., & Murray, C. J. L. (2017). Factors associated with increases in US health care spending, 1996–2013. JAMA, 318(17), 1668–1678. [Google Scholar] [Crossref]
5. Futoma, J., Morris, J., & Lucas, J. (2015). A comparison of models for predicting early hospital readmissions. Journal of Biomedical Informatics, 56, 229–238. [Google Scholar] [Crossref]
6. Goto, T., Hirayama, A., Faridi, M. K., Camargo, C. A., & Hasegawa, K. (2019). Machine learning for prediction of emergency department return visits. Academic Emergency Medicine, 26(5), 546–555. [Google Scholar] [Crossref]
7. Lundberg, S. M., & Lee, S. I. (2017). A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems, 30, 4765–4774. [Google Scholar] [Crossref]
8. Lundberg, S. M., Erion, G., Chen, H., DeGrave, A., Prutkin, J. M., Nair, B., & Lee, S. I. (2020). From local explanations to global understanding with explainable AI for trees. Nature Machine Intelligence, 2(1), 56–67. [Google Scholar] [Crossref]
9. Manning, W. G., & Mullahy, J. (2001). Estimating log models: To transform or not to transform? Journal of Health Economics, 20(4), 461–494. [Google Scholar] [Crossref]
10. Obermeyer, Z., & Emanuel, E. J. (2016). Predicting the future Big data, machine learning, and clinical medicine. New England Journal of Medicine, 375(13), 1216–1219. [Google Scholar] [Crossref]
11. Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453. [Google Scholar] [Crossref]
12. Rajpurkar, P., Irvin, J., Ball, R. L., Zhu, K., Yang, B., Mehta, H., & Ng, A. Y. (2017). CheXNet: Radiologist-level pneumonia detection on chest X-rays with deep learning. arXiv preprint arXiv:1711.05225. [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- Assessment of the Role of Artificial Intelligence in Repositioning TVET for Economic Development in Nigeria
- Teachers’ Use of Assure Model Instructional Design on Learners’ Problem Solving Efficacy in Secondary Schools in Bungoma County, Kenya
- “E-Booksan Ang Kaalaman”: Development, Validation, and Utilization of Electronic Book in Academic Performance of Grade 9 Students in Social Studies
- Analyzing EFL University Students’ Academic Speaking Skills Through Self-Recorded Video Presentation
- Major Findings of The Study on Total Quality Management in Teachers’ Education Institutions (TEIs) In Assam – An Evaluative Study