Evaluating Uncertainty Quantification in Clinical Machine Learning: Calibration, Robustness, and Decision Utility under Distribution Shift
Authors
Department of Statistics, Florida State University, Tallahassee, FL (USA)
Department of Statistics, Florida State University, Tallahassee, FL (USA)
Department of Economics, Carleton University, Ottawa, Ontario (Canada)
Article Information
DOI: 10.51584/IJRIAS.2026.11070085
Subject Category: Statistics
Volume/Issue: 11/7 | Page No: 1241-1252
Publication Timeline
Submitted: 2026-07-28
Accepted: 2026-07-23
Published: 2026-08-04
Abstract
Machine learning models deployed in clinical decision support systems almost universally produce point predictions without any accompanying measure of uncertainty. In high-stakes healthcare settings, this is not merely a technical limitation: a miscalibrated prediction can directly influence patient management decisions with real consequences for safety and outcomes. Several uncertainty quantification (UQ) methods have been proposed to address this gap, including conformal prediction, Bayesian neural networks (BNNs), and Monte Carlo (MC) dropout; however, their comparative evaluation has predominantly been conducted under idealised conditions that do not reflect clinical deployment, where patient populations, treatment practices, and data recording procedures change over time. We present a rigorous empirical framework for comparing these three UQ approaches on two clinical prediction tasks, in-hospital mortality and 30-day readmission, using 74,829 ICU admissions from the MIMIC-IV database. Methods are assessed across three dimensions: calibration quality (ECE, ACE, Brier score), robustness under temporal distribution shift, and clinical decision utility via net benefit analysis. All experiments are replicated across five independent seeds, with comparisons made using Wilcoxon signed-rank tests with Holm-Bonferroni correction. Under standard evaluation conditions, all three methods achieve similar discriminative performance (AUROC 0.836-0.844 for mortality; 0.637-0.641 for readmission). Under temporal shift, BNN calibration degrades most sharply on the readmission task (ΔECE = 0.011 ± 0.002) compared with MC Dropout (ΔECE = 0.002 ± 0.003), while AUROC paradoxically improves for all methods, demonstrating that discriminative and calibration performance can decouple under distribution shift. Conformal prediction maintains near-nominal empirical coverage on the mortality task (0.886 ± 0.002) but shows notable violations on readmission, raising practical concerns about exchangeability assumptions in deployed systems. These findings support a more demanding evaluation standard for UQ in clinical machine learning, one that moves beyond static i.i.d. benchmarks toward temporally robust, decision-aware assessment.
Keywords
uncertainty quantification; conformal prediction; Bayesian neural networks; Monte Carlo dropout
Downloads
References
1. Obermeyer, Z., & Emanuel, E. J. (2016). Predicting the future: Big data, machine learning, and clinical medicine. New England Journal of Medicine, 375(13), 1216-1219. [Google Scholar] [Crossref]
2. Vovk, V., Gammerman, A., & Shafer, G. (2005). Algorithmic learning in a random world. Springer. [Google Scholar] [Crossref]
3. Blundell, C., Cornebise, J., Kavukcuoglu, K., & Wierstra, D. (2015). Weight uncertainty in neural networks. Proceedings of the 32nd International Conference on Machine Learning, 1613-1622. [Google Scholar] [Crossref]
4. Gal, Y., & Ghahramani, Z. (2016). Dropout as a Bayesian approximation: Representing model uncertainty in deep learning. Proceedings of the 33rd International Conference on Machine Learning, 1050-1059. [Google Scholar] [Crossref]
5. Finlayson, S. G., Subbaswamy, A., Singh, K., Bowers, J., Kupke, A., Zittrain, J., Kohane, I. S., & Saria, S. (2021). The clinician and dataset shift in artificial intelligence. New England Journal of Medicine, 385(3), 283-286. [Google Scholar] [Crossref]
6. Nestor, B., McDermott, M. B., Boag, W., Berner, G., Naumann, T., Hughes, M. C., Goldenberg, A., & Ghassemi, M. (2019). Feature robustness in non-stationary health records. Proceedings of Machine Learning for Health (ML4H). [Google Scholar] [Crossref]
7. Nixon, J., Dusenberry, M. W., Zhang, L., Jerfel, G., & Tran, D. (2019). Measuring calibration in deep learning. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. [Google Scholar] [Crossref]
8. Vaicenavicius, J., Widmann, D., Andersson, C., Lindsten, F., Roll, J., & Schön, T. (2019). Evaluating model calibration in classification. Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics, 3459-3467. [Google Scholar] [Crossref]
9. Hüllermeier, E., & Waegeman, W. (2021). Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods. Machine Learning, 110(3), 457-506. [Google Scholar] [Crossref]
10. Ovadia, Y., Fertig, E., Ren, J., Nado, Z., Sculley, D., Nowozin, S., Dillon, J., Lakshminarayanan, B., & Snoek, J. (2019). Can you trust your model's uncertainty? Evaluating predictive uncertainty under dataset shift. Advances in Neural Information Processing Systems, 32. [Google Scholar] [Crossref]
11. Papadopoulos, H., Proedrou, K., Vovk, V., & Gammerman, A. (2002). Inductive confidence machines for regression. Proceedings of the 13th European Conference on Machine Learning, 345-356. [Google Scholar] [Crossref]
12. Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1-3. [Google Scholar] [Crossref]
13. Kingma, D. P., Salimans, T., & Welling, M. (2015). Variational dropout and the local reparameterization trick. Advances in Neural Information Processing Systems, 28, 2575-2583. [Google Scholar] [Crossref]
14. Kingma, D. P., & Ba, J. (2014). Adam: A method for stochastic optimisation. arXiv preprint arXiv:1412.6980. [Google Scholar] [Crossref]
15. Johnson, A. E. W., Bulgarelli, L., Shen, L., Gayles, A., Shammout, A., Horng, S., Pollard, T. J., Hao, S., Moody, B., Gow, B., Lehman, L. H., Celi, L. A., & Mark, R. G. (2023). MIMIC-IV, a freely accessible electronic health record dataset. Scientific Data, 10(1), 1. [Google Scholar] [Crossref]
16. Goldberger, A. L., Amaral, L. A. N., Glass, L., Hausdorff, J. M., Ivanov, P. C., Mark, R. G., Mietus, J. E., Moody, G. B., Peng, C. K., & Stanley, H. E. (2000). PhysioBank, PhysioToolkit, and PhysioNet. Circulation, 101(23), e215-e220. [Google Scholar] [Crossref]
17. Lakshminarayanan, B., Pritzel, A., & Blundell, C. (2017). Simple and scalable predictive uncertainty estimation using deep ensembles. Advances in Neural Information Processing Systems, 30, 6402-6413. [Google Scholar] [Crossref]
18. Tibshirani, R. J., Barber, R. F., Candès, E., & Ramdas, A. (2019). Conformal prediction under covariate shift. Advances in Neural Information Processing Systems, 32. [Google Scholar] [Crossref]
19. Sensoy, M., Kaplan, L., & Kandemir, M. (2018). Evidential deep learning to quantify classification uncertainty. Advances in Neural Information Processing Systems, 31. [Google Scholar] [Crossref]
20. Adisa, I. T. (2026). An integrated framework for explainable, fair, and observable hospital readmission prediction: Development and validation on MIMIC-IV. International Journal of Research and Innovation in Applied Science, 11(6). https://doi.org/10.51584/IJRIAS.2026.11060154 [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- The Net Relative Run-Ratio Method (NRRR), a Foolproof Technique to Replace the Net Run Rate (NRR) Method in Evaluating the Authority of Match-Wins
- Statistical Role of CB-SEM Vs PLS-SEM in the Field of Social Science
- Predictive Modelling and Statistical Analysis of Housing Prices in Lagos State, Nigeria
- Collocational Patterns of Guru in American Business vs. Spiritual Discourse
- A Comparative Analysis of Heuristic and Dynamic Algorithms for Route Optimization in Johor’s Delivery Hubs