A Systematic Analysis of Performance Evaluation Metrics in Machine Learning Models

Authors

Muhammad Tella

Department of Management and Information Technology, Abubakar Tafawa Balewa University, Bauchi (Nigeria)

Mahmud Ahmed Usman

Department of Management and Information Technology, Abubakar Tafawa Balewa University, Bauchi (Nigeria)

Kabiru Ibrahim Musa

Department of Management and Information Technology, Abubakar Tafawa Balewa University, Bauchi (Nigeria)

Article Information

DOI: 10.51584/IJRIAS.2026.11010070

Subject Category: Machine Learning

Volume/Issue: 11/1 | Page No: 834-840

Publication Timeline

Submitted: 2026-01-17

Accepted: 2026-01-25

Published: 2026-02-06

Abstract

Machine Learning (ML) has been a critical computational paradigm that has shaped contemporary applications in such domains as finance, healthcare, and cybersecurity, such that its performance evaluation cannot be less critical. However, its selection and interpretation of metrics has remained inconsistent, often leading to misleading conclusions. This study presents a systematic analysis of the most commonly used performance evaluation metrics in ML, integrating conceptual taxonomy, mathematical definitions, and empirical assessment under controlled perturbations. There are three dimensions to ML performance evaluation metrics categorization: robustness, discrimination, and calibration. Experiment conducted on classification and regression, and using synthetic datasets and benchmarks, evaluate threshold variation, class imbalance and label noise. Results obtained showed that no single metric captures model performance comprehensively and widely used metrics may yield conflicting or misleading assessments under certain conditions. Also, context-aware selection and multi-dimensional reporting were necessary for reliable evaluation. By empirically linking metric behaviour to data characteristics, this study provides guidance for context-aware metric selection and reporting that is not only standardized but also evidence-based.

Keywords

Machine Learning, Robustness, Calibration, Evaluation Framework, Regression

Downloads

References

1. T. M. Mitchell, "Does machine learning really work?," AI Magazine, vol. 18, no. 3, pp. 12-20, 1997. [Google Scholar] [Crossref]

2. J. L. Crawley, Pattern Recognition and Machine Learning, New York: Springer, 2006. [Google Scholar] [Crossref]

3. D. Chicco and G. Jurman, "The advantages of the Matthews correlation coefficient (MCC) over F1 score andaccuracy in binary classification evaluation," BMC Genomics, vol. 21, no. 1, pp. 1-13, 2020. [Google Scholar] [Crossref]

4. C. Guo, G. Pleiss, Y. Sun and K. Q. Weinberger, "On Calibration of Modern Neural Networks," in International Conference on Machine Learning, 2017. [Google Scholar] [Crossref]

5. M. Grandini, E. Bagli and G. Visani, "Metrics for Multi-Class Classification: An Overview," arXiv preprint, 2020. [Google Scholar] [Crossref]

6. Y. Ovadia, E. Fertig, J. Ren, Z. Nado, D. Sculley, S. Nowozin, J. V. Dillon, B. Lakshminarayanan and J. Snoek, "Can You Trust Your Model’s Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift," Advances in neural information processing systems, 2019. [Google Scholar] [Crossref]

7. T. Faucett, "An introduction to ROC analysis," Pattern recognition letters, vol. 27, no. 8, pp. 861-874, 2006. [Google Scholar] [Crossref]

8. T. Saito and M. Rehmsmeier, "The Precision-Recall Plot Is More Informative than the ROCPlot WhenEvaluating Binary Classifiers on Imbalanced Datasets," PloS one, vol. 10, no. 3, 2015. [Google Scholar] [Crossref]

9. M. Minderer, J. Djolonga, R. Romijnders, F. Hubis, X. Zhai, N. Houlsby, D. Tran and M. Lucic, "Revisiting the Calibration of Modern Neural Networks," Advances in neural information processing systems, vol. 34, pp. 15682-15694, 2021. [Google Scholar] [Crossref]

10. M. Xiong, A. Deng, P. W. Koh, J. Wu, S. Li, J. Xu and B. Hooi, "Proximity-informed calibration for deep neural networks," Advances in Neural Information Processing Systems , vol. 36 , pp. 68511-68538, 2023. [Google Scholar] [Crossref]

11. M. Schulze, N. Ebert, L. Reichardt and O. Wasenm, "Classifier Ensemble for Efficient Uncertainty Calibration of Deep Neural Networks for Image Classification," arXiv preprint, 2025. [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles