A Systematic Analysis of Performance Evaluation Metrics in Machine Learning Models
Authors
Department of Management and Information Technology, Abubakar Tafawa Balewa University, Bauchi (Nigeria)
Department of Management and Information Technology, Abubakar Tafawa Balewa University, Bauchi (Nigeria)
Department of Management and Information Technology, Abubakar Tafawa Balewa University, Bauchi (Nigeria)
Article Information
DOI: 10.51584/IJRIAS.2026.11010070
Subject Category: Machine Learning
Volume/Issue: 11/1 | Page No: 834-840
Publication Timeline
Submitted: 2026-01-17
Accepted: 2026-01-25
Published: 2026-02-06
Abstract
Machine Learning (ML) has been a critical computational paradigm that has shaped contemporary applications in such domains as finance, healthcare, and cybersecurity, such that its performance evaluation cannot be less critical. However, its selection and interpretation of metrics has remained inconsistent, often leading to misleading conclusions. This study presents a systematic analysis of the most commonly used performance evaluation metrics in ML, integrating conceptual taxonomy, mathematical definitions, and empirical assessment under controlled perturbations. There are three dimensions to ML performance evaluation metrics categorization: robustness, discrimination, and calibration. Experiment conducted on classification and regression, and using synthetic datasets and benchmarks, evaluate threshold variation, class imbalance and label noise. Results obtained showed that no single metric captures model performance comprehensively and widely used metrics may yield conflicting or misleading assessments under certain conditions. Also, context-aware selection and multi-dimensional reporting were necessary for reliable evaluation. By empirically linking metric behaviour to data characteristics, this study provides guidance for context-aware metric selection and reporting that is not only standardized but also evidence-based.
Keywords
Machine Learning, Robustness, Calibration, Evaluation Framework, Regression
Downloads
References
1. T. M. Mitchell, "Does machine learning really work?," AI Magazine, vol. 18, no. 3, pp. 12-20, 1997. [Google Scholar] [Crossref]
2. J. L. Crawley, Pattern Recognition and Machine Learning, New York: Springer, 2006. [Google Scholar] [Crossref]
3. D. Chicco and G. Jurman, "The advantages of the Matthews correlation coefficient (MCC) over F1 score andaccuracy in binary classification evaluation," BMC Genomics, vol. 21, no. 1, pp. 1-13, 2020. [Google Scholar] [Crossref]
4. C. Guo, G. Pleiss, Y. Sun and K. Q. Weinberger, "On Calibration of Modern Neural Networks," in International Conference on Machine Learning, 2017. [Google Scholar] [Crossref]
5. M. Grandini, E. Bagli and G. Visani, "Metrics for Multi-Class Classification: An Overview," arXiv preprint, 2020. [Google Scholar] [Crossref]
6. Y. Ovadia, E. Fertig, J. Ren, Z. Nado, D. Sculley, S. Nowozin, J. V. Dillon, B. Lakshminarayanan and J. Snoek, "Can You Trust Your Model’s Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift," Advances in neural information processing systems, 2019. [Google Scholar] [Crossref]
7. T. Faucett, "An introduction to ROC analysis," Pattern recognition letters, vol. 27, no. 8, pp. 861-874, 2006. [Google Scholar] [Crossref]
8. T. Saito and M. Rehmsmeier, "The Precision-Recall Plot Is More Informative than the ROCPlot WhenEvaluating Binary Classifiers on Imbalanced Datasets," PloS one, vol. 10, no. 3, 2015. [Google Scholar] [Crossref]
9. M. Minderer, J. Djolonga, R. Romijnders, F. Hubis, X. Zhai, N. Houlsby, D. Tran and M. Lucic, "Revisiting the Calibration of Modern Neural Networks," Advances in neural information processing systems, vol. 34, pp. 15682-15694, 2021. [Google Scholar] [Crossref]
10. M. Xiong, A. Deng, P. W. Koh, J. Wu, S. Li, J. Xu and B. Hooi, "Proximity-informed calibration for deep neural networks," Advances in Neural Information Processing Systems , vol. 36 , pp. 68511-68538, 2023. [Google Scholar] [Crossref]
11. M. Schulze, N. Ebert, L. Reichardt and O. Wasenm, "Classifier Ensemble for Efficient Uncertainty Calibration of Deep Neural Networks for Image Classification," arXiv preprint, 2025. [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- A Machine Learning Model for Predicting the Risk of Developing Diabetes - T2DM Using Real-World Data from Kilifi, Kenya
- AI-Powered Facial Recognition Attendance System Using Deep Learning and Computer Vision
- A Comprehensive Review on Brain Tumour Segmentation Using Deep Learning Approach
- A Scalable Retrieval-Augmented Generation Pipeline for Domain-Specific Knowledge Applications
- Predictive Maintenance in Semiconductor Manufacturing Using Machine Learning on Imbalanced Dataset