A Comparative Analysis of Machine Learning Algorithms for the Early Prediction of Diabetes with an Evaluation of Class-Imbalance Handling
Authors
Department of Computer Science, College of Computing and Information Sciences, Caleb University, KM 15, Ikorodu-Itoikin Road, Imota, P.M.B. 21238, Lagos State, Nigeria (Nigeria)
Department of Computer Science, College of Computing and Information Sciences, Caleb University, KM 15, Ikorodu-Itoikin Road, Imota, P.M.B. 21238, Lagos State, Nigeria (Nigeria)
Department of Computer Science, College of Computing and Information Sciences, Caleb University, KM 15, Ikorodu-Itoikin Road, Imota, P.M.B. 21238, Lagos State, Nigeria (Nigeria)
Department of Computer Science, College of Computing and Information Sciences, Caleb University, KM 15, Ikorodu-Itoikin Road, Imota, P.M.B. 21238, Lagos State, Nigeria (Nigeria)
Department of Computer Science, College of Computing and Information Sciences, Caleb University, KM 15, Ikorodu-Itoikin Road, Imota, P.M.B. 21238, Lagos State, Nigeria (Nigeria)
Department of Computer Science, College of Computing and Information Sciences, Caleb University, KM 15, Ikorodu-Itoikin Road, Imota, P.M.B. 21238, Lagos State, Nigeria (Nigeria)
Department of Computer Science, College of Computing and Information Sciences, Caleb University, KM 15, Ikorodu-Itoikin Road, Imota, P.M.B. 21238, Lagos State, Nigeria (Nigeria)
Article Information
Publication Timeline
Submitted: 2026-08-19
Accepted: 2026-08-21
Published: 2026-08-26
Abstract
Although early detection of diabetes significantly lowers its harm, machine learning models for screening are often assessed based on overall accuracy, a metric that is deceptive given the high-class imbalance that characterizes clinical data. In this work, we compared and evaluated the performance of five algorithms: logistic regression, naive Bayes, support vector machine (SVM), decision tree, and extreme gradient boosting (XGBoost) to early predict diabetes in many patients. Diabetes affected about 13.9% of the 253,680 records that were looked at in the 2015 Behavioral Risk Factor Surveillance Systems Diabetes Health Indicators dataset and to deal with this problem, people followed a process called CRISP-DM and they also used something called the Synthetic Minority Oversampling Technique or SMOTE for short. They used these things to retrain each model with the data, which was not balanced. Each model was then looked at using a few different measures, including the F1-score, accuracy, precision, recall and the area, under the ROC curve to see how well each Diabetes model was working. The Diabetes models were evaluated to see how well they were doing. Ten-fold cross-validation and a held-out test set were both used to validate the results.In the absence of imbalance handling, most models diagnosed less than one in five cases of diabetes with an accuracy of roughly 86%. While the decision tree improved somewhat and XGBoost was essentially unaffected, applying SMOTE increased the recall of the linear and probabilistic models from below 0.17 to above 0.76 at the expense of accuracy and precision. The best predictors were found to be high blood pressure, overall health, and high cholesterol. The study comes to the conclusion that headline accuracy is an unreliable guide in imbalanced medical prediction, that imbalance handling can change a model's practical usefulness, and that this benefit is strongly algorithm-dependent, meaning that the decision to resample should be based on the algorithm and the screening priorities rather than being applied consistently.
Keywords
Diabetes, detection, Synthetic Minority Oversampling Technique (SMOTE), machine learning, comparative analysis, XGBoost, Logistic regression, support vector machine (SVM)
Downloads
References
1. Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16, 321–357. [Google Scholar] [Crossref]
2. https://doi.org/10.1613/jair.953 [Google Scholar] [Crossref]
3. Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). Association for Computing Machinery. https://doi.org/10.1145/2939672.2939785 [Google Scholar] [Crossref]
4. Chou, C.-Y., Hsu, D.-Y., & Chou, C.-H. (2023). Predicting the onset of diabetes with machine learning methods. Journal of Personalized Medicine, 13(3), Article 406. https://doi.org/10.3390/jpm13030406 [Google Scholar] [Crossref]
5. Kavakiotis, I., Tsave, O., Salifoglou, A., Maglaveras, N., Vlahavas, I., & Chouvarda, I. (2017). Machine learning and data mining methods in diabetes research. Computational and Structural Biotechnology Journal, 15, 104–116. https://doi.org/10.1016/j.csbj.2016.12.005 [Google Scholar] [Crossref]
6. Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, E. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830. [Google Scholar] [Crossref]
7. Salmi, M., Atif, D., Oliva, D., Abraham, A., & Ventura, S. (2024). Handling imbalanced medical datasets: Review of a decade of research. Artificial Intelligence Review, 57(10), Article 273. [Google Scholar] [Crossref]
8. https://doi.org/10.1007/s10462-024-10884-2 [Google Scholar] [Crossref]
9. Sun, H., Saeedi, P., Karuranga, S., Pinkepank, M., Ogurtsova, K., Duncan, B. B., Stein, C., Basit, A., Chan, J. C. N., Mbanya, J. C., Pavkov, M. E., Ramachandaran, A., Wild, S. H., James, S., Herman, W. H., Zhang, P., Bommer, C., Kuo, S., Boyko, E. J., & Magliano, D. J. (2022). IDF Diabetes Atlas: Global, regional and country-level diabetes prevalence estimates for 2021 and projections for 2045. Diabetes Research and Clinical Practice, 183, Article 109119. [Google Scholar] [Crossref]
10. https://doi.org/10.1016/j.diabres.2021.109119 [Google Scholar] [Crossref]
11. Wang, S., Chen, R., Wang, S., Kong, D., Cao, R., Lin, C., Feng, W., Wang, Y., & Xu, H. (2023). Comparative study on risk prediction model of type 2 diabetes based on machine learning theory: A cross-sectional study. BMJ Open, 13(8), Article e069018. https://doi.org/10.1136/bmjopen-2022-069018 [Google Scholar] [Crossref]
12. Yadu, S., Chandra, R., & Sinha, V. K. (2024). Comparing different machine learning techniques in predicting diabetes on early stage. Engineering Proceedings, 62(1), Article 20. [Google Scholar] [Crossref]
13. https://doi.org/10.3390/engproc2024062020 [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- Assessment of the Role of Artificial Intelligence in Repositioning TVET for Economic Development in Nigeria
- Teachers’ Use of Assure Model Instructional Design on Learners’ Problem Solving Efficacy in Secondary Schools in Bungoma County, Kenya
- “E-Booksan Ang Kaalaman”: Development, Validation, and Utilization of Electronic Book in Academic Performance of Grade 9 Students in Social Studies
- Analyzing EFL University Students’ Academic Speaking Skills Through Self-Recorded Video Presentation
- Major Findings of The Study on Total Quality Management in Teachers’ Education Institutions (TEIs) In Assam – An Evaluative Study