An Explainable SMOTE-Enhanced Fuzzy Ensemble Decision Tree Framework for Intelligent Cancer Prediction and Classification with Hyperparameter Turning

Authors

Sylvester I. Ele

Department of Software Engineering, University of Calabar, Nigeria (Nigeria)

Oko Sunday Adi

Department of Computer Science, University of Calabar, Nigeria (Nigeria)

Article Information

DOI: 10.51584/IJRIAS.2026.11060252

Subject Category: Computer Science

Volume/Issue: 11/6 | Page No: 3278-3288

Publication Timeline

Submitted: 2026-06-27

Accepted: 2026-07-02

Published: 2026-07-15

Abstract

Early detection of cervical cancer remains a critical challenge in healthcare, particularly in low- and middle-income countries where limited access to screening and diagnostic facilities contributes to high mortality rates. Although machine learning techniques have shown considerable promise in supporting cancer diagnosis, their effectiveness is often hindered by class imbalance, data uncertainty, and limited model interpretability. This study proposes an explainable SMOTE-enhanced fuzzy ensemble decision tree framework with hyperparameter optimization for intelligent cervical cancer prediction and classification. The framework integrates Synthetic Minority Oversampling Technique (SMOTE), fuzzy logic, multiple fuzzy decision trees (Fuzzy ID3, Fuzzy C4.5, and Fuzzy CART), and ensemble learning models, including Random Forest, Gradient Boosting, XGBoost, LightGBM, and AdaBoost, to improve predictive accuracy, robustness, and clinical interpretability. The Cervical Cancer (Risk Factors) dataset obtained from the UCI Machine Learning Repository, comprising 858 patient records and 36 clinical and demographic attributes, was employed for experimentation. Data preprocessing involved missing-value imputation, normalization, fuzzification of selected clinical variables using triangular membership functions, and class balancing through SMOTE. Hyperparameter tuning was performed using Bayesian optimization, grid search, and randomized search techniques. Model performance was evaluated using accuracy, precision, recall, F1-score, ROC-AUC, confusion matrices, and ROC curves. Experimental results demonstrated that ensemble methods consistently outperformed conventional decision tree algorithms. The proposed framework achieved its best performance with Gradient Boosting, attaining an accuracy of 97.09%, precision of 87.50%, recall of 63.64%, F1-score of 73.68%, and a ROC-AUC of 94.35%, while XGBoost and Random Forest achieved the highest discriminative capabilities with ROC-AUC values of 97.35% and 97.21%, respectively. The incorporation of SMOTE significantly improved minority-class detection, whereas fuzzy logic enhanced the handling of uncertainty inherent in medical data. The findings confirm that the proposed explainable framework provides an effective, robust, and interpretable decision-support mechanism for early cervical cancer diagnosis and has strong potential for broader applications in intelligent healthcare systems.

Keywords

Cervical cancer prediction, explainable artificial intelligence, SMOTE, fuzzy decision trees, hyperparameter optimization.

Downloads

References

1. Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P.(2002). SMOTE: Synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16, 321–357. [Google Scholar] [Crossref]

2. Breiman, L. (2001). Random Forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/A:1010933404324. [Google Scholar] [Crossref]

3. Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 785–794. [Google Scholar] [Crossref]

4. Fernández, A., Garcia, S., Galar, M., Prati, R. C., Krawczyk, B., & Herrera, F. (2018). Learning from imbalanced data sets. Springer. https://link.springer.com/book/10.1007/978-3-319-98074-4. [Google Scholar] [Crossref]

5. Friedman, J. H. (2001). Greedy Function Approximation: A Gradient Boosting Machine. Annals of Statistics, 29(5), 1189–1232. [Google Scholar] [Crossref]

6. Haixiang, G., Yijing, L., Shang, J., Mingyun, G., Yuanyue, H., & Bing, G. (2017). Learning from Class-Imbalanced Data: Review of Methods and Applications. Expert Systems with Applications, 73, 220–239. [Google Scholar] [Crossref]

7. Janikow, C. Z. (1998). Fuzzy decision trees: Issues and methods. IEEE Transactions on Systems, Man, and Cybernetics, 28(1), 1–14. [Google Scholar] [Crossref]

8. Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., & Liu, T. Y. (2017). LightGBM: A highly efficient gradient boosting decision tree. Advances in Neural Information Processing Systems, 30. [Google Scholar] [Crossref]

9. Kourou, K., Exarchos, T. P., Exarchos, K. P., Karamouzis, M. V., & Fotiadis, D. I. (2015). Machine learning applications in cancer prognosis and prediction. Computational and Structural Biotechnology Journal, 13, 8–17. [Google Scholar] [Crossref]

10. Quinlan, J. R. (1993). C4.5: Programs for machine learning. Morgan Kaufmann. [Google Scholar] [Crossref]

11. World Health Organization. (2023). Cervical cancer. https://www.who.int/news-room/fact-sheets/detail/cervical-cancer [Google Scholar] [Crossref]

12. Zadeh, L. A. (1965). Fuzzy sets. Information and Control, 8(3), 338–353. [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles