Adversarial Attacks on Machine Learning–Based Cybersecurity Classifiers: A Systematic Analysis and Defense Framework
Authors
Department of Cybersecurity Evangel University Akaeze, Ebonyi State (Nigeria)
Caritas University Enugu Amorji Nike, Enugu State (Nigeria)
Article Information
DOI: 10.51244/IJRSI.2026.1307000430
Subject Category: Computer Science
Volume/Issue: 13/7 | Page No: 5867-5898
Publication Timeline
Submitted: 2026-08-08
Accepted: 2026-08-13
Published: 2026-08-24
Abstract
Machine learning and deep learning have become essential components of modern cybersecurity because of their ability to detect malicious activities, classify network traffic, identify malware, recognize phishing attempts, and support automated incident response. However, machine learning–based cybersecurity classifiers are vulnerable to adversarial attacks in which attackers deliberately manipulate data, features, model inputs, or training processes to cause misclassification or evade detection. This study systematically analyzes adversarial attacks against machine learning–based cybersecurity classifiers and proposes a comprehensive defense framework to improve their robustness and reliability. The study examines major attack categories, including evasion, data poisoning, model extraction, inference, backdoor, and adversarial example attacks. It also analyzes attack surfaces, threat models, and the consequences of adversarial manipulation in intrusion detection, malware detection, phishing classification, and other AI-enabled cybersecurity systems. The findings indicate that adversarial attacks can significantly reduce detection performance, increase false-negative and false-positive rates, manipulate decision boundaries, and undermine trust in automated security systems. A defense-in-depth framework is proposed, incorporating secure data management, adversarial training, robust feature engineering, model validation, ensemble learning, anomaly detection, explainable AI, continuous monitoring, human oversight, and regular security auditing. The study concludes that no single defense mechanism can provide complete protection against adversarial machine learning attacks. Therefore, resilient AI-based cybersecurity requires a layered approach that protects data, features, models, inference processes, and the entire machine learning lifecycle.
Keywords
Adversarial Machine Learning, Cybersecurity Classifiers, Machine Learning Security, Evasion Attacks, Data Poisoning, Intrusion Detection, Adversarial Examples, Robust AI, Deep Learning, Cyber Defense
Downloads
References
1. Zhao, P., Zhu, W., Jiao, P., Gao, D., & Wu, O. (2025). Data poisoning in deep learning: A survey. arXiv preprint arXiv:2503.22759. DOI: 10.48550/arXiv.2503.22759. [Google Scholar] [Crossref]
2. Kumar, R. S. S., Nyström, M., Lambert, J., Marshall, A., Goertzel, M., Comissoneru, A., Swann, M., & Xia, S. (2020). Adversarial machine learning—Industry perspectives. In 2020 IEEE Security and Privacy Workshops (SPW) (pp. 69–75). IEEE. DOI: 10.1109/SPW50608.2020.00028. [Google Scholar] [Crossref]
3. Alotaibi, A., & Rassam, M. A. (2023). Adversarial machine learning attacks against intrusion detection systems: A survey on strategies and defense. Future Internet, 15(2), 62. DOI: 10.3390/fi15020062. [Google Scholar] [Crossref]
4. Yan, S., Ren, J., Wang, W., Sun, L., Zhang, W., & Yu, Q. (2023). A survey of adversarial attack and defense methods for malware classification in cybersecurity. IEEE Communications Surveys & Tutorials, 25(1), 467–496. DOI: 10.1109/COMST.2022.3225137. [Google Scholar] [Crossref]
5. Saini, S., Chennamaneni, A., & Sawyerr, B. (2024). A review of the duality of adversarial learning in network intrusion: Attacks and countermeasures. arXiv preprint arXiv:2412.13880. DOI: 10.48550/arXiv.2412.13880. [Google Scholar] [Crossref]
6. Zhou, S., Liu, C., Ye, D., Zhu, T., Zhou, W., & Yu, P. S. (2022). Adversarial attacks and defenses in deep learning: From a perspective of cybersecurity. ACM Computing Surveys, 55(8), 1–39. DOI: 10.1145/3547330. [Google Scholar] [Crossref]
7. Khaleel, Y. L., Habeeb, M. A., Albahri, A. S., Al-Quraishi, T., Albahri, O. S., & Alamoodi, A. H. (2024). Network and cybersecurity applications of defense in adversarial attacks: A state-of-the-art using machine learning and deep learning methods. Journal of Intelligent Systems, 33(1), 20240153. DOI: 10.1515/jisys-2024-0153. [Google Scholar] [Crossref]
8. Wang, Y., Sun, T., Li, S., Yuan, X., Ni, W., Hossain, E., & Poor, H. V. (2023). Adversarial attacks and defenses in machine learning-empowered communication systems and networks: A contemporary survey. IEEE Communications Surveys & Tutorials, 25(4), 2245–2298. DOI: 10.1109/COMST.2023.3288437. [Google Scholar] [Crossref]
9. Ibitoye, O., Shafiq, O., & Matrawy, A. (2019). Analyzing adversarial attacks against deep learning for intrusion detection in IoT networks. In Proceedings of the IEEE Global Communications Conference (GLOBECOM). DOI: 10.1109/GLOBECOM38437.2019.9014337. [Google Scholar] [Crossref]
10. Albahri, A. S., et al. (2024). Fuzzy decision-making framework for explainable golden multi-machine learning models for real-time adversarial attack detection in vehicular ad hoc networks. Information Fusion, 105, 102208. DOI: 10.1016/j.inffus.2023.102208. [Google Scholar] [Crossref]
11. Biggio, B., & Roli, F. (2018). Wild patterns: Ten years after the rise of adversarial machine learning. Pattern Recognition, 84, 317–331. DOI: 10.1016/j.patcog.2018.07.023. [Google Scholar] [Crossref]
12. Carlini, N., & Wagner, D. (2017). Towards evaluating the robustness of neural networks. In 2017 IEEE Symposium on Security and Privacy (pp. 39–57). DOI: 10.1109/SP.2017.49. [Google Scholar] [Crossref]
13. Papernot, N., McDaniel, P., Goodfellow, I., Jha, S., Celik, Z. B., & Swami, A. (2017). Practical black-box attacks against machine learning. In Proceedings of the 2017 ACM Asia Conference on Computer and Communications Security (pp. 506–519). DOI: 10.1145/3052973.3053009. [Google Scholar] [Crossref]
14. Demetrio, L., Biggio, B., Lagorio, G., Roli, F., & Armando, A. (2021). Functionality-preserving black-box optimization of adversarial Windows malware. IEEE Transactions on Information Forensics and Security, 16, 3469–3477. DOI: 10.1109/TIFS.2021.3097241. [Google Scholar] [Crossref]
15. Kolosnjaji, B., Demontis, A., Biggio, B., Maiorca, D., Giacinto, G., Roli, F., & Eckert, C. (2018). Adversarial malware binaries: Evading deep learning for malware detection in executables. In 2018 26th European Signal Processing Conference (pp. 533–537). DOI: 10.23919/EUSIPCO.2018.8553350. [Google Scholar] [Crossref]
16. Suciu, O., Coull, S. E., & Johns, J. (2019). Exploring adversarial examples in malware detection. In 2019 IEEE Security and Privacy Workshops (pp. 8–14). DOI: 10.1109/SPW.2019.00014. [Google Scholar] [Crossref]
17. Rosenberg, I., Shabtai, A., Rokach, L., & Elovici, Y. (2018). Generic black-box end-to-end attack against state-of-the-art API call based malware classifiers. In Research in Attacks, Intrusions, and Defenses (pp. 490–510). DOI: 10.1007/978-3-030-00470-5_23. [Google Scholar] [Crossref]
18. Apruzzese, G., Colajanni, M., Ferretti, L., Marchetti, M., & Guido, A. (2018). On the effectiveness of machine and deep learning for cyber security. In 2018 10th International Conference on Cyber Conflict (pp. 371–390). DOI: 10.23919/CYCON.2018.8405026. [Google Scholar] [Crossref]
19. Sommer, R., & Paxson, V. (2010). Outside the closed world: On using machine learning for network intrusion detection. In 2010 IEEE Symposium on Security and Privacy (pp. 305–316). DOI: 10.1109/SP.2010.25. [Google Scholar] [Crossref]
20. Yin, C., Zhu, Y., Fei, J., & He, X. (2017). A deep learning approach for intrusion detection using recurrent neural networks. IEEE Access, 5, 21954–21961. DOI: 10.1109/ACCESS.2017.2762418. [Google Scholar] [Crossref]
21. National Institute of Standards and Technology (NIST). (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0). DOI: 10.6028/NIST.AI.100-1. [Google Scholar] [Crossref]
22. National Institute of Standards and Technology (NIST). (2024). Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations. DOI: 10.6028/NIST.AI.100-2. [Google Scholar] [Crossref]
23. Okoli, U. I., Obi, O. C., Adewusi, A. O., & Abrahams, T. O. (2024). Machine learning in cybersecurity: A review of threat detection and defense mechanisms. World Journal of Advanced Research and Reviews, 21(1), 2286–2295. DOI: 10.30574/wjarr.2024.21.1.0315. [Google Scholar] [Crossref]
24. Girhepuje, S., Verma, A., & Raina, G. (2024). A survey on offensive AI within cybersecurity. arXiv preprint arXiv:2410.03566. DOI: 10.48550/arXiv.2410.03566. [Google Scholar] [Crossref]
25. Zhang, W. E., Sheng, Q. Z., Alhazmi, A., & Li, C. (2020). Adversarial attacks on deep-learning models in natural language processing. ACM Transactions on Intelligent Systems and Technology, 11(3), 1–41. DOI: 10.1145/3374217. [Google Scholar] [Crossref]
26. Zügner, D., Akbarnejad, A., & Günnemann, S. (2018). Adversarial attacks on neural networks for graph data. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (pp. 2847–2856). DOI: 10.1145/3219819.3220078. [Google Scholar] [Crossref]
27. Alzubaidi, L., et al. (2024). MEFF: A model ensemble feature fusion approach for tackling adversarial attacks in medical imaging. Intelligent Systems with Applications, 22, 200355. DOI: 10.1016/j.iswa.2024.200355. [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- What the Desert Fathers Teach Data Scientists: Ancient Ascetic Principles for Ethical Machine-Learning Practice
- Comparative Analysis of Some Machine Learning Algorithms for the Classification of Ransomware
- Comparative Performance Analysis of Some Priority Queue Variants in Dijkstra’s Algorithm
- Transfer Learning in Detecting E-Assessment Malpractice from a Proctored Video Recordings.
- Dual-Modal Detection of Parkinson’s Disease: A Clinical Framework and Deep Learning Approach Using NeuroParkNet