AdverShield-LLM: Adversarial Robustness Certification for IoT-Integrated Retrieval-Augmented Generation via Randomized Smoothing

Authors

Yasser Samir Hadi

Department of Computer Systems Techniques, Administrative Polytechnic College- Baghdad, Middle Technical University, Iraq (Iraq)

Article Information

DOI: 10.51244/IJRSI.2026.1305000230

Subject Category: IoT

Volume/Issue: 13/5 | Page No: 2565-2591

Publication Timeline

Submitted: 2026-05-14

Accepted: 2026-05-19

Published: 2026-06-11

Abstract

The emergence of the Internet of Things (IoT) ecosystems, Retrieval-Augmented Generation (RAG) systems have become commonplace and provide a means for embedding dynamically retrieved external knowledge in the response from a Large Language Model (LLM). While potentially helpful, IoT-enabled RAG pipelines present significant adversarial threats such as poisoning passages into the IoT knowledge base, altering dense retrieval embeddings, and conducting indirect prompt injection attacks via the inputs through the IoT sensors, all of which can impact the fidelity of generated responses and compromise the trustworthiness of the system. Current defenses are based mostly on heuristic filtering or empirical adversarial training, and are not known to be robustly certified or are fragile under adaptive adversaries.
In response to these challenges, this article introduces a new certified defense framework named AdverShield-LLM to combine the randomized smoothing technique with a multi-granular noise injection mechanism well-suited to the distributed and low latency requirements of RAG systems in the IoT domain. AdverShield-LLM consists of three synergistic modules: (i) Passage-Level Smoothed Aggregation (PLSA) module which certifies the robustness of RAG retrieval against bounded corpus poisoning under an isolate-then-smooth paradigm, (ii) Token-Adaptive Gaussian Defense (TAGD) layer that certifies LLM generation against indirect prompt injection by propagating l_2-norm perturbation bounds through the transformer attention stack, and (iii) IoT-Aware Certified Radius Scheduler (IACRS) that dynamically schedules noise budgets among constrained edge nodes while preserving the certified radius.
AdverShield-LLM is evaluated on three IoT security benchmarks—MS-RAG-IoT, NQ-Adversarial and IoTQA-Poison—with extensive experiments showing its certified accuracy is 81.4% under l_2 perturbation radius σ=0.50 compared to the strongest baseline RobustRAG which reported +9.3% accuracy, and reduced the attack success rate from 74.2% to 8.6% against PoisonedRAG. Moreover, AdverShield-LLM ensures the accuracy of clean answers within 2.1% of the undefended RAG accuracy, proving that certified robustness does not compromise the utility of RAGs in resource-limited IoT environments.

Keywords

adversarial robustness certification, IoT security, retrieval-augmented generation, randomized smoothing.

Downloads

References

1. B. C. Das, M. H. Amini, and Y. Wu, "Security and privacy challenges of large language models: A survey," ACM Comput. Surv., vol. 57, no. 6, Art. no. 152, pp. 1–39, Feb. 2025, https://dl.acm.org/doi/full/10.1145/3712001 [Google Scholar] [Crossref]

2. Y. Dong, R. Mu, Y. Zhang, S. Sun, T. Zhang, C. Wu, G. Jin, Y. Qi, J. Hu, J. Meng, S. Bensalem, and X. Huang, "Safeguarding large language models: A survey," Artif. Intell. Rev., vol. 58, no. 12, Art. no. 382, Oct. 2025, https://link.springer.com/article/10.1007/s10462-025-11389-2 [Google Scholar] [Crossref]

3. P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-T. Yih, T. Rocktäschel, S. Riedel, and D. Kiela, "Retrieval-augmented generation for knowledge-intensive NLP tasks," in Advances in Neural Information Processing Systems (NeurIPS), vol. 33, 2020, pp. 9459–9474. https://proceedings.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html [Google Scholar] [Crossref]

4. Zou, Wei, Runpeng Geng, Binghui Wang, and Jinyuan Jia. "{PoisonedRAG}: Knowledge corruption attacks to {Retrieval-Augmented} generation of large language models." In 34th USENIX Security Symposium (USENIX Security 25), pp. 3827-3844. 2025. https://www.usenix.org/conference/usenixsecurity25/presentation/zou-poisonedrag [Google Scholar] [Crossref]

5. B. Zhang, H. Xin, M. Fang, Z. Liu, B. Yi, T. Li, and Z. Liu, "Traceback of poisoning attacks to retrieval-augmented generation," in Proc. ACM Web Conf. 2025 (WWW '25), Sydney, NSW, Australia, Apr. 2025, pp. 2085–2097, https://doi.org/10.1145/3696410.3714756 [Google Scholar] [Crossref]

6. Y. Li, P. Eustratiadis, S. Lupart, and E. Kanoulas, "Unsupervised corpus poisoning attacks in continuous space for dense retrieval," in Proc. 48th Int. ACM SIGIR Conf. Res. Dev. Inf. Retr. (SIGIR '25), Padua, Italy, Jul. 2025, pp. 2452–2462, https://dl.acm.org/doi/abs/10.1145/3726302.3730110 [Google Scholar] [Crossref]

7. Q. Guo, S. Pang, X. Jia, Y. Liu, and Q. Guo, "Efficient generation of targeted and transferable adversarial examples for vision-language models via diffusion models," IEEE Trans. Inf. Forensics Security, vol. 20, pp. 1333–1348, 2024, https://ieeexplore.ieee.org/abstract/document/10812818 [Google Scholar] [Crossref]

8. Y.-A. Liu, R. Zhang, J. Guo, M. de Rijke, W. Chen, Y. Fan, and X. Cheng, "Black-box adversarial attacks against dense retrieval models: A multi-view contrastive learning method," in Proc. 32nd ACM Int. Conf. Inf. Knowl. Manage. (CIKM '23), Birmingham, United Kingdom, Oct. 2023, pp. 1647–1656, https://dl.acm.org/doi/abs/10.1145/3583780.3614793 [Google Scholar] [Crossref]

9. S. Wang, T. Zhu, B. Liu, M. Ding, D. Ye, W. Zhou, and P. Yu, "Unique security and privacy threats of large language models: A comprehensive survey," ACM Comput. Surv., vol. 58, no. 4, Art. no. 83, pp. 1–36, Oct. 2025, https://dl.acm.org/doi/full/10.1145/3764113. [Google Scholar] [Crossref]

10. Q. Tang, J. Qian, X. Du, S. Wang, and H. Liu, "Advancing adversarial and LLM robustness in trustworthy AI: A comprehensive survey," Artif. Intell. Rev., 2026, https://link.springer.com/article/10.1007/s10462-026-11558-x [Google Scholar] [Crossref]

11. J. Deng, H. Hong, A. Palmer, X. Zhou, J. Bi, K. Mahmood, Y. Hong, and D. Aguiar, "Certifying adapters: Enabling and enhancing the certification of classifier adversarial robustness," in Proc. 2025 Int. Joint Conf. Neural Netw. (IJCNN), Rome, Italy, Jun.–Jul. 2025, pp. 1–8, https://ieeexplore.ieee.org/abstract/document/11228719 [Google Scholar] [Crossref]

12. J. Cohen, E. Rosenfeld, and Z. Kolter, "Certified adversarial robustness via randomized smoothing," in Proc. 36th Int. Conf. Mach. Learn. (ICML), Long Beach, CA, USA, Jun. 2019, pp. 1310–1320. https://proceedings.mlr.press/v97/cohen19c.html [Google Scholar] [Crossref]

13. Y. Tao, Y. Shen, H. Zhang, Y. Shen, L. Wang, C. Shi, and S. Du, "Robustness of large language models against adversarial attacks," in Proc. 4th Int. Conf. Artif. Intell., Robot., Commun. (ICAIRC), Xiamen, China, 2024, pp. 182–185.https://ieeexplore.ieee.org/abstract/document/10900215 [Google Scholar] [Crossref]

14. H. Zhang, W. Shao, H. Liu, Y. Ma, P. Luo, Y. Qiao, N. Zheng, and K. Zhang, "B-avibench: Toward evaluating the robustness of large vision-language model on black-box adversarial visual-instructions," IEEE Trans. Inf. Forensics Security, vol. 20, pp. 1434–1446, 2024. https://ieeexplore.ieee.org/abstract/document/10816024 [Google Scholar] [Crossref]

15. O. Muliarevych, "Enhancing system security: LLM-driven defense against prompt injection vulnerabilities," in Proc. IEEE 17th Int. Conf. Adv. Trends Radioelectron., Telecommun. Comput. Eng. (TCSET), Lviv-Slavske, Ukraine, 2024, pp. 420–423. https://ieeexplore.ieee.org/abstract/document/10755823 [Google Scholar] [Crossref]

16. X. Yin, C. Ni, and S. Wang, "Multitask-based evaluation of open-source LLM on software vulnerability," IEEE Trans. Softw. Eng., vol. 50, no. 11, pp. 3071–3087, Nov. 2024. https://ieeexplore.ieee.org/abstract/document/10706805 [Google Scholar] [Crossref]

17. Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen et al., "A survey on evaluation of large language models," ACM Trans. Intell. Syst. Technol., vol. 15, no. 3, pp. 1–45, 2024. https://dl.acm.org/doi/full/10.1145/3641289 [Google Scholar] [Crossref]

18. J. Shi, Z. Yuan, Y. Liu, Y. Huang, P. Zhou, L. Sun, and N. Z. Gong, "Optimization-based prompt injection attack to LLM-as-a-judge," in Proc. ACM SIGSAC Conf. Comput. Commun. Security (CCS), Salt Lake City, UT, USA, Oct. 2024, pp. 660–674. https://dl.acm.org/doi/abs/10.1145/3658644.3690291 [Google Scholar] [Crossref]

19. X. Zhang, C. Zhang, T. Li, Y. Huang, X. Jia, M. Hu, J. Zhang, Y. Liu, S. Ma, and C. Shen, "Jailguard: A universal detection framework for prompt-based attacks on LLM systems," ACM Trans. Softw. Eng. Methodol., vol. 35, no. 1, pp. 1–40, 2025. https://dl.acm.org/doi/full/10.1145/3724393 [Google Scholar] [Crossref]

20. E. Mathew, "Enhancing security in large language models: A comprehensive review of prompt injection attacks and defenses," Authorea Preprints, 2024. https://www.techrxiv.org/doi/full/10.36227/techrxiv.172954263.32914470 [Google Scholar] [Crossref]

21. N. O. Jaffal, M. Alkhanafseh, and D. Mohaisen, "Large language models in cybersecurity: A survey of applications, vulnerabilities, and defense techniques," AI, vol. 6, no. 9, p. 216, 2025. https://doi.org/10.3390/ai6090216 [Google Scholar] [Crossref]

22. Joshi and S. Baidya, "Securing the cognitive layer: A survey on security threats, defenses, and privacy-preserving architectures for LLM-IoT integration," J. Cybersecurity Privacy, vol. 6, no. 2, p. 63, 2026. https://doi.org/10.3390/jcp6020063 [Google Scholar] [Crossref]

23. M. Kumar, J. K. Samriya, G. K. Walia, P. Verma, H. Wu, and S. S. Gill, "Blockchain empowered secure federated learning for consumer IoT applications in cloud-edge collaborative environment," IEEE Trans. Consumer Electron., vol. 71, no. 2, pp. 3986–3996, 2025. https://ieeexplore.ieee.org/abstract/document/10849591 [Google Scholar] [Crossref]

24. K. Peng, P. Xiao, S. Wang, and V. C. M. Leung, "SCOF: Security-aware computation offloading using federated reinforcement learning in industrial internet of things with edge computing," IEEE Trans. Services Comput., vol. 17, no. 4, pp. 1780–1792, 2024. https://ieeexplore.ieee.org/abstract/document/10473157 [Google Scholar] [Crossref]

25. W. Issa, N. Moustafa, B. Turnbull, N. Sohrabi, and Z. Tari, "Blockchain-based federated learning for securing internet of things: A comprehensive survey," ACM Comput. Surv., vol. 55, no. 9, pp. 1–43, 2023. https://dl.acm.org/doi/full/10.1145/3560816 [Google Scholar] [Crossref]

26. S. G. AboulEla and R. F. Kashef, "Leveraging large language models, graph neural networks, and explainable AI for revolutionizing the next-generation network intrusion detection systems," J. Intell. Inf. Syst., vol. 63, no. 5, pp. 1807–1835, 2025. https://link.springer.com/article/10.1007/s10844-025-00964-2 [Google Scholar] [Crossref]

27. V. Padmavathi and R. Saminathan, "A federated edge intelligence framework with trust based access control for secure and privacy preserving IoT systems," Sci. Rep., vol. 15, no. 1, p. 35832, 2025. https://www.nature.com/articles/s41598-025-19712-1 [Google Scholar] [Crossref]

28. S. Sattarpour, A. Barati, and H. Barati, "EBIDS: Efficient BERT-based intrusion detection system in the network and application layers of IoT," Cluster Comput., vol. 28, no. 2, p. 138, 2025. https://link.springer.com/article/10.1007/s10586-024-04775-y [Google Scholar] [Crossref]

29. Z. Xiang, D. J. Miller, and G. Kesidis, "Detection of backdoors in trained classifiers without access to the training set," IEEE Trans. Neural Netw. Learn. Syst., vol. 33, no. 3, pp. 1177–1191, Mar. 2022. https://ieeexplore.ieee.org/abstract/document/9296553 [Google Scholar] [Crossref]

30. S. A. Sharaf and S. Nooh, "Identifying significant features in adversarial attack detection framework using federated learning empowered medical IoT network security," Sci. Rep., vol. 15, no. 1, p. 31485, 2025. https://www.nature.com/articles/s41598-025-14913-0 [Google Scholar] [Crossref]

31. S. Salim, N. Moustafa, and A. Almorjan, "Responsible deep-federated-learning-based threat detection for satellite communications," IEEE Internet Things J., vol. 12, no. 5, pp. 4807–4819, 2025. https://ieeexplore.ieee.org/abstract/document/10847856 [Google Scholar] [Crossref]

32. C. Wang, H. Li, W. Song, and Y. Lin, "Retrieval-augmented generation: A survey of security challenges and countermeasures," in Proc. 11th IEEE Int. Conf. Privacy Comput. Data Security (PCDS), 2025, pp. 210–217 https://ieeexplore.ieee.org/abstract/document/11172756 [Google Scholar] [Crossref]

33. V. Karpukhin, B. Oguz, S. Min, P. Lewis, L. Wu, S. Edunov, D. Chen, and W. Yih, "Dense passage retrieval for open-domain question answering," in Proc. Conf. Empirical Methods Natural Lang. Process. (EMNLP), Online, Nov. 2020, pp. 6769–6781.https://aclanthology.org/2020.emnlp-main.550/ [Google Scholar] [Crossref]

34. . V. Padmavathi and R. Saminathan, "A federated edge intelligence framework with trust based access control for secure and privacy preserving IoT systems," Sci. Rep., vol. 15, no. 1, p. 35832, 2025. https://www.nature.com/articles/s41598-025-19712-1 [Google Scholar] [Crossref]

35. Y. Lu, X. Huang, Y. Dai, S. Maharjan, and Y. Zhang, "Blockchain and federated learning for privacy-preserved data sharing in industrial IoT," IEEE Trans. Ind. Informat., vol. 16, no. 6, pp. 4177–4186, Jun. 2020. https://ieeexplore.ieee.org/abstract/document/8843900 [Google Scholar] [Crossref]

36. Y. Shi, K. Wei, L. Shen, J. Li, X. Wang, B. Yuan, and S. Guo, "Efficient federated learning with enhanced privacy via lottery ticket pruning in edge computing," IEEE Trans. Mobile Comput., vol. 23, no. 10, pp. 9946–9958, Oct. 2024. https://ieeexplore.ieee.org/abstract/document/10452820 [Google Scholar] [Crossref]

37. Y. Zhang, Y. Li, L. Cui, D. Cai, L. Liu, T. Fu, X. Huang et al., "Siren's song in the AI ocean: A survey on hallucination in large language models," Comput. Linguist., vol. 51, no. 4, pp. 1373–1418, 2025. https://doi.org/10.1162/COLI.a.16 [Google Scholar] [Crossref]

38. Q. Wu, K. He, and X. Chen, "Personalized federated learning for intelligent IoT applications: A cloud-edge based framework," IEEE Open J. Comput. Soc., vol. 1, pp. 35–44, 2020. https://ieeexplore.ieee.org/abstract/document/9090366 [Google Scholar] [Crossref]

39. X. Jia, Y. Chen, X. Mao, R. Duan, J. Gu, R. Zhang, H. Xue, Y. Liu, and X. Cao, "Revisiting and exploring efficient fast adversarial training via LAW: Lipschitz regularization and auto weight averaging," IEEE Trans. Inf. Forensics Security, vol. 19, pp. 8125–8139, 2024. https://ieeexplore.ieee.org/abstract/document/10574880 [Google Scholar] [Crossref]

40. E. Dai, T. Zhao, H. Zhu, J. Xu, Z. Guo, H. Liu, J. Tang, and S. Wang, "A comprehensive survey on trustworthy graph neural networks: Privacy, robustness, fairness, and explainability," Mach. Intell. Res., vol. 21, no. 6, pp. 1011–1061, 2024. https://link.springer.com/article/10.1007/s11633-024-1510-8 [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles