Enhancing Phishing URL Detection Using a Two-Level Rule-Based Framework Combining Lexical and RDAP Registration Features

Authors

Wan Afifie Aliff Bin Wan Abdullah

ADTEC JTM Kampus Selandar, 77500 Selandar, Melaka (Malaysia)

Zulkiflee Muslim

Faculty of Artificial Intelligence and Cyber Security, Universiti Teknikal Malaysia Melaka (UTeM), Hang Tuah Jaya, 76100 Durian Tunggal, Melaka (Malaysia)

Haniza Nahar

Faculty of Information and Communication Technology, Universiti Teknikal Malaysia Melaka (UTeM), Hang Tuah Jaya, 76100 Durian Tunggal, Melaka (Malaysia)

Radzi Motsidi

School of Technical Foundation and Diploma Studies, Universiti Teknikal Malaysia Melaka (UTeM), Hang Tuah Jaya, 76100 Durian Tunggal, Melaka (Malaysia)

Article Information

DOI: 10.47772/IJRISS.2026.100800025

Subject Category: Computer Science

Volume/Issue: 10/8 | Page No: 330-347

Publication Timeline

Submitted: 2026-08-14

Accepted: 2026-08-19

Published: 2026-08-25

Abstract

Phishing remains one of the most persistent cyber threats, and almost every campaign ultimately depends on a deceptive Uniform Resource Locator (URL). Existing defences face a structural trade-off: blacklists are reactive and cannot cover newly registered domains during the zero-hour window, while machine-learning detectors, although accurate, are opaque, feature-hungry, and often depend on page content or full DNS telemetry that many organisations cannot collect. This study proposes and evaluates a lightweight, fully interpretable two-level rule-based framework that fuses lexical URL features with domain registration evidence retrieved through the Registration Data Access Protocol (RDAP). Level 1 scores each URL using five transparent lexical rules derived from training-set distributions of domain length, number of dots, number of hyphens, number of digits, and URL entropy. Level 2 applies three RDAP rules covering domain age, days to expiry, and a missing-registration-data flag, targeting the young, short-lived, and poorly documented domains that characterise phishing infrastructure. The two levels are combined through logical OR and AND decision fusion and evaluated on a balanced, held-out set of 400 URLs drawn from a curated corpus of 800. Level 1 achieved 95.50% accuracy (precision 0.9789, recall 0.9300); Level 2 achieved perfect recall (1.0000) at 0.8969 precision; OR fusion preserved perfect recall; and AND fusion delivered the best overall result at 96.50% accuracy with perfect precision, zero false positives, and a Matthews Correlation Coefficient of 0.9323. A confusion-matrix decomposition further shows that the false-positive sets of the two levels are completely disjoint, confirming that lexical and registration evidence fail independently. Exploiting this, a cascaded implementation of AND fusion reproduces identical decisions while issuing RDAP queries for only 47.5% of URLs, a 52.5% reduction in external lookups.

Keywords

Phishing detection, URL lexical features, RDAP, rule-based classification, decision fusion, interpretable security

Downloads

References

1. Safi, A., & Singh, S. (2023). A systematic literature review on phishing website detection techniques. Journal of King Saud University – Computer and Information Sciences, 35(2). https://doi.org/10.1016/j.jksuci.2023.01.004 [Google Scholar] [Crossref]

2. Al-Qahtani, A. F., & Cresci, S. (2022). The COVID-19 scamdemic: A survey of phishing attacks and their countermeasures during COVID-19. IET Information Security, 16(5). https://doi.org/10.1049/ise2.12073 [Google Scholar] [Crossref]

3. Kaur, B. (2025). Social engineering attacks in the digital age. Theseus. https://www.theseus.fi/handle/10024/896456 [Google Scholar] [Crossref]

4. Muhammad Saeed Liaqat, Mumtaz, G., Rasheed, N., & Mubeen, Z. (2024). Exploring phishing attacks in the AI age: A comprehensive literature review. Journal of Computing & Biomedical Informatics, 7(02). [Google Scholar] [Crossref]

5. Alnemari, S., & Alshammari, M. (2023). Detecting phishing domains using machine learning. Applied Sciences, 13(8), 4649. https://doi.org/10.3390/app13084649 [Google Scholar] [Crossref]

6. Madupati, B. (2024). Cyber attacks in the remote work era: An analysis of phishing, ransomware, and mitigation strategies. International Journal of Science and Research, 13(9), 703–708. https://doi.org/10.21275/sr24903073257 [Google Scholar] [Crossref]

7. Akter, T. (2025). A taxonomy and multi-layered defense framework for generative AI-powered phishing campaigns on social media platforms. University of Turku. https://www.utupub.fi/handle/10024/182730 [Google Scholar] [Crossref]

8. Nagunwa, T. (2024). Detection of phishing websites hosted in name server flux networks using machine learning. Journal of Computer Science, 20(1), 10–32. https://doi.org/10.3844/jcssp.2024.10.32 [Google Scholar] [Crossref]

9. Sun, X., & Liu, Z. (2023). Domain generation algorithms detection with feature extraction and domain center construction. PLOS ONE, 18(1), e0279866. https://doi.org/10.1371/journal.pone.0279866 [Google Scholar] [Crossref]

10. Aljabri, M., Altamimi, H. S., Albelali, S. A., Al-Harbi, M., Alhuraib, H. T., Alotaibi, N. K., Alahmadi, A. A., Alhaidari, F., Mohammad, R. M. A., & Salah, K. (2022). Detecting malicious URLs using machine learning techniques: Review and research directions. IEEE Access, 10, 121395–121417. https://doi.org/10.1109/access.2022.3222307 [Google Scholar] [Crossref]

11. Jadhav, S. (2023). Phishing website detector using ML. International Journal for Research in Applied Science and Engineering Technology, 11(5), 5509–5514. https://doi.org/10.22214/ijraset.2023.52872 [Google Scholar] [Crossref]

12. Fernando, M., Mahmood, A., Jabed, M., & He, Z. (2025). Phishlex: A real-time machine learning model for zero-day phishing detection by systematizing URL techniques. https://doi.org/10.2139/ssrn.5205908 [Google Scholar] [Crossref]

13. Vyawhare, C. R., Totare, R. Y., Sonawane, P. S., & Deshmukh, P. B. (2022). Machine learning system for malicious website detection: A literature review. International Journal for Research in Applied Science and Engineering Technology, 10(5), 56–61. https://doi.org/10.22214/ijraset.2022.42050 [Google Scholar] [Crossref]

14. Choo, E., Nabeel, M., Kim, D., De Silva, R., Yu, T., & Khalil, I. (2024). A large scale study and classification of VirusTotal reports on phishing and malware URLs. ACM SIGMETRICS Performance Evaluation Review, 52(1), 55–56. https://doi.org/10.1145/3673660.3655042 [Google Scholar] [Crossref]

15. Arun, A., & Abosata, N. (2024). Next generation of phishing attacks using AI powered browsers. arXiv. https://doi.org/10.48550/arxiv.2406.12547 [Google Scholar] [Crossref]

16. Wong, A., Abuadbba, A., Almashor, M., & Kanhere, S. (2022). PhishClone: Measuring the efficacy of cloning evasion attacks. arXiv. https://doi.org/10.48550/arXiv.2209.01582 [Google Scholar] [Crossref]

17. Abuadbba, A., Wang, S., Almashor, M., Ahmed, M. E., Gaire, R., Camtepe, S., & Nepal, S. (2022). Towards web phishing detection limitations and mitigation. arXiv. https://doi.org/10.48550/arXiv.2204.00985 [Google Scholar] [Crossref]

18. Thakur, K., Ali, M. L., Obaidat, M. A., & Kamruzzaman, A. (2023). A systematic review on deep-learning-based phishing email detection. Electronics, 12(21), 4545. https://doi.org/10.3390/electronics12214545 [Google Scholar] [Crossref]

19. Magdy, S., Abouelseoud, Y., & Mikhail, M. (2022). Efficient spam and phishing emails filtering based on deep learning. Computer Networks, 206, 108826. https://doi.org/10.1016/j.comnet.2022.108826 [Google Scholar] [Crossref]

20. Pandey, P., & Mishra, N. (2023). Phish-Sight: A new approach for phishing detection using dominant colors on web pages and machine learning. International Journal of Information Security. https://doi.org/10.1007/s10207-023-00672-4 [Google Scholar] [Crossref]

21. Lübbers, H. (2023). Investigating phishing attacks using the Registration Data Access Protocol (RDAP). CEUR Workshop Proceedings, 3631. https://ceur-ws.org/Vol-3631/paper6.pdf [Google Scholar] [Crossref]

22. Loffredo, M., & Martinelli, M. (2024). RFC 9536: Registration Data Access Protocol (RDAP) reverse search. RFC Editor. https://www.rfc-editor.org/rfc/rfc9536.pdf [Google Scholar] [Crossref]

23. Harrison, T., & Singh, J. (2026). RFC 9910: Registration Data Access Protocol (RDAP) Regional Internet Registry (RIR) search. RFC Editor. https://www.rfc-editor.org/rfc/rfc9910.pdf [Google Scholar] [Crossref]

24. Newton, A. (2026, February 10). The current state of RDAP. APNIC Blog. https://blog.apnic.net/2026/02/10/the-current-state-of-rdap/ [Google Scholar] [Crossref]

25. Agarwal, S., & Vasek, M. (2025). Examining newly registered phishing domains at scale. Workshop on the Economics of Information Security. https://discovery.ucl.ac.uk/id/eprint/10209951/ [Google Scholar] [Crossref]

26. Chiang, P.-H., & Tsai, S.-C. (2024). Detection of malicious domains with concept drift using ensemble learning. IEEE Transactions on Network and Service Management, 21(6), 6796–6809. https://doi.org/10.1109/TNSM.2024.3435516 [Google Scholar] [Crossref]

27. Hassaoui, M., Hanini, M., & El Kafhali, S. (2024). Data science in cybersecurity to detect malware-based domain generation algorithm: Improvement, challenges, and prospects. Journal of Computational and Cognitive Engineering, 3(3), 213–225. https://doi.org/10.47852/bonviewjcce42022875 [Google Scholar] [Crossref]

28. Thein, T. T., Shiraishi, Y., & Morii, M. (2023). Malicious domain detection based on decision tree. IEICE Transactions on Information and Systems, E106.D(9), 1490–1494. https://doi.org/10.1587/transinf.2022ofl0002 [Google Scholar] [Crossref]

29. Hung, N. V., & Shone, N. (2025). Detecting DGA-managed thingbots using DNS traffic. Journal of Control Engineering and Applied Informatics, 27(2). https://doi.org/10.61416/ceai.v27i2.9302 [Google Scholar] [Crossref]

30. Ma, S., Pang, T., Cui, R., & Yang, D. (2024). A malicious domain detection method based on DNS logs. 283–288. https://doi.org/10.1109/icbctis64495.2024.00051 [Google Scholar] [Crossref]

31. Machmeier, S., & Heuveline, V. (2024). Detecting DNS tunnelling and data exfiltration using dynamic time warping. 2024 8th Cyber Security in Networking Conference (CSNet), 83–91. https://doi.org/10.1109/csnet64211.2024.10851475 [Google Scholar] [Crossref]

32. Kibreab Adane, Beyene, B., & Abebe, M. (2023). Single and hybrid-ensemble learning-based phishing website detection: Examining impacts of varied nature datasets and informative feature selection technique. Digital Threats: Research and Practice. https://doi.org/10.1145/3611392 [Google Scholar] [Crossref]

33. Kayode-Ajala, O. (2023). Applying machine learning algorithms for detecting phishing websites: Applications of SVM, KNN, decision trees, and random forests. International Journal of Information and Cybersecurity. [Google Scholar] [Crossref]

34. Choudhary, T., Mhapankar, S., Bhddha, R., Kharuk, A., & Patil, R. (2023). A machine learning approach for phishing attack detection. Journal of Artificial Intelligence and Technology, 3(3), 108–113. https://doi.org/10.37965/jait.2023.0197 [Google Scholar] [Crossref]

35. Awasthi, A., & Goel, N. (2022). Phishing website prediction using base and ensemble classifier techniques with cross-validation. Cybersecurity, 5(1). https://doi.org/10.1186/s42400-022-00126-9 [Google Scholar] [Crossref]

36. Kocyigit, E., Korkmaz, M., Sahingoz, O. K., & Diri, B. (2024). Enhanced feature selection using genetic algorithm for machine-learning-based phishing URL detection. Applied Sciences, 14(14). https://doi.org/10.3390/app14146081 [Google Scholar] [Crossref]

37. Chong, J. C., Sim, N. Y., Khoh, C. W., & Yi, L. T. (2025). Phishing attack detection on URLs using KNN, RF, DT with GA and k-fold cross validation approach. International Journal of Research and Innovation in Social Science, 9(1), 1623–1641. https://doi.org/10.47772/IJRISS.2025.9010134 [Google Scholar] [Crossref]

38. Ari Kustiawan, Y., & Ghauth, K. I. (2025). Evaluating the impact of feature engineering in phishing URL detection: A comparative study of URL, HTML, and derived features. IEEE Access, 13, 126756–126768. https://doi.org/10.1109/access.2025.3579223 [Google Scholar] [Crossref]

39. Sabir, B., Babar, M. A., Gaire, R., & Abuadbba, A. (2022). Reliability and robustness analysis of machine learning based phishing URL detectors. IEEE Transactions on Dependable and Secure Computing, 1–18. https://doi.org/10.1109/tdsc.2022.3218043 [Google Scholar] [Crossref]

40. Shafin, S. S. (2024). An explainable feature selection framework for web phishing detection with machine learning. Data Science and Management. https://doi.org/10.1016/j.dsm.2024.08.004 [Google Scholar] [Crossref]

41. Balogun, A. O., Adewole, K. S., Raheem, M. O., Akande, O. N., Usman-Hamza, F. E., Mabayoje, M. A., Akintola, A. G., Asaju-Gbolagade, A. W., Jimoh, M. K., Jimoh, R. G., & Adeyemo, V. E. (2021). Optimized decision forest for website phishing detection. Lecture Notes in Networks and Systems, 568–582. https://doi.org/10.1007/978-3-030-90321-3_47 [Google Scholar] [Crossref]

42. Santelices, R. B. (2025). A students' perspective on cybersecurity awareness and education. International Journal of Research and Innovation in Social Science, 9(IIIS), 7976–7988. https://doi.org/10.47772/IJRISS.2025.903SEDU0597 [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles