Evaluation Metrics for Deep Learning–Based Semantic Search: A Critical Review with a User Satisfaction Perspective

Authors

ADESANYA Adetola Joel

Department of Computer Science, Lead City University, Ibadan, Oyo State, Nigeria. (Nigeria)

AYOADE Akintayo Michael

Department of Computer Science, Lead City University, Ibadan, Oyo State, Nigeria. (Nigeria)

Folarin Israel Bolaji

Department of Computer Science, Lead City University, Ibadan, Oyo State, Nigeria. (Nigeria)

Article Information

DOI: 10.51244/IJRSI.2026.1305000044

Subject Category: Computer Science

Volume/Issue: 13/5 | Page No: 479-490

Publication Timeline

Submitted: 2026-05-04

Accepted: 2026-05-16

Published: 2026-05-26

Abstract

Background: Semantic search, driven by deep learning models like BERT and Sentence-BERT (SBERT), has greatly improved information retrieval. It has shifted from matching keywords to capturing the context of search and user intent. However, to evaluate how effective these systems are, traditional system-focused metrics such as precision, recall, Mean Reciprocal Rank (MRR), and Normalized Discounted Cumulative Gain (NDCG) are still used. These metrics do not adequately reflect user experience. They often overlook important behavioral and contextual factors such as user engagement, search satisfaction, relevance perception, and interaction quality in real-world environments. This review examines existing evaluation metrics for deep learning-based semantic search. It identifies their strengths and limitations, as well as how well they capture real-world user satisfaction. It also explores helpful ways to incorporate user-centered approaches into the evaluation of these systems.
Method: A critical review approach was used, synthesizing literature from 2020 to 2025 across databases like IEEE Xplore, ACM Digital Library, Scopus, and Google Scholar. Studies on semantic search evaluation, deep learning-based retrieval, and user-centered metrics were thematically analyzed for information. The reviewed studies were selected using predefined inclusion and exclusion criteria, and the analysis categorized evaluation methods into traditional and user-centered approaches.
Findings: The review finds that while traditional metrics provide reproducibility and comparability, they fail to capture important aspects of user experience such as clarity, usability, and satisfaction. Emerging user-oriented alternatives like click-through rates, dwell time, and satisfaction surveys offer valuable insights, but they remain secondary, fragmented, and lack standardization. The review highlights an ongoing gap between the leaderboard performance of search systems and their real-world utility. The review further reveals that many high-performing semantic retrieval systems achieve strong benchmark scores while still failing to fully satisfy users in practical search scenarios.
Conclusion: Semantic search evaluation must change from traditional, system-focused measures to hybrid metrics that integrate algorithmic precision with user-centered awareness. By combining these traditional metrics with behavioral signals and subjective feedback, future evaluation methods can ensure that semantic search systems are not only technically sound but also practical, usable, and satisfying for end-users. The study therefore recommends the development of standardized hybrid evaluation frameworks capable of balancing retrieval accuracy with measurable user experience indicators.

Keywords

Semantic search, information retrieval, evaluation metrics, deep learning

Downloads

References

1. L. Wang, N. Yang, X. Huang, L. Yang, R. Majumder, and F. Wei, "Multilingual e5 text embeddings: A technical report," arXiv preprint, arXiv:2402.05672, 2024. [Google Scholar] [Crossref]

2. X. Li and J. Li, "Angle-optimized text embeddings," arXiv preprint, arXiv:2309.12871, 2023. [Google Scholar] [Crossref]

3. A. Esteva, A. Kale, R. Paulus, K. Hashimoto, W. Yin, D. Radev, and R. Socher, “COVID-19 information retrieval with deep-learning based semantic search, question answering, and abstractive summarization,” NPJ Digit. Med., vol. 4, no. 1, p. 68, 2021. [Google Scholar] [Crossref]

4. M.-K. Ghali, A. Farrag, D. Won, and Y. Jin, “Enhancing knowledge retrieval with in-context learning and semantic search through generative AI,” Knowl.-Based Syst., 2025, Art. no. 113047. [Google Scholar] [Crossref]

5. T. Hellert, J. Montenegro, and A. Pollastro, "PhysBERT: A text embedding model for physics scientific literature," APL Mach. Learn., vol. 2, no. 4, 2024. [Google Scholar] [Crossref]

6. L. Wang, N. Yang, X. Huang, L. Yang, R. Majumder, and F. Wei, "Improving text embeddings with large language models," arXiv preprint, arXiv:2401.00368, 2023. [Google Scholar] [Crossref]

7. A. Moffat, "Batch evaluation metrics in information retrieval: Measures, scales, and meaning," IEEE Access, vol. 10, pp. 105564–105577, 2022. [Google Scholar] [Crossref]

8. L. L. de Oliveira, D. S. Vargas, A. M. A. Alexandre, F. C. Cordeiro, D. D. S. M. Gomes, M. D. C. Rodrigues, ... and V. P. Moreira, "Evaluating and mitigating the impact of OCR errors on information retrieval," Int. J. Digit. Libr., vol. 24, no. 1, pp. 45–62, 2023. [Google Scholar] [Crossref]

9. X. Li, J. Jin, Y. Zhou, Y. Zhang, P. Zhang, Y. Zhu, and Z. Dou, "From matching to generation: A survey on generative information retrieval," ACM Trans. Inf. Syst., vol. 43, no. 3, pp. 1–62, 2025. [Google Scholar] [Crossref]

10. A. Salemi and H. Zamani, "Evaluating retrieval quality in retrieval-augmented generation," in Proc. 47th Int. ACM SIGIR Conf. Res. Develop. Inf. Retrieval, July 2024, pp. 2395–2400. [Google Scholar] [Crossref]

11. W. Zhang, M. Zhang, S. Wu, J. Pei, Z. Ren, M. de Rijke, ... and P. Ren, "Excluir: Exclusionary neural information retrieval," in Proc. AAAI Conf. Artif. Intell., vol. 39, no. 12, Apr. 2025, pp. 13295–13303. [Google Scholar] [Crossref]

12. M. Gaur, K. Gunaratna, V. Srinivasan, and H. Jin, "Iseeq: Information seeking question generation using dynamic meta-information retrieval and knowledge graphs," in Proc. AAAI Conf. Artif. Intell., vol. 36, no. 10, June 2022, pp. 10672–10680. [Google Scholar] [Crossref]

13. N. Girdhar, M. Coustaty, and A. Doucet, "Digitizing history: transitioning historical paper documents to digital content for information retrieval and mining—a comprehensive survey," IEEE Trans. Comput. Social Syst., 2024. [Google Scholar] [Crossref]

14. O. Rainio, J. Teuho, and R. Klén, "Evaluation metrics and statistical tests for machine learning," Sci. Rep., vol. 14, no. 1, p. 6086, 2024. [Google Scholar] [Crossref]

15. X. Qi, Y. Zhang, S. Cao, S. Yan, and H. Su, "Human–computer interaction based on the intelligent information retrieval method for customer satisfaction in power system service," Int. J. Model. Simul. Sci. Comput., vol. 14, no. 01, p. 2341004, 2023. [Google Scholar] [Crossref]

16. A. M. Flores, M. C. Pavan, and I. Paraboni, "User profiling and satisfaction inference in public information access services," J. Intell. Inf. Syst., vol. 58, no. 1, pp. 67–89, 2022. [Google Scholar] [Crossref]

17. C. Siro, M. Aliannejadi, and M. de Rijke, "Understanding user satisfaction with task-oriented dialogue systems," in Proc. 45th Int. ACM SIGIR Conf. Res. Develop. Inf. Retrieval, 2022, pp. 2018–2023. [Google Scholar] [Crossref]

18. C. Siro, M. Aliannejadi, and M. De Rijke, "Understanding and predicting user satisfaction with conversational recommender systems," ACM Trans. Inf. Syst., vol. 42, no. 2, pp. 1–37, 2023. [Google Scholar] [Crossref]

19. C. Bauer, M. Fröbe, D. Jannach, U. Kruschwitz, P. Rosso, D. Spina, and N. Tintarev, "Overcoming methodological challenges in information retrieval and recommender systems through awareness and education," Bauer et al. [2023a], pp. 51–67. [Google Scholar] [Crossref]

20. T. E. Kim and A. Lipani, "A multi-task based neural model to simulate users in goal oriented dialogue systems," in Proc. 45th Int. ACM SIGIR Conf. Res. Develop. Inf. Retrieval, July 2022, pp. 2115–2119. [Google Scholar] [Crossref]

21. M. B. Yılmaz and K. Rızvanoğlu, "Understanding users’ behavioral intention to use voice assistants on smartphones through the integrated model of user satisfaction and technology acceptance: a survey approach," J. Eng. Des. Technol., vol. 20, no. 6, pp. 1738–1764, 2022. [Google Scholar] [Crossref]

22. C. L. Hsu and J. C. C. Lin, "Understanding the user satisfaction and loyalty of customer service chatbots," J. Retail. Consum. Serv., vol. 71, p. 103211, 2023. [Google Scholar] [Crossref]

23. J. Guo, Y. Cai, Y. Fan, F. Sun, R. Zhang, and X. Cheng, "Semantic models for the first-stage retrieval: A comprehensive review," ACM Trans. Inf. Syst. (TOIS), vol. 40, no. 4, pp. 1–42, 2022. [Google Scholar] [Crossref]

24. X. Li, K. Dong, Y. Q. Lee, W. Xia, H. Zhang, X. Dai, ... and R. Tang, "CoIR: A comprehensive benchmark for code information retrieval models," arXiv, preprint, arXiv:2407.02883, 2024. [Google Scholar] [Crossref]

25. H. W. Kim, D. H. Shin, J. Kim, G. H. Lee, and J. W. Cho, "Assessing the performance of ChatGPT's responses to questions related to epilepsy: a cross-sectional study on natural language processing and medical information retrieval," Seizure: Eur. J. Epilepsy, vol. 114, pp. 1–8, 2024. [Google Scholar] [Crossref]

26. J. Gao, C. Xiong, P. Bennett, and N. Craswell, Neural Approaches to Conversational Information Retrieval, vol. 44. Heidelberg: Springer, 2023. [Google Scholar] [Crossref]

27. N. Thakur, L. Bonifacio, M. Fröbe, A. Bondarenko, E. Kamalloo, M. Potthast, ... and J. Lin, "Systematic evaluation of neural retrieval models on the Touché 2020 argument retrieval subset of BEIR," in Proc. 47th Int. ACM SIGIR Conf. Res. Develop. Inf. Retrieval, July 2024, pp. 1420–1430. [Google Scholar] [Crossref]

28. W. Gao, H. Wang, Q. Liu, F. Wang, X. Lin, L. Yue, ... and S. Wang, "Leveraging transferable knowledge concept graph embedding for cold-start cognitive diagnosis," in Proc. 46th Int. ACM SIGIR Conf. Res. Develop. Inf. Retrieval, July 2023, pp. 983–992. [Google Scholar] [Crossref]

29. J. I. Friese and N. Fuhr, "Towards Reproducibility of Interactive Retrieval Experiments: Framework and Case Study," in Proc. Eur. Conf. Inf. Retrieval (ECIR), Apr. 2025, pp. 146–160. Cham: Springer Nature Switzerland. [Google Scholar] [Crossref]

30. A. Anand, L. Lyu, M. Idahl, Y. Wang, J. Wallat, and Z. Zhang, "Explainable information retrieval: A survey," arXiv preprint, arXiv:2211.02405, 2022. [Google Scholar] [Crossref]

31. L. Heryawan, D. Novitaningrum, K. R. Nastiti, and S. N. Mahmudah, "Medical record document search with TF-IDF and vector space model (VSM)," Int. J. Adv. Sci., Eng. Inf. Technol., vol. 14, no. 3, pp. 847–852, 2024. [Google Scholar] [Crossref]

32. F. M. Hasyim and F. Fahmi, "A literature review: The importance of term normalization in vector space mode," Literatify: Trends Libr. Develop., vol. 5, no. 1, pp. 18–28, 2024. [Google Scholar] [Crossref]

33. C. I. Nakpih, "A modified Vector Space Model for semantic information retrieval," Nat. Lang. Process. J., vol. 8, p. 100081, 2024. [Google Scholar] [Crossref]

34. D. Paseru, R. Kuera, and S. Salmon, "Comparison of Vector Space Model and BM25F Methods in Book Search," J. RESTIKOM: Riset Tek. Inform. Komput., vol. 5, no. 3, pp. 513–522, 2023. [Google Scholar] [Crossref]

35. M. A. A. Shiddiqi and A. Sanmarino, "Vector Space Model-based Information Retrieval Systems at South Sumatera Regional Libraries," J. Comput. Sci. Appl. Eng. (JOSAPEN), vol. 1, no. 2, pp. 49–53, 2023. [Google Scholar] [Crossref]

36. P. B. Bahtera and D. S. Kartawijaya, "Content Classification of the Official Website of the Ministry of Foreign Affairs of the Republic of Indonesia (MoFA RI) using Vector Space Model (VSM)," MALCOM: Indones. J. Mach. Learn. Comput. Sci., vol. 4, no. 4, pp. 1309–1319, 2024. [Google Scholar] [Crossref]

37. B. P. Zen, I. Susanto, K. Putriyani, and Sintiya, "Automatic document classification for Tempo news articles about COVID-19 based on term frequency, inverse document frequency (TF-IDF), and Vector Space Model (VSM)," in AIP Conf. Proc., vol. 2952, no. 1, July 2024, p. 060003. [Google Scholar] [Crossref]

38. J. J. M. De Araujo and A. A. J. Sinlae, "The Visualization of the Vector Space Model in Searching for Immigration News in the East Nusa Tenggara Region," in Proc. Nat. Conf. Electr. Eng., Informatics, Ind. Technol. Creative Media, vol. 3, no. 1, pp. 924–931, 2023. [Google Scholar] [Crossref]

39. Y. Wang, Y. Hou, H. Wang, Z. Miao, S. Wu, Q. Chen, ... and M. Yang, "A neural corpus indexer for document retrieval," Adv. Neural Inf. Process. Syst., vol. 35, pp. 25600–25614, 2022. [Google Scholar] [Crossref]

40. S. Bruch, C. Lucchese, M. Maistro, and F. M. Nardini, "Special Section on Efficiency in Neural Information Retrieval," ACM Trans. Inf. Syst., vol. 42, no. 5, pp. 1–4, 2024. [Google Scholar] [Crossref]

41. T. Thakur, N. Reimers, J. Daxenberger, and I. Gurevych, “BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models,” in Proc. 35th Conf. on Neural Inf. Process. Syst. (NeurIPS), 2021. [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles