A Comprehensive Review of Text Information Extraction Techniques for Images and Videos

Authors

Dr. Kapil Kumar

Assistant Professor, School of Computer Applications & Engineering, Vivek University, Bijnor, Uttar Pradesh (India)

Article Information

DOI: 10.51584/IJRIAS.2026.11060276

Subject Category: Computer Science

Volume/Issue: 11/6 | Page No: 3665-3673

Publication Timeline

Submitted: 2026-07-02

Accepted: 2026-07-07

Published: 2026-07-17

Abstract

Text information present in pictures and video contain helpful data for programmed explanation, ordering, and organizing of pictures. Extraction of this data includes discovery, restriction, following, extraction, improvement, and acknowledgment of the content from a given picture. Nonetheless, varieties of text because of contrasts in size, style, direction, and arrangement, just as low picture difference and complex foundation make the issue of programmed text extraction very testing. While thorough overviews of related issues, for example, face location, record investigation, and picture and video ordering can be discovered, the issue of text data extraction isn't very much studied. Countless methods have been proposed to address this issue, and the motivation behind this paper is to order and audit these calculations, talk about benchmark information and execution assessment, and to call attention to promising bearings for future exploration.

Keywords

Text data extraction, text location, text restriction, text following, text upgrade, OCR

Downloads

References

1. H. K. Kim, Efficient Automatic Text Location Method and Content-Based Indexing and Structuring of Video Database, Journal of Visual Communication and Image Representation 7 (4) (1996) 336-344. [Google Scholar] [Crossref]

2. H. J. Zhang, Y. Gong, S. W. Smoliar, and S. Y. Tan, Automatic Parsing of News Video, Proc. of IEEE Conference on Multimedia Computing and Systems, 1994, [Google Scholar] [Crossref]

3. Y. Cui and Q. Huang, Character Extraction of License Plates from Video, Proc. of IEEE Conference on Computer Vision and Pattern Recognition, 1997, pp. 502 –507. [Google Scholar] [Crossref]

4. T. Sato, T. Kanade, E. K. Hughes, and M. A. Smith, Video OCR for Digital News Archive, Proc. of IEEE Workshop on Content based Access of Image and Video Databases, 1998, pp. 52-60. [Google Scholar] [Crossref]

5. A. K. Jain, and Y. Zhong, Page Segmentation using Texture Analysis, Pattern Recognition, 29 (5) (1996) 743-770. [Google Scholar] [Crossref]

6. Y. Y. Tang, S. W. Lee, and C. Y. Suen, Automatic Document Processing: A Survey, Pattern Recognition, 29 (12) (1996) 1931-1952. [Google Scholar] [Crossref]

7. B. Yu, A. K. Jain, and M. Mohiuddin, Address Block Location on Complex Mail Pieces, Proc. of International Conference on Document Analysis and Recognition, 1997, pp. 897-901. [Google Scholar] [Crossref]

8. D. S. Kim and S. I. Chien, Automatic Car License Plate Extraction using Modified Generalized Symmetry Transform and Image Warping, Proc. of International Symposium on Industrial Electronics, 2001, Vol. 3, pp. 2022-2027. [Google Scholar] [Crossref]

9. J. C. Shim, C. Dorai, and R. Bolle, Automatic Text Extraction from Video for Content-based Annotation and Retrieval, Proc. of International Conference on Pattern Recognition, Vol. 1, 1998, pp. 618-620. [Google Scholar] [Crossref]

10. S. Antani, D. Crandall, A. Narasimhamurthy, V. Y. Mariano, and R. Kasturi, Evaluation of Methods for Detection and Localization of Text in Video, Proc. of the IAPR workshop on Document Analysis Systems, Rio de Janeiro, December 2000, pp. 506-514. [Google Scholar] [Crossref]

11. S. Antani, Reliable Extraction of Text From Video, PhD thesis, Pennsylvania State University, August 2001. [Google Scholar] [Crossref]

12. Z. Lu, Detection of Text Region from Digital Engineering Drawings, IEEE Transactions on Pattern Analysis and Machine Intelligence, 20 (1998) 431-439. [Google Scholar] [Crossref]

13. D. Crandall, S. Antani, and R. Kasturi, Robust Detection of Stylized Text Events in Digital Video, Proceedings of International Conference on Document Analysis and Recognition, 2001, pp. 865-869. [Google Scholar] [Crossref]

14. Yu Zhong, Hongjiang Zhang, and Anil K. Jain, Automatic Caption Localization in Compressed Video, IEEE Transactions on Pattern Analysis and Machine Intelligence, 22, (4) (2000) 385-392. [Google Scholar] [Crossref]

15. R. Lienhart and F. Stuber, Automatic Text Recognition In Digital Videos, Proc. of SPIE, 1996, pp. 180-188. [Google Scholar] [Crossref]

16. J. Ohya, A. Shio, and S. Akamatsu, Recognizing Characters in Scene Images, IEEE Transactions on Pattern Analysis and Machine Intelligence, 16 (2) (1994) 214-224. [Google Scholar] [Crossref]

17. S. H. Park, K. I. Kim, K. Jung, and H. J. Kim, Locating Car License Plates using Neural Networks, IEE Electronics Letters, 35 (17) (1999) 1475-1477. [Google Scholar] [Crossref]

18. A. K. Jain, and K. Karu, Learning Texture Discrimination Masks, IEEE Transactions on Pattern Analysis and Machine Intelligence, 18 (2) (1996) 195-205. [Google Scholar] [Crossref]

19. K. Jung, Neural network-based Text Location in Color Images, Pattern Recognition Letters, 22 (14) December (2001) 1503-1515. [Google Scholar] [Crossref]

20. B.L. Yeo, and B. Liu, Visual Content Highlighting via Automatic Extraction of Embedded Captions on MPEG Compressed Video, Proc. of SPIE, 1996, pp.142-149. [Google Scholar] [Crossref]

21. J. Zhou, D. Lopresti, and Z. Lei, OCR for World Wide Web Images, Proc. of SPIE on Document Recognition IV, Vol. 3027, 1997, pp. 58-66. [Google Scholar] [Crossref]

22. J. Zhou, D. Lopresti, and T. Tasdizen, Finding Text in Color Images, Proc. of SPIE on Document Recognition V, 1998, pp. 130-140. [Google Scholar] [Crossref]

23. Y. Watanabe, Y. Okada, Y. B. Kim, and T. Takeda, Translation Camera, Proc. of International Conference on Pattern Recognition, 1998, Vol. 1, pp. 613-617. [Google Scholar] [Crossref]

24. Haritaoglu, Scene Text Extraction and Translation for Handheld Devices, Proc. of IEEE Conference on Computer Vision and Pattern Recognition, 2001, Vol. 2, pp.408-413. [Google Scholar] [Crossref]

25. G. Feng, H. Cheng, and C. Bouman, High Quality MRC Document Coding,'' Proc. of International Conference on Image Processing, Image Quality, Image Capture Systems Conference (PICS), Montreal Canada, 2001. [Google Scholar] [Crossref]

26. H. Cheng, C. A. Bouman, and J. P. Allebach, Multiscale Document Segmentation, Proc. of IS&T 50th Annual Conference, 1997, pp. 417-425, Cambridge, MA. [Google Scholar] [Crossref]

27. B. Shahraray and D. C. Gibbon, Automatic Generation of Pictorial Transcripts of Video Programs, Proc. of SPIE, 1995, Vol. 2417. [Google Scholar] [Crossref]

28. S. Fisher, R. Lienhart, and W. Effelsberg, Automatic Recognition of Film Genres, Proc. of ACM Multimedia’95, 1995, pp. 295-304, San Francisco. [Google Scholar] [Crossref]

29. Y. K. Ham, M. S. Kang, H. K. Chung, and R. H. Park, Recognition of Raised Characters for Automatic Classification of Rubber Tires, Opt. Eng., 1995, Vol. 34, pp.102-108. [Google Scholar] [Crossref]

30. S. Kim, D. Kim, Y. Ryu, and G. Kim, A Robust License-Plate Extraction Method under Complex Image Conditions, Proc. of International Conference on Pattern Recognition, 2002, Vol. 3, pp. 216-219. [Google Scholar] [Crossref]

31. S. Long, X. He, and C. Yao, “Scene Text Detection and Recognition: The Deep Learning Era,” International Journal of Computer Vision, vol. 129, no. 1, pp. 161–184, 2021. [Google Scholar] [Crossref]

32. F. Borisyuk, A. Gordo, and V. Sivakumar, “Rosetta: Large Scale System for Text Detection and Recognition in Images,” Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2019. [Google Scholar] [Crossref]

33. J. Memon, M. Sami, R. A. Khan, and M. Uddin, “Handwritten Optical Character Recognition (OCR): A Comprehensive Systematic Literature Review,” IEEE Access, vol. 8, pp. 142642–142668, 2020. [Google Scholar] [Crossref]

34. A. Baek, G. Kim, J. Lee, S. Park, D. Han, S. Yun, S. J. Oh, and H. Lee, “What Is Wrong with Scene Text Recognition Model Comparisons? Dataset and Model Analysis,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019. [Google Scholar] [Crossref]

35. A. Halder, U. Pal, P. Shivakumara, and M. Blumenstein, “A Comprehensive Review on Text Detection and Recognition in Scene Images,” Artificial Intelligence and Applications, vol. 2, 2024. [Google Scholar] [Crossref]

36. U. Maria, A. Singh, and R. Verma, “Exploring Text Recognition, Segmentation and Detection in Natural Scene Images: A Review,” Journal of Digital Systems, vol. 19, no. 2, pp. 66–82, 2024. [Google Scholar] [Crossref]

37. E. Eli, “A Comprehensive Review of Non-Latin Natural Scene Text Detection and Recognition,” Engineering Applications of Artificial Intelligence, vol. 145, 2025. [Google Scholar] [Crossref]

38. M. B. Gohil and A. A. Desai, “A Review on Text Detection and Recognition in Images and Videos Using Machine Learning and Deep Learning Techniques,” International Journal of Engineering Research & Technology (IJERT), vol. 15, no. 1, 2026. [Google Scholar] [Crossref]

39. A. K. Kashyap, M. Upadhya, V. S. Panwar, et al., “Integrated Framework Utilizing Scene Text Detection and Recognition Techniques for Enhancing Point of Interest Extraction from Name Boards in All Indic Languages,” Scientific Reports, vol. 16, 2026. [Google Scholar] [Crossref]

40. V. Kadha and R. Kumar, “A Deep Learning Survey of Scene Text Detection and Recognition,” Computer Vision and Image Understanding, 2026 [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles