A Hybrid Convolutional Neural Network-Vision Transformer Framework for Gujarati Scene Text Recognition in Assistive Applications
Authors
Department of Computer Science, Indus University, Ahmedabad, Gujarat (India)
G H Patel Post Graduate Department of Computer Science and Technology, Sardar Patel University, Vallabh Vidyanagar, Gujarat (India)
Article Information
DOI: 10.51244/IJRSI.2026.1307000260
Subject Category: Computer Science
Volume/Issue: 13/7 | Page No: 3569-3578
Publication Timeline
Submitted: 2026-07-24
Accepted: 2026-07-30
Published: 2026-08-12
Abstract
For visual impairments it is very difficult to read and comprehend textual information included in instructional materials, notice boards, announcements and visual representations of general information. Text-to-speech (TTS) and optical character recognition (OCR) systems are helping to make data more accessible, but it is challenging to reliably recognise Gujarati text in practical settings due to the complexity of real-world conditions like varying light levels, text background strength, font styles and character arrangement. Thus, two wearable device access projects urgently require a high-level support framework with technological skills that can distinguish Gujarati speech and provide information. Image capture, pre-processing, adaptive text segmentation, sequence modelling with BiLSTM with attention, feature extraction using a hybrid Convolutional Neural Network Vision Transformer (CNNViT) and optimisation directed by Bayesian Optimisation Reinforcement Learning (BORL) are all included in the suggested framework. While CNNs recover local Gujarati character-level representations, the Vision Transformer records global contextual associations. The second function entails identifying Gujarati text in speech and translating it into a spoken language that users with visual impairments can understand. Because speech translation is asymmetric, the use of BiLSTM and attention mechanisms improves sequential recognition performance.
Keywords
Visual impairments, Gujarati, Text Segmentation, Convolutional Neural Network
Downloads
References
1. Palarimath, S., Radhakrishnan, V., Veerasamy, D., Valsalan, P., Viswan, V. and Thumu, M.B., 2025, March. Artificial Intelligence and Machine Learning for Accessibility: Smart Reader as an Assistive Tool for the Visually Impaired. In 2025 International Conference on Machine Learning and Autonomous Systems (ICMLAS) (pp. 424-430). IEEE [Google Scholar] [Crossref]
2. Ikram, S., Bajwa, I.S., Gyawali, S., Ikram, A. and Alsubaie, N., 2025. Enhancing object detection in assistive technology for the visually impaired: A DETR-based approach. IEEE Access, 13, pp.71647-71661. [Google Scholar] [Crossref]
3. Souza, L.R.D., Francisco, R., Rosa Tavares, J.E.D. and Barbosa, J.L.V., 2025. Intelligent environments and assistive technologies for assisting visually impaired people: a systematic literature review. Universal Access in the Information Society, 24(4), pp.3021-3048. [Google Scholar] [Crossref]
4. Madaan, H., Bajaj, R., Khedar, R., Avya, V.M., Bajaj, A. and Sharma, M., 2025, September. AI-Powered Smart Glasses for Social and Spatial Awareness in the Visually Impaired. In 2025 12th International Conference on Reliability, Infocom Technologies and Optimisation (Trends and Future Directions) (ICRITO) (pp. 1-6). IEEE. [Google Scholar] [Crossref]
5. Acharya, S., Bhattacharjee, S., Das, A.K. and Ghosh, S., 2026, February. On-Device Visual Text Acquisition and Speech Synthesis for Assistive Applications Using Raspberry PI. In 2026 9th International Conference on Electronics, Materials Engineering & Nano-Technology (IEMENTech) (Vol. 9, pp. 1-6). IEEE. [Google Scholar] [Crossref]
6. Anchev, A., Savova, M. and Ivanova, G., 2025, November. Integrating Machine Learning into Assistive OCR Technology for People with Visual Impairments. In 2025, the 10th International Conference on Energy Efficiency and Agricultural Engineering (EE&AE) (pp. 1-6). IEEE. [Google Scholar] [Crossref]
7. Hegde, A.P. and Dharani, A., 2025, November. AI-Powered OCR-Based Translator with Speech Output for Partially Visual Impaired Users. In 2025, the 9th International Conference on Computational Systems and Information Technology for Sustainable Solutions (CSITSS) (pp. 1-6). IEEE. [Google Scholar] [Crossref]
8. Olawade, D.B., Bolarinwa, O.A., Adebisi, Y.A. and Shongwe, S., 2025. The role of artificial intelligence in enhancing healthcare for people with disabilities—Social Science & Medicine, 364, p.117560. [Google Scholar] [Crossref]
9. Gupta, A.M. and Sharma, H., 2025, January. Text Recognition in Indic Scripts using Deep Learning-A Comprehensive Analysis. In 2025 International Conference on Intelligent Systems and Computational Networks (ICISCN) (pp. 1-6). IEEE. [Google Scholar] [Crossref]
10. Limbachiya, K. and Sharma, A., 2025. A Two-Level Multi-Branch Convolutional Neural Network Framework for Handwritten Gujarati Character Recognition. IEEE Access, 13, pp.205733-205752. [Google Scholar] [Crossref]
11. Patel, D., Maniar, H. and Patel, J., 2025, December. CRNN with BiLSTM-CTC for Robust Recognition of Gujarati Handwritten Words and Lines. In International Conference on Soft Computing and its Engineering Applications (pp. 317-332). Cham: Springer Nature Switzerland. [Google Scholar] [Crossref]
12. Pragada, P. and CH, D.R., 2025. Design of an Iterative Method for Telugu Language OCR Leveraging Transformer-Based Models and Adaptive Thresholding. SN Computer Science, 6(4), p.346. [Google Scholar] [Crossref]
13. Anakpluek, N., Pasanta, W., Chantharasukha, L., Chokratansombat, P., Kanjanakaew, P. and Siriborvornratanakul, T., 2025. Improved Tesseract optical character recognition performance on Thai document datasets. Big Data Research, 39, p.100508. [Google Scholar] [Crossref]
14. Ingavale, M. and Patil, J.K., 2025, March. Assessment of CNN Models for Optical Character Recognition of Modi Lipi Script. In 2025 3rd International Conference on Smart Systems for applications in Electrical Sciences (ICSSES) (pp. 1-6). IEEE. [Google Scholar] [Crossref]
15. Ghai, D., Saxena, S., Dhingra, G. and Tripathi, S.L., 2025. A comprehensive review on performance-based comparative analysis, categorization, classification and mapping of text extraction system techniques for images. Multimedia Tools and Applications, 84(5), pp.2327-2484. [Google Scholar] [Crossref]
16. Nguyen, T.N., Burie, J.C., Le, T.L. and Schweyer, A.V., 2025. Text line segmentation approach combining deep learning model and traditional image processing techniques-application to transliteration of Cham manuscripts. Multimedia Tools and Applications, 84(32), pp.39143-39169. [Google Scholar] [Crossref]
17. Pawar, B.S., Patil, C.H., Jabde, M.K. and Mali, S., 2025, December. Morphological and Contour-Based Handwritten English Word Segmentation on IAM Dataset. In International Conference on Soft Computing and its Engineering Applications (pp. 489-499). Cham: Springer Nature Switzerland. [Google Scholar] [Crossref]
18. Goel, N. and Kaushik, A., 2025, July. Prescription-to-Text Conversion Using Ensemble Faster R-CNN with Optical Character Recognition. In 2025 6th International Conference on Data Intelligence and Cognitive Informatics (ICDICI) (pp. 455-460). IEEE. [Google Scholar] [Crossref]
19. Janani, M., Laya, H. and Mathishree, B., 2025, March. OCR Text Recognition and Recommendation Using Machine Learning. In 2025 International Conference on Visual Analytics and Data Visualization (ICVADV) (pp. 804-807). IEEE. [Google Scholar] [Crossref]
20. Karimov, N., Isakov, R., Ismailov, T., Madraimov, A., Bozorbekov, A. and Jumaniyozov, Y., 2025, April. Handwritten Text Recognition in Ancient Manuscripts Using Convolutional Neural Networks (CNN). In 2025 International Conference on Computational Innovations and Engineering Sustainability (ICCIES) (pp. 1-6). IEEE. [Google Scholar] [Crossref]
21. Liu, C., Jiang, Q., Peng, D., Kong, Y., Zhang, J., Xiong, L., Duan, J., Sun, C. and Jin, L., 2025. QT-TextSR: Enhancing scene text image super-resolution via efficient interaction with text recognition using a Query-aware Transformer. Neurocomputing, 620, p.129241. [Google Scholar] [Crossref]
22. Da, C., Wang, P. and Yao, C., 2026. Multi-granularity prediction with learnable fusion for scene text recognition. International Journal of Computer Vision, 134(1), p.47. [Google Scholar] [Crossref]
23. Al Kharusi, R.M., Al Abri, A.R.M., Al Husaini, Y.N. and Al Husaini, M.A., 2026, April. A Hybrid CNN–Vision Transformer Architecture for Enhanced Document Understanding in Higher Education Archiving. In 2026 International Conference on Artificial Intelligence, Systems and Emerging Technologies (ICAISET) (pp. 1-7). IEEE. [Google Scholar] [Crossref]
24. Khattab, A., Elpeltagy, M., Youness, F., Elshafei, A. and Khedr, A.Y., 2026. Integrating CNN and BiLSTM for enhanced scene text recognition. Neural Computing and Applications, 38(7), p.235. [Google Scholar] [Crossref]
25. Rakshitha, G., Rozana, I.M., Patil, R.B. and Shrivastav, P., 2025, May. Efficient Handwritten Text Recognition Using Residual Networks and BiLSTM. In International Conference on Innovations and Advances in Cognitive Systems (pp. 47-59). Cham: Springer Nature Switzerland. [Google Scholar] [Crossref]
26. Patel, P., Vegda, D.C. and Vyas, R.A., 2025, June. Gujarati Speech to English Text Conversion using attention based Sequence modeling. In 2025 International Conference on Electronics, AI and Computing (EAIC) (pp. 1-7). IEEE. [Google Scholar] [Crossref]
27. Aparna, C. and Rajchandar, K., 2026. Improving Cursive Handwriting Recognition: A Dual Model Approach Using HOG Features and VGG-19. National Academy Science Letters, pp.1-4. [Google Scholar] [Crossref]
28. Xu, C., Liang, J. and Ling, X., 2025, March. A Machine Learning-Based Calligraphy Font Recognition System Using HOG Features and Support Vector Machine. In 2025 11th International Conference on Computing and Artificial Intelligence (ICCAI) (pp. 175-180). IEEE. [Google Scholar] [Crossref]
29. Reddy, A.V., Kumar, S.P. and Pati, P.B., 2026, March. Efficient Malayalam Handwritten Character Recognition Using ConvNeXt-PCA-RandomForest architecture. In 2026 IEEE International Conference for Convergence in Computing Technology (I3CTCON) (pp. 1-5). IEEE. [Google Scholar] [Crossref]
30. Patel, C.C. and Patel, P., 2025, September. Gujarati Sign Language Character Recognition Using Neural Network and Multi-class Support Vector Machine Classifiers. In International Conference on Artificial Intelligence and Computing (pp. 381-400). Singapore: Springer Nature Singapore. [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- What the Desert Fathers Teach Data Scientists: Ancient Ascetic Principles for Ethical Machine-Learning Practice
- Comparative Analysis of Some Machine Learning Algorithms for the Classification of Ransomware
- Comparative Performance Analysis of Some Priority Queue Variants in Dijkstra’s Algorithm
- Transfer Learning in Detecting E-Assessment Malpractice from a Proctored Video Recordings.
- Dual-Modal Detection of Parkinson’s Disease: A Clinical Framework and Deep Learning Approach Using NeuroParkNet