Handwritten Word Recognition for Low-Resource Languages: A CRNN-CTC Framework for Kirundi
Authors
Anhui University of Technology, Ma'anshan, Anhui 243002 (China)
Article Information
DOI: 10.51584/IJRIAS.2026.11070019
Subject Category: Computer Science
Volume/Issue: 11/7 | Page No: 423-441
Publication Timeline
Submitted: 2026-06-26
Accepted: 2026-07-01
Published: 2026-07-28
Abstract
Handwritten Text Recognition (HTR) has experienced remarkable progress with the development of deep learning techniques. However, most existing studies focus on high-resource languages for which large annotated datasets are readily available. In contrast, low-resource languages remain largely underrepresented in handwriting recognition research due to the scarcity of handwritten corpora, linguistic resources, and benchmark datasets. This paper presents a handwritten word recognition framework for Kirundi, a low-resource Bantu language spoken primarily in Burundi. The proposed system employs a Convolutional Recurrent Neural Network (CRNN) combined with Connectionist Temporal Classification (CTC) for end-to-end sequence recognition without explicit character segmentation.
To address severe data scarcity, a small handwritten Kirundi dataset consisting of manually collected word samples was constructed and annotated. Data augmentation techniques, including rotation, translation, Gaussian noise, Gaussian blur, and elastic distortion, were applied to increase sample diversity. In addition, synthetic handwritten-style data were generated to further expand the training set. Three experimental configurations were investigated: real handwritten data only, real data with augmentation, and real data combined with augmentation and synthetic handwritten-style images.
Experimental results demonstrate that synthetic data generation improved recognition performance and reduced Character Error Rate (CER) from 0.8354 to 0.7560, corresponding to an approximate relative improvement of 9.5%. Although exact word-level recognition remained difficult because of the extremely limited dataset size, the proposed framework successfully learned meaningful sequential patterns and produced increasingly structured Kirundi-like predictions. The study establishes an initial benchmark for Kirundi handwritten word recognition and highlights the potential of synthetic data generation for low-resource handwriting recognition tasks.
Keywords
Handwritten Word Recognition; Low-Resource Languages; Deep Learning; CRNN; CTC
Downloads
References
1. Bluche, T. (2016). Deep Neural Networks for Large Vocabulary Handwritten Text Recognition. PhD Dissertation, Université Paris-Saclay. [Google Scholar] [Crossref]
2. Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep Learning. Cambridge, MA: MIT Press. [Google Scholar] [Crossref]
3. Graves, A. (2012). Supervised Sequence Labelling with Recurrent Neural Networks. Berlin, Germany: Springer. [Google Scholar] [Crossref]
4. Graves, A., Fernández, S., Gomez, F., & Schmidhuber, J. (2006). Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks. Proceedings of the 23rd International Conference on Machine Learning (ICML), 369–376.https://doi.org/10.1145/1143844.1143891 [Google Scholar] [Crossref]
5. Graves, A., Liwicki, M., Fernández, S., Bertolami, R., Bunke, H., & Schmidhuber, J. (2009). A Novel Connectionist System for Improved Unconstrained Handwriting Recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 31(5), 855–868.DOI: https://doi.org/10.1109/TPAMI.2008.137 [Google Scholar] [Crossref]
6. He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep Residual Learning for Image Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 770–778. [Google Scholar] [Crossref]
7. Hochreiter, S., & Schmidhuber, J. (1997). Long Short-Term Memory. Neural Computation, 9(8), 1735–1780.https://doi.org/10.1162/neco.1997.9.8.1735 [Google Scholar] [Crossref]
8. Kingma, D. P., & Ba, J. (2015). Adam: A Method for Stochastic Optimization. Proceedings of the 3rd International Conference on Learning Representations (ICLR). [Google Scholar] [Crossref]
9. LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep Learning. Nature, 521(7553), 436–444.https://doi.org/10.1038/nature14539 [Google Scholar] [Crossref]
10. Li, M., Lv, T., Cui, J., Xie, L., Zhuang, Y., & Xu, B. (2023). TrOCR: Transformer-Based Optical Character Recognition with Pre-Trained Models. Proceedings of the AAAI Conference on Artificial Intelligence, 37(2), 13094–13102. [Google Scholar] [Crossref]
11. Liwicki, M., & Bunke, H. (2005). IAM-OnDB—An On-Line English Sentence Database Acquired from Handwritten Text on a Whiteboard. Proceedings of the 8th International Conference on Document Analysis and Recognition (ICDAR), 956–961. [Google Scholar] [Crossref]
12. Marti, U.-V., & Bunke, H. (2002). The IAM Database: An English Sentence Database for Offline Handwriting Recognition. International Journal on Document Analysis and Recognition, 5(1), 39–46.https://doi.org/10.1007/s100320200071 [Google Scholar] [Crossref]
13. Michael, F., Antonacopoulos, A., & Gatos, B. (2019). ICDAR 2019 Competition on Handwritten Text Recognition on Historical Documents. Proceedings of the International Conference on Document Analysis and Recognition (ICDAR). [Google Scholar] [Crossref]
14. Mori, S., Suen, C. Y., & Yamamoto, K. (1992). Historical Review of OCR Research and Development. Proceedings of the IEEE, 80(7), 1029–1058.DOI: https://doi.org/10.1109/5.156468 [Google Scholar] [Crossref]
15. Paszke, A., Gross, S., Massa, F., et al. (2019). PyTorch: An Imperative Style, High-Performance Deep Learning Library. Advances in Neural Information Processing Systems, 32. [Google Scholar] [Crossref]
16. Shi, B., Bai, X., & Yao, C. (2017). An End-to-End Trainable Neural Network for Image-Based Sequence Recognition and Its Application to Scene Text Recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(11), 2298–2304.https://doi.org/10.1109/TPAMI.2016.2646371 [Google Scholar] [Crossref]
17. Simard, P. Y., Steinkraus, D., & Platt, J. C. (2003). Best Practices for Convolutional Neural Networks Applied to Visual Document Analysis. Proceedings of the 7th International Conference on Document Analysis and Recognition (ICDAR), 958–963. [Google Scholar] [Crossref]
18. Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). Attention Is All You Need. Advances in Neural Information Processing Systems, 30. [Google Scholar] [Crossref]
19. Wigington, C., Tensmeyer, C., Davis, B., Barrett, W., Price, B., & Cohen, S. (2018). Start, Follow, Read: End-to-End Full-Page Handwriting Recognition. Proceedings of the European Conference on Computer Vision (ECCV), 367–383.DOI: https://doi.org/10.1007/978-3-030-01231-1_23 [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- What the Desert Fathers Teach Data Scientists: Ancient Ascetic Principles for Ethical Machine-Learning Practice
- Comparative Analysis of Some Machine Learning Algorithms for the Classification of Ransomware
- Comparative Performance Analysis of Some Priority Queue Variants in Dijkstra’s Algorithm
- Transfer Learning in Detecting E-Assessment Malpractice from a Proctored Video Recordings.
- Dual-Modal Detection of Parkinson’s Disease: A Clinical Framework and Deep Learning Approach Using NeuroParkNet