A Novel Recurrent Convolutional Neural Network Framework for Continuous Sign Language Recognition Using Iterative Training and Multimodal Fusion
Authors
R K College of Engineering (A), Kethanakonda (V), Ibrahimpatnam (M), Vijayawada, AMARAVATI – 521 456, Andhra Pradesh, INDIA (India)
R K College of Engineering (A), Kethanakonda (V), Ibrahimpatnam (M), Vijayawada, AMARAVATI – 521 456, Andhra Pradesh, INDIA (India)
R K College of Engineering (A), Kethanakonda (V), Ibrahimpatnam (M), Vijayawada, AMARAVATI – 521 456, Andhra Pradesh, INDIA (India)
Article Information
DOI: 10.51244/IJRSI.2026.1305000127
Subject Category: Artificial Intelligence
Volume/Issue: 13/5 | Page No: 1380-1391
Publication Timeline
Submitted: 2026-05-09
Accepted: 2026-05-14
Published: 2026-06-03
Abstract
In this article, we present a novel approach to continuous Sign Language (SL) recognition using a Recurrent Convolutional Neural Network (RCNN) with an iterative training process and multimodal fusion. Our primary goal is to accurately transcribe continuous SL video streams into ordered gloss sequences, overcoming the limitations of traditional methods that rely on frame-wise labeling and Hidden Markov Models (HMMs). To address the challenges posed by limited training data, we introduce an iterative optimization process that refines gestural alignments, ensuring improved model performance across training iterations. Additionally, we incorporate a multimodal fusion strategy that combines RGB frames and optical flow data to capture both appearance and motion cues, enhancing the spatiotemporal feature representation. The experimental results demonstrate that our approach outperforms existing SL recognition methods in terms of recognition accuracy and Word Error Rate (WER), showing significant potential for real-world applications such as real-time SL translation and human-computer interaction. Our system achieves robust performance even with unsegmented video streams, making it a promising solution for continuous SL recognition tasks.
Keywords
Continuous Sign Language Recognition, Recurrent Convolutional Neural Networks (RCNN), Iterative Training, Multimodal Fusion, Word Error Rate (WER).
Downloads
References
1. S. C. Ong and S. Ranganath, “Automatic sign language analysis: A survey and the future beyond lexical meaning,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 27, no. 6, pp. 873–891, 2005. [Google Scholar] [Crossref]
2. Z. Ren, J. Yuan, J. Meng, and Z. Zhang, “Robust part-based hand gesture recognition using Kinect sensor,” IEEE Trans. Multimedia, vol. 15, no. 5, pp. 1110–1120, 2013. [Google Scholar] [Crossref]
3. C. Wang, Z. Liu, and S.-C. Chan, “Superpixel-based hand gesture recognition with Kinect depth camera,” IEEE Trans. Multimedia, vol. 17, no. 1, pp. 29–39, 2015. [Google Scholar] [Crossref]
4. P. Molchanov, X. Yang, S. Gupta, K. Kim, S. Tyree, and J. Kautz, “Online detection and classification of dynamic hand gestures with recurrent 3D convolutional neural network,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog., 2016, pp. 4207–4215. [Google Scholar] [Crossref]
5. H. Cooper, E. J. Ong, N. Pugeault, and R. Bowden, “Sign language recognition using sub-units,”J. Mach. Learning Research, vol. 13, pp. 2205–2231, 2012. [Google Scholar] [Crossref]
6. P. Molchanov, S. Gupta, K. Kim, and J. Kautz, “Hand gesture recognition with 3D convolutional neural networks,” in IEEE Conf. Comput. Vis. Pattern Recog. Workshops, 2015, pp. 1–7. [Google Scholar] [Crossref]
7. G. Zen, L. Porzi, E. Sangineto, E. Ricci, and N. Sebe, “Learning personalized models for facial expression analysis and gesture recognition,” IEEE Trans. Multimedia, vol. 18, no. 4, pp. 775–788, 2016. [Google Scholar] [Crossref]
8. G. D. Evangelidis, G. Singh, and R. Horaud, “Continuous gesture recognition from articulated poses,” in Eur. Conf. Comput. Vis. Workshops, 2014, pp. 595–607. [Google Scholar] [Crossref]
9. N. Neverova, C. Wolf, G. Taylor, and F. Nebout, “Multi-scale deep learning for gesture detection and localization,” in Eur. Conf. Comput. Vis. Workshops, 2014, pp. 474–490. [Google Scholar] [Crossref]
10. D. Wu, L. Pigou, P.-J. Kindermans, N. Le, L. Shao, J. Dambre, and J.-M. Odobez, “Deep dynamic neural networks for multimodal gesture segmentation and recognition,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 38, no. 8, pp. 1583–1597, 2016. [Google Scholar] [Crossref]
11. O. Koller, J. Forster, and H. Ney, “Continuous sign language recognition: Towards large vocabulary statistical recognition systems handling multiple signers,” Comput. Vis. Image Understand., vol. 141, pp. 108–125, 2015. [Google Scholar] [Crossref]
12. O. Koller, H. Ney, and R. Bowden, “Deep hand: how to train a CNN on 1 million hand images when your data is continuous and weakly labelled,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog., 2016, pp. 3793–3802. [Google Scholar] [Crossref]
13. O. Koller, S. Zargaran, H. Ney, and R. Bowden, “Deep sign: hybrid CNN-HMM for continuous sign language recognition,” in Proc. Brit. Mach. Vis. Conf., 2016. [Google Scholar] [Crossref]
14. U. Von Agris, M. Knorr, and K.-F. Kraiss, “The significance of facial features for automatic sign language recognition,” in 8th IEEE Int. Conf. Autom. Face Gesture Recog., 2008, pp. 1–6. [Google Scholar] [Crossref]
15. P. Buehler, A. Zisserman, and M. Everingham, “Learning sign language by watching TV (using weakly aligned subtitles),” in Proc. IEEE Conf. Comput. Vis. Pattern Recog., 2009, pp. 2961–2968. [Google Scholar] [Crossref]
16. L.-C. Wang, R. Wang, D.-H. Kong, and B.-C. Yin, “Similarity assessment model for Chinese sign language videos,” IEEE Trans. Multimedia, vol. 16, no. 3, pp. 751–761, 2014. [Google Scholar] [Crossref]
17. C. Monnier, S. German, and A. Ost, “A multi-scale boosted detector for efficient and robust gesture recognition,” in Eur. Conf. Comput. Vis. Workshops, 2014, pp. 491–502. [Google Scholar] [Crossref]
18. J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog., 2015, pp. 2625–2634. [Google Scholar] [Crossref]
19. L. Wang, Y. Xiong, Z. Wang, and Y. Qiao, “Towards good practices for very deep two-stream convnets,” arXiv preprint arXiv:1507.02159, 2015. [Google Scholar] [Crossref]
20. Y. Shi, Y. Tian, Y. Wang, and T. Huang, “Sequential deep trajectory descriptor for action recognition with three-stream CNN,” IEEE Trans. Multimedia, vol. 19, no. 7, pp. 1510–1520, 2017. [Google Scholar] [Crossref]
21. L. Pigou, A. v. d. Oord, S. Dieleman, M. M. Van Herreweghe, and J. Dambre, “Beyond temporal pooling: Recurrence and temporal convolutions for gesture recognition in video,” Int. J. Comput. Vis., pp. 1–10, 2015. [Google Scholar] [Crossref]
22. O. Koller, S. Zargaran, and H. Ney, “Re-sign: Re-aligned end-to-end sequence modelling with deep recurrent CNN-HMMs,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog., 2017. [Google Scholar] [Crossref]
23. T. Pfister, J. Charles, and A. Zisserman, “Domain-adaptive discriminative one-shot learning of gestures,” in Proc. Eur. Conf. Comput. Vis., 2014, pp. 814–829. [Google Scholar] [Crossref]
24. D. Wu and L. Shao, “Leveraging hierarchical parametric networks for skeletal joints based action segmentation and recognition,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog., 2014, pp. 724–731. [Google Scholar] [Crossref]
25. Kondragunta Rama Krishnaiah and H. Harish. “Performance Characterization and Comparative Study of DSR and OLSR Routing Protocols in Mobile Ad Hoc Networks.” International Journal on Recent and Innovation Trends in Computing and Communication 11, no. 11 (December 2023): 1842–1846. [Google Scholar] [Crossref]
26. Kondragunta Rama Krishnaiah and H. Harish. “Understanding Emotions with Deep Learning: A Multimodal Approach for Detecting Speech and Facial Expressions.” Journal of Computational Analysis and Applications 32, no. 1 (2024): 661–673. https://doi.org/10.48047/jocaaa.2024.32.01.21 [Google Scholar] [Crossref]
27. Kondragunta Rama Krishnaiah, H. Harish, and Manjunath B. E. “Character-Level Convolutional Neural Networks for Cyberbullying Detection: A Robust Approach to Handling Noisy Social Media Text.” International Journal of Intelligent Systems and Applications in Engineering 12, no. 14s (February 2024): 746–752. [Google Scholar] [Crossref]
28. Kondragunta Rama Krishnaiah and H. Harish. “Image-Based Real Estate Appraisal: Leveraging Mask R-CNN for Damage Detection and Severity Estimation.” International Journal of Intelligent Systems and Applications in Engineering 12, no. 21s (March 2024): 4961–4969. [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- The Role of Artificial Intelligence in Revolutionizing Library Services in Nairobi: Ethical Implications and Future Trends in User Interaction
- ESPYREAL: A Mobile Based Multi-Currency Identifier for Visually Impaired Individuals Using Convolutional Neural Network
- Comparative Analysis of AI-Driven IoT-Based Smart Agriculture Platforms with Blockchain-Enabled Marketplaces
- AI-Based Dish Recommender System for Reducing Fruit Waste through Spoilage Detection and Ripeness Assessment
- SEA-TALK: An AI-Powered Voice Translator and Southeast Asian Dialects Recognition