Emotion Recognition Using Machine Learning and Computer Vision: A Hybrid CNN–ViT Multimodal Framework

Authors

Deepa Chandrashekhar Rathod

Department of Computer Science and Engineering, Jain Institute of Technology, Davangere (India)

Bhumika B K

Department of Computer Science and Engineering, Jain Institute of Technology, Davangere (India)

Arpitha G A

Department of Computer Science and Engineering, Jain Institute of Technology, Davangere (India)

Brunda U Jajur

Department of Computer Science and Engineering, Jain Institute of Technology, Davangere (India)

Usha K

Department of Computer Science and Engineering, Jain Institute of Technology, Davangere (India)

Article Information

DOI: 10.51244/IJRSI.2026.1304000249

Subject Category: Computer Science

Volume/Issue: 13/4 | Page No: 2922-2933

Publication Timeline

Submitted: 2026-04-22

Accepted: 2026-04-27

Published: 2026-05-19

Abstract

Automatic recognition of human emotions from facial expressions and multimodal signals constitutes a foundational challenge in affective computing and human–computer interaction, with broad applications spanning healthcare monitoring, autonomous vehicle safety, educational technology, and social robotics. Despite remarkable progress driven by deep learning, particularly convolutional neural networks (CNNs), recurrent neural networks (RNNs), and Vision Transformers (ViT), achieving robust emotion recognition in unconstrained, real-world environments remains an open problem. This paper presents a comprehensive synthesis of over twenty-five state-of-the-art studies on facial and multimodal emotion recognition, encompassing CNN-based systems trained on FER2013, CK+, RAF-DB, and AffectNet; transformer-based hybrid architectures; and multimodal fusion systems integrating facial, speech, and electroencephalography (EEG) cues evaluated on RAVDESS, IEMOCAP, CMU-MOSEI, eNTERFACE'05, and MAHNOB-HCI. Building upon these insights, this work proposes a novel Hybrid CNN–ViT Multimodal Emotion Recognition (HCV-MER) framework comprising: (i) a squeeze-and-excitation ResNet combined with a Vision Transformer facial backbone incorporating region-specific attention over eyes and mouth; (ii) a lightweight temporal aggregation unit for video-level inference; and (iii) a cross-modal attention fusion module integrating facial and speech streams. Experimental evaluations target FER2013, RAF-DB, CK+, and RAVDESS using TensorFlow, PyTorch, and OpenCV. Expected improvements over baseline CNN architectures range from five to ten percentage points on challenging in-the-wild benchmarks. The paper further analyzes unresolved challenges including cross-domain generalization, demographic fairness, micro-expression recognition, and privacy-preserving deployment.

Keywords

Facial emotion recognition; Convolutional neural networks

Downloads

References

1. R. Pereira et al., "Systematic Review of Emotion Detection with Computer Vision and Deep Learning," Sensors, vol. 24, 2024. [Google Scholar] [Crossref]

2. M. Karnati et al., "Understanding Deep Learning Techniques for Recognition of Human Emotions Using Facial Expressions: A Comprehensive Survey," IEEE Transactions on Instrumentation and Measurement, vol. 72, pp. 1–28, 2023. [Google Scholar] [Crossref]

3. E. S. Agung et al., "Image-based facial emotion recognition using convolutional neural network on Emognition dataset," Scientific Reports, vol. 14, 2024. [Google Scholar] [Crossref]

4. S. Taru et al., "Emotion Recognition System from Facial Expressions Using Machine Learning," Journal of Artificial Intelligence and Capsule Networks, vol. 7, 2025. [Google Scholar] [Crossref]

5. C. L. Jiménez et al., "Multimodal Emotion Recognition on RAVDESS Dataset Using Transfer Learning," Sensors, vol. 21, no. 22, 2021. [Google Scholar] [Crossref]

6. Cirneanu et al., "New Trends in Emotion Recognition Using Image Analysis by Neural Networks, A Systematic Review," Sensors, vol. 23, 2023. [Google Scholar] [Crossref]

7. M. Ranjani et al., "Emotion Recognition Using CNN," in Proc. ICOEI, 2025. [Google Scholar] [Crossref]

8. D. Canedo and A. J. R. Neves, "Facial Expression Recognition Using Computer Vision: A Systematic Review," Applied Sciences, vol. 9, no. 21, p. 4678, 2019. [Google Scholar] [Crossref]

9. G. Juniawan et al., "Real-Time Facial Emotion Detection Application with Image Processing Based on CNN," International Journal of Electrical Engineering, Mathematics and Computer Science, vol. 3, no. 1, 2024. [Google Scholar] [Crossref]

10. D. Mamieva et al., "Multimodal Emotion Detection via Attention-Based Fusion of Extracted Facial and Speech Features," Sensors, vol. 23, 2023. [Google Scholar] [Crossref]

11. Z. Huang et al., "A study on computer vision for facial emotion recognition," Scientific Reports, vol. 13, 2023. [Google Scholar] [Crossref]

12. P. S. Tomar et al., "Fusing facial and speech cues for enhanced multimodal emotion recognition," International Journal of Information Technology, vol. 16, 2024. [Google Scholar] [Crossref]

13. D. Shukla et al., "Human Face Detection and Emotion Recognition Using OpenCV through AI," in Proc. IEMECON, 2024. [Google Scholar] [Crossref]

14. Yushchenko et al., "Evaluating CNN, RNN, and Vision Transformer for Emotion Recognition: Strengths and Weaknesses," in Proc. IEEE eStream, 2025. [Google Scholar] [Crossref]

15. R. Raj and I. Demirkol, "An improved facial emotion recognition system using CNN for optimization of human robot interaction," Scientific Reports, vol. 15, 2025. [Google Scholar] [Crossref]

16. B. C. Ko, "A Brief Review of Facial Emotion Recognition Based on Visual Information," Sensors, vol. 18, no. 2, p. 401, 2018. [Google Scholar] [Crossref]

17. B. K. Jha et al., "Human Emotion Detection and Face Recognition System," International Journal on Engineering Technology, vol. 6, 2025. [Google Scholar] [Crossref]

18. S. E. A. Hassan et al., "Survey on Emotion Recognition Using Deep Learning," in Proc. ICEEM, 2025. [Google Scholar] [Crossref]

19. J.-Y. Pan et al., "Multimodal Emotion Recognition Based on Facial Expressions, Speech, and EEG," IEEE Open Journal of Engineering in Medicine and Biology, vol. 4, 2023. [Google Scholar] [Crossref]

20. Chaudhari et al., "ViTFER: Facial Emotion Recognition with Vision Transformers," Applied System Innovation, vol. 5, no. 6, p. 117, 2022. [Google Scholar] [Crossref]

21. Dosovitskiy et al., "An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale," in Proc. ICLR, 2021. [Google Scholar] [Crossref]

22. J. Hu et al., "Squeeze-and-Excitation Networks," in Proc. IEEE CVPR, 2018. [Google Scholar] [Crossref]

23. K. He et al., "Deep Residual Learning for Image Recognition," in Proc. IEEE CVPR, 2016. [Google Scholar] [Crossref]

24. G. Qiuqiang et al., "PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition," IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 28, pp. 2880–2894, 2020. [Google Scholar] [Crossref]

25. R. Livingstone and F. Russo, "The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS)," PLOS ONE, vol. 13, no. 5, 2018. [Google Scholar] [Crossref]

26. P. Lucey et al., "The Extended Cohn-Kanade Dataset (CK+): A complete expression dataset for action unit and emotion-specified expression," in Proc. IEEE CVPRW, 2010. [Google Scholar] [Crossref]

27. S. Li et al., "Reliable Crowdsourcing and Deep Locality-Preserving Learning for Expression Recognition in the Wild," in Proc. IEEE CVPR, 2017. [Google Scholar] [Crossref]

28. Mollahosseini et al., "AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild," IEEE Transactions on Affective Computing, vol. 10, no. 1, pp. 18–31, 2019. [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles