Deep Learning-Based Face Detection, Feature Extraction, and Face Recognition from Video: A Comprehensive Review
Authors
Smt. Tanuben & Dr. ManubhaiTrivedi College of Information Science, Surat (India)
Article Information
DOI: 10.51244/IJRSI.2026.1307000275
Subject Category: Learning
Volume/Issue: 13/7 | Page No: 3763-3775
Publication Timeline
Submitted: 2026-07-29
Accepted: 2026-08-03
Published: 2026-08-14
Abstract
Face recognition has become one of the most prominent biometric technologies due to its extensive applications in surveillance, access control, authentication, human-computer interaction, and intelligent security systems. Recent advances in deep learning, particularly Convolutional Neural Networks (CNNs), have significantly improved the accuracy and robustness of face detection, feature extraction, and face recognition under challenging real-world conditions. This paper presents a comprehensive review of deep learning-based techniques employed across the complete face recognition pipeline. A comparative analysis of existing studies is presented to highlight the evolution of deep learning techniques and their effectiveness in improving recognition accuracy and computational efficiency. The review also discusses widely used benchmark datasets, performance evaluation metrics, and the major challenges encountered in unconstrained environments, such as pose variation, illumination changes, occlusion, facial expressions, aging, and low-resolution imagery. Finally, emerging research directions, including lightweight deep learning models, attention mechanisms, Vision Transformers, self-supervised learning, explainable artificial intelligence, and real-time video-based face recognition, are outlined to provide insights for future research. This review serves as a comprehensive reference for researchers and practitioners seeking a thorough understanding of recent developments and future trends in deep learning-based face recognition.
Keywords
Face Detection, Face Recognition, Feature Extraction, Deep Learning, Convolutional Neural Networks (CNNs), Video-Based Face Recognition, Biometrics, Computer Vision
Downloads
References
1. Zheng, J., Ranjan, R., Chen, C. H., Chen, J. C., Castillo, C. D., & Chellappa, R. (2020). An automatic system for unconstrained video-based face recognition. IEEE Transactions on Biometrics, Behavior, and Identity Science, 2(3), 194-209. [Google Scholar] [Crossref]
2. Wang, M., & Deng, W. (2021). Deep face recognition: A survey. Neurocomputing, 429, 215-244. [Google Scholar] [Crossref]
3. Guo, G., & Zhang, N. (2019). A survey on deep learning based face recognition. Computer vision and image understanding, 189, 102805. [Google Scholar] [Crossref]
4. Parkhi, O., Vedaldi, A., & Zisserman, A. (2015). Deep face recognition. In BMVC 2015-Proceedings of the British Machine Vision Conference 2015. British Machine Vision Association. [Google Scholar] [Crossref]
5. Schroff, F., Kalenichenko, D., & Philbin, J. (2015). Facenet: A unified embedding for face recognition and clustering. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 815-823). [Google Scholar] [Crossref]
6. He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770-778). [Google Scholar] [Crossref]
7. Deng, Jiankang, JiaGuo, NiannanXue, and Stefanos Zafeiriou. "Arcface: Additive angular margin loss for deep face recognition." In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 4690-4699. 2019. [Google Scholar] [Crossref]
8. Chen, S., Liu, Y., Gao, X., & Han, Z. (2018, August). Mobile Facenets: Efficient CNNs for accurate real-time face verification on mobile devices. In Chinese conference on biometric recognition (pp. 428-438). Cham: Springer International Publishing. [Google Scholar] [Crossref]
9. Taigman, Y., Yang, M., Ranzato, M. A., & Wolf, L. (2014). Deepface: Closing the gap to human-level performance in face verification. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1701-1708). [Google Scholar] [Crossref]
10. Yang, L., & Shami, A. (2020). On hyperparameter optimization of machine learning algorithms: Theory and practice. Neurocomputing, 415, 295-316. [Google Scholar] [Crossref]
11. Turk, M., & Pentland, A. (1991). Eigenfaces for recognition. Journal of cognitive neuroscience, 3(1), 71-86. [Google Scholar] [Crossref]
12. Viola, P., & Jones, M. (2001, December). Rapid object detection using a boosted cascade of simple features. In Proceedings of the 2001 IEEE computer society conference on computer vision and pattern recognition. CVPR 2001 (Vol. 1, pp. I-I). IEEE. [Google Scholar] [Crossref]
13. Ojala, T., Pietikäinen, M., & Harwood, D. (1996). A comparative study of texture measures with classification based on featured distributions. Pattern recognition, 29(1), 51-59. [Google Scholar] [Crossref]
14. Lowe, D. G. (2004). Distinctive image features from scale-invariant keypoints. International journal of computer vision, 60(2), 91-110. [Google Scholar] [Crossref]
15. Dalal, N., & Triggs, B. (2005, June). Histograms of oriented gradients for human detection. In 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR'05) (Vol. 1, pp. 886-893). IEEE. [Google Scholar] [Crossref]
16. Ranjan, R., Sankaranarayanan, S., Bansal, A., Bodla, N., Chen, J. C., Patel, V. M., & Chellappa, R. (2018). Deep learning for understanding faces: Machines may be just as good, or better, than humans. IEEE Signal Processing Magazine, 35(1), 66-83. [Google Scholar] [Crossref]
17. Zhang, K., Zhang, Z., Li, Z., &Qiao, Y. (2016). Joint face detection and alignment using multitask cascaded convolutional networks. IEEE signal processing letters, 23(10), 1499-1503. [Google Scholar] [Crossref]
18. Yang, S., Luo, P., Loy, C. C., & Tang, X. (2015). From facial parts responses to face detection: A deep learning approach. In Proceedings of the IEEE international conference on computer vision (pp. 3676-3684). [Google Scholar] [Crossref]
19. Yang, S., Luo, P., Loy, C. C., & Tang, X. (2016). Wider face: A face detection benchmark. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 5525-5533). [Google Scholar] [Crossref]
20. Zhu, C., Zheng, Y., Luu, K., & Savvides, M. (2017). CMS-RCNN: contextual multi-scale region-based cnn for unconstrained face detection. In Deep learning for biometrics (pp. 57-79). Cham: Springer International Publishing. [Google Scholar] [Crossref]
21. Zhang, S., Zhu, X., Lei, Z., Shi, H., Wang, X., & Li, S. Z. (2017). S3fd: Single shot scale-invariant face detector. In Proceedings of the IEEE international conference on computer vision (pp. 192-201). [Google Scholar] [Crossref]
22. Li, J., Wang, Y., Wang, C., Tai, Y., Qian, J., Yang, J., & Huang, F. (2019). DSFD: dual shot face detector. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 5060-5069). [Google Scholar] [Crossref]
23. He, Y., Xu, D., Wu, L., Jian, M., Xiang, S., & Pan, C. (2019). LFFD: A light and fast face detector for edge devices. arXiv preprint arXiv:1904.10633. [Google Scholar] [Crossref]
24. Albiero, V., Chen, X., Yin, X., Pang, G., & Hassner, T. (2021). img2pose: Face alignment and detection via 6dof, face pose estimation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 7617-7627). [Google Scholar] [Crossref]
25. Belhumeur, P. N., Hespanha, J. P., & Kriegman, D. J. (1996, April). Eigenfaces vs. fisherfaces: Recognition using class specific linear projection. In European conference on computer vision (pp. 43-58). Berlin, Heidelberg: Springer Berlin Heidelberg. [Google Scholar] [Crossref]
26. Bartlett, M. S., Movellan, J. R., & Sejnowski, T. J. (2002). Face recognition by independent component analysis. IEEE Transactions on neural networks, 13(6), 1450-1464. [Google Scholar] [Crossref]
27. Liu, C., & Wechsler, H. (2002). Gabor feature based classification using the enhanced fisher linear discriminant model for face recognition. IEEE Transactions on Image processing, 11(4), 467-476. [Google Scholar] [Crossref]
28. Bay, H., Tuytelaars, T., & Van Gool, L. (2006, May). Surf: Speeded up robust features. In European conference on computer vision (pp. 404-417). Berlin, Heidelberg: Springer Berlin Heidelberg. [Google Scholar] [Crossref]
29. LeCun, Y., Bottou, L., Bengio, Y., &Haffner, P. (1998). Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11), 2278-2324. [Google Scholar] [Crossref]
30. Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25. [Google Scholar] [Crossref]
31. Simonyan, K., & Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556. [Google Scholar] [Crossref]
32. Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., & Rabinovich, A. (2015). Going deeper with convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1-9). [Google Scholar] [Crossref]
33. Huang, G., Liu, Z., Van Der Maaten, L., & Weinberger, K. Q. (2017). Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 4700-4708). [Google Scholar] [Crossref]
34. Howard, A. G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., & Adam, H. (2017). Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861. [Google Scholar] [Crossref]
35. Khalajzadeh, H., Mansouri, M., & Teshnehlab, M. (2013, November). Face recognition using convolutional neural network and simple logistic classifier. In Soft computing in industrial applications: proceedings of the 17th online world conference on soft computing in industrial applications (pp. 197-207). Cham: Springer International Publishing. [Google Scholar] [Crossref]
36. Ding, C., & Tao, D. (2017). Trunk-branch ensemble convolutional neural networks for video-based face recognition. IEEE transactions on pattern analysis and machine intelligence, 40(4), 1002-1014. [Google Scholar] [Crossref]
37. Ferraz, C. T., & Saito, J. H. (2018, October). A comprehensive analysis of local binary convolutional neural network for fast face recognition in surveillance video. In Proceedings of the 24th Brazilian Symposium on Multimedia and the Web (pp. 265-268). [Google Scholar] [Crossref]
38. Yi, X., & Luo, X. (2018, November). A system for real-time detecting and recognizing object person. In Proceedings of the 4th International Conference on Communication and Information Processing (pp. 67-71). [Google Scholar] [Crossref]
39. Zhai, Y., & He, D. (2019, February). Video-based face recognition based on deep convolutional neural network. In Proceedings of the 2019 International Conference on Image, Video and Signal Processing (pp. 23-27). [Google Scholar] [Crossref]
40. Acuña-Escobar, D., Ibarra-Fiallo, J., & Intriago-Pazmiño, M. (2019, December). Real-time face identification from video surveillance cameras. In Proceedings of the Second International Conference on Data Science, E-Learning and Information Systems (pp. 1-4). [Google Scholar] [Crossref]
41. Mendes, P. R. C., Busson, A. J. G., Colcher, S., Schwabe, D., Guedes, Á. L. V., & Laufer, C. (2020, November). A cluster-matching-based method for video face recognition. In Proceedings of the Brazilian Symposium on Multimedia and the Web (pp. 97-104). [Google Scholar] [Crossref]
42. Jiwei, X., Jiyuan, X., Yi, F., &Dongfang, C. (2020, June). Research on video face retrieval method based on deep learning and key frame. In Proceedings of the 2020 4th International Conference on Digital Signal Processing (pp. 75-80). [Google Scholar] [Crossref]
43. Patel, B. R., & Desai, A. A. (2026). Real-time face identification from video in an uncontrolled environment using CNN. International Journal of Innovative Technology and Exploring Engineering, 15(3), 1–7. https://doi.org/10.35940/IJITEE.A3710.15030226 [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- Employee Training Efficacy in Banking Sector
- Multidimensional Predictors of University Students’ Examination Performance: A Quantitative Analysis of Psychological, Environmental, And Skill-Based Factors
- Digital Learning Innovation: Evaluating the Effectiveness of MOOC and Politicbox in Enhancing Students Understanding
- An Investigation of Group Work through Mcclelland’s Theory
- The Impact of a Practical Flipped Process-Genre Writing Model on Moroccan EFL High School Students’ Writing Performance