Deep Learning-Based Video Facial Recognition System Using MTCNN and VGG-16 with Transfer Learning for Character Identification
Keywords:
EspañolAbstract
Automatic face recognition has evolved significantly with the use of deep learning techniques, enabling more robust systems that can withstand variations in lighting, pose, and expression. This paper presents the implementation of a face recognition system for video recordings using the MTCNN face detection algorithm, which is used for its accuracy in identifying and locating faces in uncontrolled environments with varying lighting conditions. After face detection, the VGG-16 model was used as a deep feature extractor, leveraging its architecture to employ transfer learning techniques to learn hierarchical and discriminative representations of the face. The system processes the video frames using the OpenCV library, then uses the MTCNN network for facial identification, and finally assigns identities through a classifier trained, as previously mentioned, with the VGG-16 model. In recent studies, authors report that convolutional neural networks consistently outperform traditional methods due to their ability to automate the extraction of relevant features and handle complex environmental conditions. Preliminary results indicate that the MTCNN–VGG-16 integration forms an effective strategy for facial recognition in real-world scenarios, demonstrating high accuracy and stability under diverse visual conditions. This work provides an accessible and reproducible architecture for potential applications in video analytics, surveillance, and intelligent automationReferences
Almabdy, S., & Elrefaei, L. (2019). Deep Convolutional Neural Network-Based Approaches for Face Recognition. Applied Sciences, 9(20), 4397. https://doi.org/10.3390/app9204397
Ardiawan, M. I., & Negarara, G. P. K. (2024). A Comparative Analysis of FaceNet, VGGFace, and GhostFaceNets Face Recognition Algorithms For Potential Criminal Suspect Identification. Journal of Applied Artificial Intelligence, 5(2), 34-49. https://doi.org/10.48185/jaai.v5i2.1237
Athira, K. A., & Divya Udayan, J. (2024). Temporal Fusion of Time-Distributed VGG-16 and LSTM for Precise Action Recognition in Video Sequences. Procedia Computer Science, 233, 892-901. https://doi.org/10.1016/j.procs.2024.03.278
Cao, K., Rong, Y., Li, C., Tang, X., & Loy, C. C. (2018). Pose-robust face recognition via deep residual equivariant mapping. arXiv. https://arxiv.org/abs/1803.00839
Dar, S. A., & Palanivel, S. (2021). Performance evaluation of convolutional neural networks (CNNs) and VGG on real time face recognition system. Advances in Science, Technology and Engineering Systems Journal, 6(2), 956–964. https://doi.org/10.25046/aj0602109
Ding, Y., Cheng, Y., Cheng, X., Li, B., You, X., & Yuan, X. (2017). Noise-resistant network: a deep-learning method for face recognition under noise. EURASIP Journal on Image and Video Processing, 2017, Article 43. https://doi.org/10.1186/s13640-017-0188-z
Fuad, Md. T. H., Fime, A. A., Sikder, D., Iftee, Md. A. R., Rabbi, J., Al-Rakhami, M. S., Gumaei, A., Sen, O., Fuad, M., & Islam, Md. N. (2021). Recent Advances in Deep Learning Techniques for Face Recognition. IEEE Access, 9, 99112-99142. https://doi.org/10.1109/ACCESS.2021.3096136
Gilligan, V. (Creador). (2008–2013). Breaking Bad [Serie de televisión]. High Bridge Productions; Gran Via Productions; Sony Pictures Television.
Gupta, J., Pathak, S., & Kumar, G. (2022). Deep Learning (CNN) and Transfer Learning: A Review. Journal of Physics: Conference Series, 2273(1), 012029. https://doi.org/10.1088/1742-6596/2273/1/012029
Hassanat, A. B., Albustanji, A., Tarawneh, A. S., Alrashidi, M., Alharbi, H., Alanazi, M., Alghamdi, M., Alkhazi, I. S., & Prasath, V. B. S. (2021). Deep learning for identification and face, gender, expression recognition under constraints. arXiv. https://arxiv.org/abs/2111.01930
Hosna, A., Merry, E., Gyalmo, J., Alom, Z., Aung, Z., & Azim, M. A. (2022). Transfer learning: A friendly introduction. Journal of Big Data, 9(1), 102. https://doi.org/10.1186/s40537-022-00652-w
Kortli, Y., Jridi, M., Al Falou, A., & Atri, M. (2020). Face Recognition Systems: A Survey. Sensors, 20(2), 342. https://doi.org/10.3390/s20020342
Krichen, M. (2023). Convolutional Neural Networks: A Survey. Computers, 12(8), 151. https://doi.org/10.3390/computers12080151
Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). ImageNet Classification with Deep Convolutional Neural Networks. In Advances in neural information processing systems (Vol. 25).
Li, L., Mu, X., Li, S., & Peng, H. (2020). A Review of Face Recognition Technology. IEEE Access, 8, 139110-139120. https://doi.org/10.1109/ACCESS.2020.3011028
Nguyen-Meidine, L. T., Granger, E., Kiran, M., & Blais-Morin, L.-A. (2017, November). A comparison of CNN-based face and head detectors for real-time video surveillance applications. In 2017 Seventh International Conference on Image Processing Theory, Tools and Applications (IPTA) (pp. 1–7). IEEE. https://doi.org/10.1109/IPTA.2017.8310113
Orozco, C. I., Xamena, E., Buemi, M. E., & Berlles, J. J. (2020). Human Action Recognition in Videos Using a Robust CNN–LSTM Approach. Ciencia y Tecnología, (20), 23–36. https://dialnet.unirioja.es/servlet/articulo?codigo=7763841
Parkhi, O. M., Vedaldi, A., & Zisserman, A. (2015). Deep Face Recognition. En Proceedings of the British Machine Vision Conference (BMVC 2015) (pp. 41.1-41.12). https://doi.org/10.5244/C.29.41
Paulchamy, B., Yahya, A., Chinnasamy, N., & Kasilingam, K. (2025). Facial Expression Recognition Through Transfer Learning: Integration of VGG16, ResNet, and AlexNet with a Multiclass Classifier. Acadlore Transactions on AI and Machine Learning, 4(1), 25–39. https://doi.org/10.56578/ataiml040103
Qawaqneh, Z., Mallouh, A. A., & Barkana, B. D. (2017). Deep Convolutional Neural Network for Age Estimation based on VGG-Face Model (No. arXiv:1709.01664). arXiv. https://doi.org/10.48550/arXiv.1709.01664
Sanabria Moyano, J. E., Roa Avella, M. del P., Lee Pérez, O. I., Sanabria Moyano, J. E., Roa Avella, M. del P., & Lee Pérez, O. I. (2022). Tecnología de reconocimiento facial y sus riesgos en los derechos humanos. Revista Criminalidad, 64(3), 61-78. https://doi.org/10.47741/17943108.366
Sanchez, W., & Sigua, E. (s. f.). Reconocimiento facial en video usando Deep learning.
Sarabu, A., & Santra, A. K. (2021). Human Action Recognition in Videos using Convolution Long Short-Term Memory Network with Spatio-Temporal Networks. Emerging Science Journal, 5(1), 25–33. https://doi.org/10.28991/esj-2021-01254
Southwest Minzu University, Key Laboratory of Electronic and Information Engineering, State Ethnic Affairs Commission, Chengdu, 610041, China, Ku, H., Dong, W., & Southwest Minzu University, Key Laboratory of Electronic and Information Engineering, State Ethnic Affairs Commission, Chengdu, 610041, China. (2020). Face Recognition Based on MTCNN and Convolutional Neural Network. Frontiers in Signal Processing, 4(1). https://doi.org/10.22606/fsp.2020.41006
Vimal, C., & Shirivastava, N. (2022). Face and Face-mask Detection System using VGG-16 Architecture based on Convolutional Neural Network. International Journal of Computer Applications, 183(50), 16-21. https://doi.org/10.5120/ijca2022921700
Wang, K. (2024). An Exploration of Face Recognition Methods Based on LBP Algorithm and PCA Analysis. Proceedings of the 2024 AETR Conference on Artificial Intelligence & Emerging Trends in Research and Methodologies. https://madison-proceedings.com/index.php/aetr/article/view/2043
Xie, Y., Wang, H., & Guo, S. (2020). Research on MTCNN face recognition system in low computing power scenarios. Journal of Internet Technology, 21(5), 1463–1475. https://jit.ndhu.edu.tw/article/view/2380/2396
Zhang, K., Zhang, Z., Li, Z., & Qiao, Y. (2016). Joint face detection and alignment using multitask cascaded convolutional networks. IEEE Signal Processing Letters, 23(10), 1499–1503. https://doi.org/10.1109/LSP.2016.2603342
Zhong, Y., Deng, W., Hu, J., Zhao, D., Li, X., & Wen, D. (2021). SFace: Sigmoid-constrained hypersphere loss for robust face recognition. IEEE Transactions on Image Processing, 30, 2587–2598. https://doi.org/10.1109/TIP.2020.3048632