Video Deepfake Detection using Hybrid CNN-Transformer Models

Authors

  • Aryan Raj Galgotias University, India
  • Prince Singh Galgotias University, India
  • Kajal Gupta Galgotias University, India

DOI:

https://doi.org/10.65890/dmp-lncse.ICICCS26.183

Keywords:

Deepfake Detection, Convolutional Neural Networks (CNN), Transformer Models, Video Forensics, Faceforensics++, DFDC, Multimedia Security

Abstract

Deepfakes on a mass scale have been enabled by advancements in deep generative models. Conventional deepfake methods based on manually designed features or independent CNN models are found to have limitations in generalising across different types of modification methods and levels of compression. In this paper, a combined CNN-Transformer approach for exact deepfakes video detection has been used. In this combined approach, the Transformer component has been used for representing relations among different lifelike videos. On the other hand, for spatial mapping of registered artefacts in deepfake images, a CNN component has been used. This combined approach has been found to perform better at detecting deepfakes than other models when tested on the FaceForensics++, DFDC, and Celeb-DF image datasets. Bias in data consideration in testing and research ethics has been considered in this paper for future research on exact multimedia forensics.

Downloads

Published

26-07-2026

Conference Proceedings Volume

Section

Articles

How to Cite

Raj, A. ., Singh, P. ., & Gupta, K. . (2026). Video Deepfake Detection using Hybrid CNN-Transformer Models. DMPedia Lecture Notes in Computer Science & Engineering, ICICCS26, 25-29. https://doi.org/10.65890/dmp-lncse.ICICCS26.183