Video Deepfake Detection using Hybrid CNN-Transformer Models
DOI:
https://doi.org/10.65890/dmp-lncse.ICICCS26.183Keywords:
Deepfake Detection, Convolutional Neural Networks (CNN), Transformer Models, Video Forensics, Faceforensics++, DFDC, Multimedia SecurityAbstract
Deepfakes on a mass scale have been enabled by advancements in deep generative models. Conventional deepfake methods based on manually designed features or independent CNN models are found to have limitations in generalising across different types of modification methods and levels of compression. In this paper, a combined CNN-Transformer approach for exact deepfakes video detection has been used. In this combined approach, the Transformer component has been used for representing relations among different lifelike videos. On the other hand, for spatial mapping of registered artefacts in deepfake images, a CNN component has been used. This combined approach has been found to perform better at detecting deepfakes than other models when tested on the FaceForensics++, DFDC, and Celeb-DF image datasets. Bias in data consideration in testing and research ethics has been considered in this paper for future research on exact multimedia forensics.
Downloads
Published
Conference Proceedings Volume
Section
License
Copyright (c) 2026 DMPedia Lecture Notes in Computer Science & Engineering

This work is licensed under a Creative Commons Attribution 4.0 International License.