Home / Journals / CMC / Online First / doi:10.32604/cmc.2026.084902
Special Issues
Table of Content

Open Access

ARTICLE

Vision Transformer–Based Deepfake Detection Across Multiple Generation Methods: A Transfer Learning Approach

Ahmad Raza1,*, Abdul Basit1,*, Syed Muqtar Ahmed2, Zeeshan Ahmad Arfeen3,*, Muhammad I. Masud4, Muhammad Farid Zamir5, Mehreen Kausar Azam6, Touqeer Ahmed Jumani7
1 Department of Information and Communication Engineering, The Islamia University of Bahawalpur (IUB), Bahawalpur, Pakistan
2 Department of Software Engineering, College of Engineering, University of Business and Technology, Jeddah, Saudi Arabia
3 Department of Electrical Engineering, Faculty of Engineering and Technology, The Islamia University of Bahawalpur, Bahawalpur, Pakistan
4 Department of Electrical Engineering, College of Engineering, University of Business and Technology, Jeddah, Saudi Arabia
5 Directorate of IT, The Islamia University of Bahawalpur, Bahawalpur, Pakistan
6 Department of Industrial Manufacturing Engineering, Pakistan Navy Engineering College, National University of Sciences and Technology (NUST), Karachi, Pakistan
7 College of Engineering, A’Sharqiyah University, Ibra, Oman
* Corresponding Author: Ahmad Raza. Email: email; Abdul Basit. Email: email; Zeeshan Ahmad Arfeen. Email: email

Computers, Materials & Continua https://doi.org/10.32604/cmc.2026.084902

Received 01 May 2026; Accepted 14 July 2026; Published online 12 August 2026

Abstract

The development of deepfake technologies is a threat to digital media authentication and cybersecurity infrastructure. The current paper proposes a method for detecting manipulated images of faces based on the Vision Transformer architecture. We fine-tune a pre-trained ViT-Base-Patch16-224 model based on this well-curated dataset of 12,137 face images, which includes an almost equal number of real and synthetic face images using a variety of different generation methods. The data set contains real-life photographs of CelebA and FFHQ, along with artificial samples of the publicly available Kaggle repositories (FaceForensics++, Celeb-DF, and DFDC) and 600 self-collected photos (300 real-life photographs of personal cell phones and 300 artificial ones created with the help of modern tools) to make it closer to real-life use. Methodology involves assessment of the quality of data, systematic preprocessing by ImageNet normalization, data augmentation, and stratification of a 70-20-10 partitioning of the data. AdamW optimization using a learning rate of 2 × 10–5 was used with 8 epochs. The test accuracy of the system was 99.01%, the precision was 98.85%, and the recall was 99.18%, with 12 misclassifications. The results demonstrate that Vision Transformers can effectively model global image dependencies for detecting Deepfakes. The complete methodology is documented in this paper to ensure reproducibility.

Keywords

Deep learning; deepfake detection; digital forensics; facial manipulation; image authentication; machine learning; vision transformer
  • 301

    View

  • 31

    Download

  • 1

    Like

Share Link