Home / Journals / CMC / Online First / doi:10.32604/cmc.2026.081091
Special Issues
Table of Content

Open Access

ARTICLE

Seeing through Deepfakes: An Explainable Multi-Task Detection Framework with Deep Learning and Large Language Models

Jiyeong Park1, Sercan Yeşilköy1, Doyeon Lim1, Huiryeong Park1, Eunseo Lee1, Mohsen Ali Alawami1,*, Ki-Woong Park2,*
1 Division of Computer Engineering, Hankuk University of Foreign Studies, Yongin, Republic of Korea
2 Department of Information Security, Sejong University, Seoul, Republic of Korea
* Corresponding Author: Mohsen Ali Alawami. Email: email; Ki-Woong Park. Email: email

Computers, Materials & Continua https://doi.org/10.32604/cmc.2026.081091

Received 23 February 2026; Accepted 24 June 2026; Published online 24 July 2026

Abstract

The recent increase in deepfake content has significantly increased cyber threats. Although numerous deepfake detection technologies have achieved high accuracy, there are limits to clarifying the rationale behind their detection decisions. To bridge the gap, in our study, we leverage the combination of Explainable Artificial Intelligence (XAI) and Large Language Models (LLMs) to deliver clear, consistent, and understandable interpretations of deepfake detection outcomes. To do that, we integrate XAI and LLMs to visually represent detection rationales and automatically generate coherent natural-language explanations. During the implementation of our method, we developed a multi-task learning framework based on a Convolutional Neural Network (CNN) combined with a Long Short-Term Memory (LSTM) architecture, trained simultaneously for binary classification (real or fake) and multi-class classification (identifying specific deepfake techniques) using the FaceForensics++ dataset. The CNN architecture considered for this model included ResNeXt50-32x4d, EfficientNet-b0, and Xception, with EfficientNet-b0 ultimately selected as the optimal detection model based on superior performance metrics on the FaceForensics++ dataset. EfficientNet-b0 achieved the best performance, with an accuracy of 95.14% for binary classification and 94.29% for multi-class classification. Additionally, a web-based interface was developed to allow intuitive inspection of the detection results, combining visual and textual explanations to enhance interpretability.

Keywords

Deepfake detection; multi-task learning; XAI; LLM; CNN-LSTM architecture
  • 360

    View

  • 60

    Download

  • 0

    Like

Share Link