Open Access
ARTICLE
Seeing through Deepfakes: An Explainable Multi-Task Detection Framework with Deep Learning and Large Language Models
1 Division of Computer Engineering, Hankuk University of Foreign Studies, Yongin, Republic of Korea
2 Department of Information Security, Sejong University, Seoul, Republic of Korea
* Corresponding Authors: Mohsen Ali Alawami. Email: ; Ki-Woong Park. Email:
Computers, Materials & Continua 2026, 89(1), 18 https://doi.org/10.32604/cmc.2026.081091
Received 23 February 2026; Accepted 24 June 2026; Issue published 13 August 2026
Abstract
The recent increase in deepfake content has significantly increased cyber threats. Although numerous deepfake detection technologies have achieved high accuracy, there are limits to clarifying the rationale behind their detection decisions. To bridge the gap, in our study, we leverage the combination of Explainable Artificial Intelligence (XAI) and Large Language Models (LLMs) to deliver clear, consistent, and understandable interpretations of deepfake detection outcomes. To do that, we integrate XAI and LLMs to visually represent detection rationales and automatically generate coherent natural-language explanations. During the implementation of our method, we developed a multi-task learning framework based on a Convolutional Neural Network (CNN) combined with a Long Short-Term Memory (LSTM) architecture, trained simultaneously for binary classification (real or fake) and multi-class classification (identifying specific deepfake techniques) using the FaceForensics++ dataset. The CNN architecture considered for this model included ResNeXt50-32x4d, EfficientNet-b0, and Xception, with EfficientNet-b0 ultimately selected as the optimal detection model based on superior performance metrics on the FaceForensics++ dataset. EfficientNet-b0 achieved the best performance, with an accuracy of 95.14% for binary classification and 94.29% for multi-class classification. Additionally, a web-based interface was developed to allow intuitive inspection of the detection results, combining visual and textual explanations to enhance interpretability.Keywords
Cite This Article
Copyright © 2026 The Author(s). Published by Tech Science Press.This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.


Submit a Paper
Propose a Special lssue
View Full Text
Download PDF
Downloads
Citation Tools