Open Access iconOpen Access

ARTICLE

Dual-Stream Facial Emotion Recognition with Self-Supervised Pre-Training and Evidential Uncertainty

Rashid Jahangir1,*, Nazik Alturki2, Mohammed Alreshoodi3

1 Department of Computer Science, COMSATS University Islamabad, Vehari Campus, Vehari, Pakistan
2 Department of Information Systems, College of Computer and Information Sciences, Princess Nourah bint Abdulrahman University, Riyadh, Saudi Arabia
3 Unit of Scientific Research, Applied College, Qassim University, Buraydah, Saudi Arabia

* Corresponding Author: Rashid Jahangir. Email: email

(This article belongs to the Special Issue: Machine Learning and Deep Learning-Based Pattern Recognition, 2nd Edition)

Computer Modeling in Engineering & Sciences 2026, 148(3), 37 https://doi.org/10.32604/cmes.2026.086137

Abstract

Facial emotion recognition (FER) remains difficult in real-world settings. Inter-subject variability, lighting changes, occlusion, and class imbalance all limit performance. Most FER systems rely on one convolutional or transformer backbone. This narrows the features available for classification. This paper presents Dual-Stream FERNet. It is a carefully evaluated integration of an EfficientNetV2-S backbone with a Swin Transformer Tiny backbone, joined by a learnable sigmoid-gated fusion module. Before fine-tuning, both branches undergo SimCLR-style self-supervised pre-training on two augmented views. This gives a stronger initialization without extra labels. An Evidential Deep Learning head then produces class probabilities and Dirichlet-parameterized uncertainty together. The model is tested on two benchmarks, KDEF and CK+. Under subject-disjoint 5-fold cross-validation, the model reached 93.84 ± 1.73% accuracy on KDEF and 93.07 ± 1.22% on CK+. The model is compared against fair, SSL-matched baselines on KDEF, ResNet50, EfficientNet-B0, and Swin-Small. The model’s real advantage is calibrated uncertainty, not higher accuracy. On KDEF the model runs at 459.67 FPS (NVIDIA RTX 4070, FP32, batch size 16, batch-1 median latency 18.96 ms). Grad-CAM shows the model attending to facial regions tied to FACS action units on both datasets. Overall, the model matches strong single-stream baselines in accuracy and adds calibrated uncertainty on top.

Keywords

Facial emotion recognition; dual-stream CNN-transformer; EfficientNetV2-S; swin transformer; sigmoid-gated feature fusion; self-supervised learning

Cite This Article

APA Style
Jahangir, R., Alturki, N., Alreshoodi, M. (2026). Dual-Stream Facial Emotion Recognition with Self-Supervised Pre-Training and Evidential Uncertainty. Computer Modeling in Engineering & Sciences, 148(3), 37. https://doi.org/10.32604/cmes.2026.086137
Vancouver Style
Jahangir R, Alturki N, Alreshoodi M. Dual-Stream Facial Emotion Recognition with Self-Supervised Pre-Training and Evidential Uncertainty. Comput Model Eng Sci. 2026;148(3):37. https://doi.org/10.32604/cmes.2026.086137
IEEE Style
R. Jahangir, N. Alturki, and M. Alreshoodi, “Dual-Stream Facial Emotion Recognition with Self-Supervised Pre-Training and Evidential Uncertainty,” Comput. Model. Eng. Sci., vol. 148, no. 3, pp. 37, 2026. https://doi.org/10.32604/cmes.2026.086137



cc Copyright © 2026 The Author(s). Published by Tech Science Press.
This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
  • 232

    View

  • 78

    Download

  • 0

    Like

Share Link