Home / Journals / CMES / Online First / doi:10.32604/cmes.2026.086137
Special Issues
Table of Content

Open Access

ARTICLE

Dual-Stream Facial Emotion Recognition with Self-Supervised Pre-Training and Evidential Uncertainty

Rashid Jahangir1,*, Nazik Alturki2, Mohammed Alreshoodi3
1 Department of Computer Science, COMSATS University Islamabad, Vehari Campus, Vehari, Pakistan
2 Department of Information Systems, College of Computer and Information Sciences, Princess Nourah bint Abdulrahman University, Riyadh, Saudi Arabia
3 Unit of Scientific Research, Applied College, Qassim University, Buraydah, Saudi Arabia
* Corresponding Author: Rashid Jahangir. Email: email
(This article belongs to the Special Issue: Machine Learning and Deep Learning-Based Pattern Recognition, 2nd Edition)

Computer Modeling in Engineering & Sciences https://doi.org/10.32604/cmes.2026.086137

Received 25 May 2026; Accepted 07 September 2026; Published online 17 September 2026

Abstract

Facial emotion recognition (FER) remains difficult in real-world settings. Inter-subject variability, lighting changes, occlusion, and class imbalance all limit performance. Most FER systems rely on one convolutional or transformer backbone. This narrows the features available for classification. This paper presents Dual-Stream FERNet. It is a carefully evaluated integration of an EfficientNetV2-S backbone with a Swin Transformer Tiny backbone, joined by a learnable sigmoid-gated fusion module. Before fine-tuning, both branches undergo SimCLR-style self-supervised pre-training on two augmented views. This gives a stronger initialization without extra labels. An Evidential Deep Learning head then produces class probabilities and Dirichlet-parameterized uncertainty together. The model is tested on two benchmarks, KDEF and CK+. Under subject-disjoint 5-fold cross-validation, the model reached 93.84 ± 1.73% accuracy on KDEF and 93.07 ± 1.22% on CK+. The model is compared against fair, SSL-matched baselines on KDEF, ResNet50, EfficientNet-B0, and Swin-Small. The model’s real advantage is calibrated uncertainty, not higher accuracy. On KDEF the model runs at 459.67 FPS (NVIDIA RTX 4070, FP32, batch size 16, batch-1 median latency 18.96 ms). Grad-CAM shows the model attending to facial regions tied to FACS action units on both datasets. Overall, the model matches strong single-stream baselines in accuracy and adds calibrated uncertainty on top.

Keywords

Facial emotion recognition; dual-stream CNN-transformer; EfficientNetV2-S; swin transformer; sigmoid-gated feature fusion; self-supervised learning
  • 60

    View

  • 22

    Download

  • 0

    Like

Share Link