Open Access
ARTICLE
Dual-Stream Facial Emotion Recognition with Self-Supervised Pre-Training and Evidential Uncertainty
1 Department of Computer Science, COMSATS University Islamabad, Vehari Campus, Vehari, Pakistan
2 Department of Information Systems, College of Computer and Information Sciences, Princess Nourah bint Abdulrahman University, Riyadh, Saudi Arabia
3 Unit of Scientific Research, Applied College, Qassim University, Buraydah, Saudi Arabia
* Corresponding Author: Rashid Jahangir. Email:
(This article belongs to the Special Issue: Machine Learning and Deep Learning-Based Pattern Recognition, 2nd Edition)
Computer Modeling in Engineering & Sciences 2026, 148(3), 37 https://doi.org/10.32604/cmes.2026.086137
Received 25 May 2026; Accepted 07 September 2026; Issue published 28 September 2026
Abstract
Facial emotion recognition (FER) remains difficult in real-world settings. Inter-subject variability, lighting changes, occlusion, and class imbalance all limit performance. Most FER systems rely on one convolutional or transformer backbone. This narrows the features available for classification. This paper presents Dual-Stream FERNet. It is a carefully evaluated integration of an EfficientNetV2-S backbone with a Swin Transformer Tiny backbone, joined by a learnable sigmoid-gated fusion module. Before fine-tuning, both branches undergo SimCLR-style self-supervised pre-training on two augmented views. This gives a stronger initialization without extra labels. An Evidential Deep Learning head then produces class probabilities and Dirichlet-parameterized uncertainty together. The model is tested on two benchmarks, KDEF and CK+. Under subject-disjoint 5-fold cross-validation, the model reached 93.84 ± 1.73% accuracy on KDEF and 93.07 ± 1.22% on CK+. The model is compared against fair, SSL-matched baselines on KDEF, ResNet50, EfficientNet-B0, and Swin-Small. The model’s real advantage is calibrated uncertainty, not higher accuracy. On KDEF the model runs at 459.67 FPS (NVIDIA RTX 4070, FP32, batch size 16, batch-1 median latency 18.96 ms). Grad-CAM shows the model attending to facial regions tied to FACS action units on both datasets. Overall, the model matches strong single-stream baselines in accuracy and adds calibrated uncertainty on top.Keywords
Automatic facial emotion recognition (FER) is a core problem in computer vision and affective computing. Reading emotional states from facial images supports applications such as online gaming, patient monitoring, mental health assessment, and driver fatigue detection [1]. Yet robust FER under realistic, unconstrained conditions is still an open problem [2]. The same emotion can look different across individuals due to differences in facial muscle structure, cultural display rules, and expressiveness. Real-world images add further difficulties. Lighting changes, occlusion from accessories, and non-frontal poses all distort the spatial layout of facial action units [3,4]. Class imbalance compounds these problems, since neutral and happy expressions appear far more often than fear or disgust in most datasets, biasing classifiers trained with standard cross-entropy (CE) loss toward the majority classes [5].
Early FER methods relied on hand-crafted descriptors—Local Binary Patterns (LBP), Histogram of Oriented Gradients (HOG), and Gabor filter banks—paired with Support Vector Machines or Random Forest classifiers. These methods are interpretable but require careful feature engineering and generalize poorly across imaging conditions. Deep Convolutional Neural Networks (CNNs) advanced the field by learning hierarchical representations directly from pixels [6]. Still, CNNs trained with plain cross-entropy loss face three recurring problems in FER. They overfit easily when labeled data is scarce, they produce overconfident softmax scores that poorly reflect true uncertainty, and they rely on a single backbone that captures only one type of feature representation. Standard convolutional architectures extract local convolutional features but lack any self-attention mechanism, while transformer architectures capture global self-attention relationships but lack an inductive bias toward local convolutional structure. Neither captures both simultaneously [7,8].
Transformer-based models have recently shown strong FER performance by capturing long-range dependencies between distant facial action units [9,10]. But pure transformers need large labeled datasets and heavy compute. Lightweight CNNs, such as EfficientNet variants, run faster but often miss global relationships between distant facial regions [11]. Combining both paradigms in one framework remains underexplored for FER [4,8]. Self-supervised learning (SSL) offers one way to ease data scarcity by pre-training backbones on unlabeled data through pretext tasks. SimCLR-style contrastive learning [12] trains a network to pull together two augmented views of the same image in embedding space, producing representations that transfer well to downstream classification. For FER specifically, SSL pre-training on the target dataset’s unlabeled images adapts the backbone more closely to facial image statistics than generic ImageNet pre-training does [13,14].
Standard FER classifiers also cannot quantify uncertainty. A model that assigns 95% confidence to “happy” for a subtly ambiguous micro-expression looks identical, on paper, to one that assigns 95% confidence to a clearly happy face. Evidential Deep Learning (EDL) addresses this by modeling class probabilities as a Dirichlet distribution parameterized by evidence, where low total evidence signals high uncertainty [15]. This uncertainty can be used during training to down-weight mislabeled or ambiguous samples, and during deployment to flag cases for human review [16]. This paper proposes Dual-Stream FERNet, a unified architecture that addresses all four challenges above. The contributions of this study are as follows.
• A dual-stream backbone pairing EfficientNetV2-S (speed, local CNN features) with Swin-T (global self-attention features), fused via a learnable sigmoid-gated feature fusion module that computes a per-dimension gating vector and produces a dynamically weighted combination of both branch outputs—a gated weighted sum rather than a query-key-value attention operation.
• A dual-branch SimCLR-style SSL pre-training phase applied jointly to both backbones as lightweight domain adaptation of ImageNet pre-trained representations toward the FER distribution, prior to supervised fine-tuning.
• An Evidential Deep Learning (EDL) classification head replacing softmax with Dirichlet-parameterized evidence outputs, providing calibrated uncertainty estimates used both as an inference signal and as a per-sample loss-weighting mechanism during training.
The remainder of this paper is organized as follows. Section 2 reviews related work on deep learning-based FER. Section 3 describes the proposed Dual-Stream FERNet architecture, SSL pre-training strategy, and training procedure. Section 4 presents experimental results and discussion. Section 5 concludes with directions for future work.
The evolution of FER systems spans three broad phases, hand-crafted feature methods, shallow neural networks, and deep learning-based approaches. This review focuses on deep learning-based FER methods, organized around three technical dimensions central to current research, backbone architecture, training strategy, and uncertainty modelling.
2.1 Single-Backbone CNN and Hybrid Architectures
Early deep learning FER systems established that CNNs trained end-to-end on facial image data substantially outperform hand-crafted feature approaches. A hierarchical deep neural network structure for FER was proposed [17] and achieved 96.5% accuracy on CK+ and 91.3% on JAFFE. The proposed method improved upon prior baselines by 1.3% and 1.5%, respectively. The approach reclassified top erroneous predictions using a hierarchical structure but operated on a single backbone without attention mechanisms. A hybrid CNN-RNN architecture was introduced [18] to jointly learn spatial features via CNN and temporal sequence patterns via RNN on the MMI and JAFFE datasets. This architecture achieved 94.91% accuracy with six hidden layers. An extension of this work replaced the temporal dataset with CK+ and incorporated deep residual blocks with a fully connected network. The extended research [19] achieved 93.24% and 95.23% accuracy on CK+ and JAFFE, respectively. These results demonstrate the benefit of combining spatial feature extraction with sequential modelling. However, the study did not address the limitations of a single convolutional stream.
A CNN combined with image edge detection was evaluated [20] on a mixed Fer-2013/LFW dataset. The model achieved 88.56% average accuracy with training speed 1.5 times faster than competing methods. An attentional convolutional network (Deep-Emotion) incorporating a spatial attention mechanism was applied to FER-2013, CK+, JAFFE, and FERG datasets. This demonstrated that explicit attention mechanisms substantially improved over standard CNN feature aggregation. Two CNN models for FER in the wild were proposed [21], a lightweight external network for industrial deployment and a unified network combining deep feature extraction with LBP learning in parallel.
2.2 Transfer Learning and Data Augmentation
Transfer learning from large-scale image classification datasets remains the dominant way to address data scarcity in FER. Models such as VGG16, ResNet50, and InceptionV3, pre-trained on ImageNet, are commonly used as backbone initializations and then fine-tuned on FER datasets. However, the domain gap between natural image classification and facial expression recognition limits how well ImageNet-derived features transfer, particularly for subtle emotion categories such as fear and disgust.
Data augmentation has also been studied as a complementary strategy. A detailed comparison of augmentation techniques and deep learning features showed that augmentation meaningfully improves FER accuracy on the FER-2013, extended FER-2013, and AffectNet datasets [22]. Separately, fusing handcrafted and automatically extracted features on the same datasets yielded roughly 1% higher accuracy than standard approaches [23]. Both studies, however, used fixed, predefined augmentation pipelines and did not adapt augmentation strength to the target domain. A hierarchical weighted random forest (WRF) classifier for driver FER, evaluated on the CK+, KMU-FED, and MMI databases, reached 92.6% accuracy on CK+ [24].
Knowledge distillation offers another route to real-time FER deployment. A real-time emotion recognition system based on adaptive knowledge distillation [25] trains a lightweight student network to match the soft-probability outputs of a larger, pre-trained teacher model, achieving competitive accuracy at much lower inference cost. This differs fundamentally from architectural parallelism or model compression as a strategy for meeting real-time requirements.
2.3 Transformer-Based and Multi-Branch Architectures
The introduction of Vision Transformers (ViT) and hierarchical Swin Transformers has established a new direction for facial expression recognition (FER) research. In contrast to convolutional architectures, transformers employ self-attention mechanisms that capture long-range dependencies between distant facial regions [26]. This capability enables more effective modeling of the interactions among multiple action units, including brow raising, lip corner pulling, and eye widening, which collectively define facial expressions.
Temporal transformer architectures have been utilized in video-based facial expression recognition (FER). STT-Net [27] introduces a simplified temporal transformer that reduces the computational complexity associated with full self-attention across video frame sequences, while maintaining competitive accuracy on sequence-level FER benchmarks. These temporal models are particularly effective for video-sequence FER, where the dynamics of expressions provide discriminative temporal information. In contrast, the current study addresses static single-image FER, a setting in which temporal information is absent and the quality of spatial features primarily determines classification performance.
EfficientNet variants have demonstrated state-of-the-art accuracy-efficiency trade-offs on image classification benchmarks and have been successfully applied to FER. However, pure EfficientNet architectures do not explicitly model global context. To date, no published studies have investigated dual-stream architectures that combine a convolutional backbone with a transformer backbone for FER using a learnable gated fusion mechanism. Additionally, the application of self-supervised contrastive pre-training (SimCLR) to FER has been restricted to single-branch settings [28]. This study addresses both gaps by introducing a gated EfficientNetV2-S and Swin-T fusion architecture with dual-branch self-supervised learning (SSL) pre-training.
2.4 Uncertainty Modelling in FER
Standard FER classifiers produce softmax probability vectors that are systematically overconfident. They assign high probability to a single class even for genuinely ambiguous expressions. This overconfidence is problematic in safety-critical applications such as driver monitoring or clinical mental health screening, where awareness of model uncertainty is essential for appropriate system behavior. Evidential Deep Learning (EDL) provides a principled alternative by modelling the class probability vector as a Dirichlet distribution [29]. The concentration parameters of the Dirichlet are output by the network as evidence values, and the total evidence determines the precision of the distribution. Low total evidence corresponds to high uncertainty, providing a signal for genuinely ambiguous expressions. Out-of-distribution detection is not evaluated in this study and accordingly do not claim this uncertainty signal detects out-of-distribution inputs. EDL has been applied in image classification and medical imaging tasks but has not previously been applied to FER in a dual-stream gated fusion setting. No prior published study has combined Evidential Deep Learning with dual-backbone CNN-Transformer sigmoid-gated feature fusion and SimCLR-style dual-branch SSL pre-training within a single FER framework.
The proposed FER system takes a facial image as input and classifies it into one of seven or eight emotion categories, depending on the dataset. The framework operates in four stages, (a) data preparation and augmentation, (b) dual-branch self-supervised pre-training, (c) supervised fine-tuning with evidential uncertainty, and (d) evaluation and visualization. The overall pipeline is illustrated in Fig. 1.

Figure 1: Proposed Dual-Stream FERNet pipeline.
Two benchmark datasets are used for evaluation. The KDEF (Karolinska Directed Emotional Faces) [30] contains images of 70 subjects (35 male, 35 female) posed in seven emotion categories, anger, disgust, fear, happy, neutral, sad, and surprise. 2938 images across the seven classes are used in this study, drawn from a publicly available pre-sorted mirror of the dataset. Images are captured under controlled laboratory conditions with consistent illumination and frontal or near-frontal poses. The second CK+ dataset [31] is drawn from a publicly available mirror providing 981 images across seven emotion categories, anger, contempt, disgust, fear, happy, sadness, and surprise.
Model performance is evaluated using subject-disjoint stratified 5-fold cross-validation. The StratifiedGroupKFold ensured subject identity so that no subject’s images appear in both the training and test partition of any fold. For CK+, all images belonging to a given subject are likewise constrained to a single fold. Each fold trains on approximately 80% and tests on the remaining approximately 20% of subjects’ images, with the exact per-fold counts varying slightly because grouping by subject does not yield perfectly equal partitions.
Data augmentation artificially increases the effective size of the training set by applying random transformations to existing images while preserving their semantic label [19]. For training images, the following augmentation pipeline is applied sequentially.
• Random resized crop to 224
• Random horizontal flip with probability 0.5 to provide mirror-image invariance.
• Random rotation within
• Color jitter with brightness ±0.30, contrast ±0.30, saturation ±0.20, hue ±0.06. This method simulated diverse illumination and camera conditions.
• Random grayscale conversion with probability 0.05 to encourage color-invariant features.
• Normalization using ImageNet statistics (mean [0.485, 0.456, 0.406], std [0.229, 0.224, 0.225]).
Evaluation images are only resized to 224
3.3 Proposed Dual-Stream FERNet Architecture
The proposed Dual-Stream FERNet model adopted a dual-backbone design with a sigmoid-gated attention fusion module, a shared projection head, and an EDL classification head. The architecture summary is given in Table 1. The EfficientNetV2-S backbone served as the speed-optimised stream. The original classification head was removed, and the backbone generated a 1280-dimensional feature vector

This formulation allowed the model to dynamically weight each branch’s contribution on a per-sample basis. Gate values average 0.675. EfficientNetV2-S gets favored over Swin-T about two to one. This held up on real test data across two separate experimental runs. Layer normalization was applied to
High
3.4 Self-Supervised Pre-Training (Independent Per-Branch Adaptation)
Before supervised fine-tuning, both backbone branches were jointly pre-trained for two epochs using a SimCLR-style contrastive objective. Each branch trains on its own during this step. There is no shared or joint objective between the two branches at this stage. For each mini-batch, two independently augmented views
AdamW with learning rate
3.5 Supervised Fine-Tuning and Training Configuration
Following SSL pre-training, the full model (both backbones, gated fusion, shared head, and EDL head) was trained in supervised mode. Training hyperparameters were held constant across both datasets to enable fair comparison. The supervised optimizer was AdamW with learning rate
The complete evidential loss is specified as follows. Given evidence
A natural concern with a 54-million-parameter dual-backbone model is over-parameterization relative to the size of the training partitions (approximately 2300–2410 images per fold for KDEF and 780–795 images per fold for CK+). This concern is substantially mitigated by several design choices. First, both EfficientNetV2-S and Swin-T backbones are initialized from ImageNet-1K pre-trained weights rather than random initialization. Second, four regularization mechanisms act simultaneously. (a) 50% Dropout in the shared projection head. (b) AdamW weight decay
This section presents and analyses the experimental results obtained from both datasets. Sections 4.1 and 4.2 report training convergence and final performance on KDEF and CK+, respectively. Section 4.5 compares these results against baseline models. Section 4.6 provides a comparison with state-of-the-art methods. Sections 4.7 and 4.8 cover Grad-CAM visualizations and computational efficiency.
The KDEF experiment used images across seven emotion classes, with subject-disjoint stratified 5-fold cross-validation. Fig. 2 shows loss and accuracy for all 5 folds. Thin lines are individual folds. The bold line is the mean. The fold-to-fold variability is visible directly. Mean test accuracy across the 5 folds is 93.84 ± 1.73% (macro-F1 93.82%, MCC 0.928). This is 3.67 points lower than the 97.51% obtained under an image-level, non-subject-disjoint protocol. This drop in accuracy is due to the subject leakage. KDEF has multiple images per subject, and an image-level split lets the same face appear in both train and test. That inflates accuracy. Per-fold accuracy ranged from 90.99% to 95.21% as given in Table 2. Fold 3 ran all 30 epochs and scored 93.52%, lower than Fold 5, which stopped at epoch 28 and scored 95.07%. Epoch count is not the main driver of fold-to-fold variance. Subject composition of each test fold is the more likely factor.

Figure 2: KDEF training and validation curves with mean across folds.

The confusion matrix on the KDEF test set, shown in Fig. 3, reveals that all seven emotion classes are classified with high per-class accuracy. Happy and neutral achieved the highest per-class accuracy, consistent with prior literature attributing strong performance on these categories to their distinctive and less ambiguous facial configurations. Fear and disgust show slightly lower individual class scores. This reflects the established difficulty to distinguish these two categories due to their overlapping action units.

Figure 3: Confusion matrix of proposed model using KDEF dataset (mean across 5 subject-disjoint folds).
The proposed model performance and mean evidential uncertainty of 0.340 given in Table 3 indicates that the model accumulated strong evidence for its predictions on the majority of test samples. The uncertainty distribution across KDEF test images is shown in Fig. 4a. Samples with high uncertainty predominantly correspond to the boundary cases between fear and disgust, and between sad and neutral, which is consistent with the psychological literatures on emotion ambiguity. The t-SNE projection of the 512-dimensional shared feature embeddings shown in Fig. 4b uses only out-of-fold, subject-disjoint embeddings, with seed 42, perplexity 30, 1000 iterations, and PCA initialization. Cluster separability is quantified directly rather than assessed visually.


Figure 4: Qualitative analysis on the KDEF dataset. (a) Evidential uncertainty distribution across test images, and (b) t-SNE visualisation of shared feature embeddings.
The CK+ experiment used 981 images across seven emotion classes, including contempt. The same subject-disjoint 5-fold protocol is used as KDEF. Fig. 5 shows a CK+ training run. The pattern is a general property of CK+’s small folds, seen across runs. Mean test accuracy across the 5 folds is 93.07 ± 1.22% (macro-F1 90.53%, MCC 0.917) as given in Table 4. This variance is smaller than KDEF’s (±1.73%). So a smaller dataset does not always mean higher fold-to-fold instability. One fold shows why weighted metrics alone fall short. In Fold 2, weighted F1 is 93.0% but macro-F1 is only 86.5%, a 6.5-point gap. This points to at least one minority class doing much worse than the larger classes in that subject split.

Figure 5: CK+ training and validation curves, all 5 subject-disjoint folds with mean across folds.

Mean evidential uncertainty on CK+ (0.419) given in Table 5 is higher than on KDEF. The values are consistent with the harder classification problem, the smaller training set, and the inclusion of the contempt class. The confusion matrix (Fig. 6) shows contempt as the weakest class by a clear margin (78% recall), with 13% of true contempt samples misclassified as sadness, the single largest off-diagonal error in the matrix. Sadness and anger are confusable in both directions, consistent with the overlapping brow- and mouth-related action units these two expressions share. All other classes exceed 89% recall, with happy (99%) and surprise (97%) the most reliably classified.


Figure 6: Confusion matrix of proposed model using CK+ dataset (mean across 5 subject-disjoint folds).
Fig. 7 presents a qualitative analysis of the proposed Dual-Stream FERNet on the CK+ dataset, using out-of-fold, subject-disjoint predictions and embeddings only. The uncertainty distribution is right-skewed, with most predictions concentrated below

Figure 7: Qualitative analysis on the CK+ dataset. (a) Evidential uncertainty distribution across test images, and (b) t-SNE visualisation of shared feature embeddings.
A well-calibrated classifier produces confidence scores that reflect the true probability of a correct prediction. This matters particularly in safety-critical FER applications, where the system must not only classify emotions but also reliably signal its own uncertainty. For the proposed model, this study reports Expected Calibration Error (ECE), Adaptive ECE, Negative Log-Likelihood (NLL), and the multiclass Brier score. These metrics are computed on out-of-fold predictions for both KDEF and CK+, rather than a single dataset and a single 15-bin ECE as in the original submission. ECE measures the average gap between predicted confidence and empirical accuracy across confidence bins, with lower values indicating better calibration. NLL measures the average log-loss of the predicted probability assignments. The multiclass Brier score computes the mean squared error between predicted probability vectors and one-hot encoded labels, where lower values indicate both better calibration and higher accuracy.
On KDEF out-of-fold predictions, the model reaches ECE


Figure 8: Reliability diagrams, out-of-fold predictions. (a) KDEF. (b) CK+.
To quantify the individual contribution of each proposed component, a four-variant ablation study is conducted on KDEF, with every variant trained on identical subject-disjoint folds and an identical seed (5 folds

4.5 Comparison with Baseline Models
Table 8 compares the proposed Dual-Stream FERNet against baseline models. All models use the same protocol, augmentation, ImageNet initialization, target-domain SSL pre-training, optimizer budget (AdamW,

4.6 Comparison with State-of-the-Art
Table 9 presents a comparative analysis of the proposed Dual-Stream FERNet against recent FER methods. No method listed in Table 9 was evaluated under an identical subject-disjoint split, class configuration, and evaluation protocol to the present study. The table is therefore presented as contextual only, and no claim of superiority over these methods is made. Under the corrected protocol, the proposed model achieves 93.84 ± 1.73% on KDEF and 93.07 ± 1.22% on CK+.
4.7 Grad-CAM Interpretability Analysis
This study computed Grad-CAM on the EfficientNetV2-S branch of the proposed model for KDEF. Grad-CAM needs a spatial (2D) activation map to compute gradients. After each backbone’s features are pooled and passed through the sigmoid gate and fusion layer, the fused representation is a flat vector with no spatial structure left. So Grad-CAM cannot be applied to the fused model directly. Standard practice for multi-branch models is to compute Grad-CAM separately at each backbone’s own last spatial layer. This study reports the EfficientNetV2-S branch. Fig. 9 shows the resulting heatmaps. Fear, sad, and surprise all concentrate on the central nose/mouth region. Surprise’s hotspot sits directly over the open mouth, a plausible match since the mouth is that expression’s most visible feature. Neutral gives a broad, diffuse activation over most of the central face, not a sharp hotspot, consistent with neutral having no single dominant action unit. Angry is the exception. Its hotspot sits near the eye/cheek boundary, not the central mouth/nose region the other four share, and it is more diffuse than the fear/sad/surprise maps.

Figure 9: KDEF Grad-CAM visualisations (EfficientNetV2-S branch).
Table 10 reports the full computational profile of the proposed Dual-Stream FERNet, measured on an NVIDIA GeForce RTX 4070 GPU (12 GB VRAM) with CUDA 12.1, PyTorch 2.5.1+cu121, in FP32 (single-precision floating point). Five warm-up batches are executed before timing begins to ensure GPU memory is fully allocated and caching effects have stabilised. The reported inference time includes image preprocessing (resize to

The proposed dual-stream model is more expensive than any single-backbone baseline on every axis measured. Its batch-1 median latency (18.96 ms) is roughly 1.6–4.6
This paper combines EfficientNetV2-S and Swin-Tiny through sigmoid-gated fusion, SimCLR-style dual-branch self-supervised pre-training, and Evidential Deep Learning for uncertainty-aware classification. On KDEF, the model reaches 93.84 ± 1.73% accuracy at 459.67 FPS (batch-1 median latency 18.96 ms). On CK+, it reaches 93.07 ± 1.22% accuracy. Grad-CAM on both datasets shows attention on facial regions tied to FACS action units. This is qualitative evidence only. It does not show the model causally relies on those regions. The model’s highest-uncertainty predictions fall on inter-class boundary cases, ones that are genuinely ambiguous even to people. CK+ performance also reflects the added difficulty of imbalanced minority classes such as contempt. These results position Dual-Stream FERNet as a controlled-setting baseline for future FER work on backbone fusion, SSL pre-training, and uncertainty modeling.
This study is limited to controlled, posed expressions. In-the-wild testing (RAF-DB, FER-2013, AffectNet, SFEW) is needed before any claim about specific deployments, such as driver monitoring or clinical mental-health screening. KDEF and CK+ are both controlled datasets with limited demographic diversity. Performance may not transfer across age, ethnicity, gender presentation, disability, or culture. Given this paper’s motivating discussion of clinical and driver-monitoring use, larger-scale evaluation is needed before any deployment. Future work has three priorities. First, evaluation on large-scale, unconstrained FER benchmarks (RAF-DB, AffectNet, FER-2013) to test generalization beyond controlled lab conditions. Second, continual learning methods for online adaptation without catastrophic forgetting. Third, extending SSL pre-training to 50–100 epochs on the unlabeled target domain. Fourth, testing SSL and the EDL head separately to see how much each one actually contributes.
Acknowledgement: Not applicable.
Funding Statement: This work is supported by Princess Nourah bint Abdulrahman University, Riyadh, Saudi Arabia through the Researchers Supporting Project PNURSP2026R333.
Author Contributions: The authors confirm contribution to the paper as follows. Rashid Jahangir handled formal analysis, investigation, methodology, and writing the original draft. Nazik Alturki handled visualization, methodology, writing review and editing, and funding acquisition. Mohammed Alreshoodi handled analysis, methodology, writing review and editing, and software. All authors reviewed and approved the final version of the manuscript.
Availability of Data and Materials: KDEF is available by request from Karolinska Institutet at https://kdef.se/download-2/. CK+ is available by request from the original maintainers at https://ckplus.jeffcohn.net.
Ethics Approval: KDEF was obtained directly from Karolinska Institutet. CK+ was obtained directly from the official CK+ distribution portal, request ID CKP-000027. Neither dataset involved new data collection from any person.
Conflicts of Interest: The authors declare no conflicts of interest.
References
1. Kaur M, Kumar M. Facial emotion recognition: a comprehensive review. Expert Syst. 2024;41(10):e13670. [Google Scholar]
2. Kopalidis T, Solachidis V, Vretos N, Daras P. Advances in facial expression recognition: a survey of methods, benchmarks, models, and datasets. Information. 2024;15(3):135. [Google Scholar]
3. Fei Z, Zhang B, Zhou W, Li X, Zhang Y, Fei M. Global multi-scale extraction and local mixed multi-head attention for facial expression recognition in the wild. Neurocomputing. 2025;622(3):129323. doi:10.1016/j.neucom.2024.129323. [Google Scholar] [CrossRef]
4. Li N, Huang Y, Wang Z, Fan Z, Li X, Xiao Z. Enhanced hybrid vision transformer with multi-scale feature integration and patch dropping for facial expression recognition. Sensors. 2024;24(13):4153. doi:10.3390/s24134153. [Google Scholar] [CrossRef]
5. Li Q, Fu K, Liu J, Li Y, Ren Q, Xu K, et al. Optimizing class imbalance in facial expression recognition using dynamic intra-class clustering. Biomimetics. 2025;10(5):296. doi:10.3390/biomimetics10050296. [Google Scholar] [CrossRef]
6. Nemavhola A, Chibaya C, Viriri S. A systematic review of CNN architectures, databases, performance metrics, and applications in face recognition. Information. 2025;16(2):107. doi:10.3390/info16020107. [Google Scholar] [CrossRef]
7. Song D, Liu C. A facial expression recognition network using hybrid feature extraction. PLoS One. 2025;20(1):e0312359. doi:10.1371/journal.pone.0312359. [Google Scholar] [CrossRef]
8. Vatcharaphrueksadee A, Maliyaem M, Sawakchart P. Hybrid models for facial emotion recognition and intensity detection: generalization across human and cartoon faces using CNN and vision transformer. J Adv Inform Technol. 2025;16(4):478–90. [Google Scholar]
9. Shao Z, Chen B, Zhou Y, Shi X, Li C, Ma L, et al. Constrained and directional ensemble attention for facial action unit detection. Pattern Recognit. 2026;169(11):111904. doi:10.1016/j.patcog.2025.111904. [Google Scholar] [CrossRef]
10. Ezzameli K. Vision transformer-based facial emotion recognition. IAENG Int J Comput Sci. 2026;53(1):410. [Google Scholar]
11. Wei LLX, Sani NS. Enhanced facial expression recognition based on ResNet50 with a convolutional block attention module. Int J Adv Comput Sci Appl. 2025;16(1):695–711. [Google Scholar]
12. Balachandran G, Ranjith S, Chenthil T, Jagan G. Facial expression-based emotion recognition across diverse age groups: a multi-scale vision transformer with contrastive learning approach. J Comb Optim. 2025;49(1):11. doi:10.1007/s10878-024-01241-8. [Google Scholar] [CrossRef]
13. Yan L, Yang J, Xia J, Gao R, Zhang L, Wan J, et al. Self-supervised extracted contrast network for facial expression recognition. Multimed Tools Appl. 2025;84(15):14977–96. doi:10.1007/s11042-024-19556-3. [Google Scholar] [CrossRef]
14. Zhu A, Jia X, Yang L, Zhou H, Su W. DUAL: a dual-stage approach for facial expression recognition based on contrastive learning. Int J Intell Syst. 2025;2025:7401168. [Google Scholar]
15. Yoon T, Kim H. Uncertainty estimation by density aware evidential deep learning. In: Proceedings of the 41st International Conference on Machine Learning; 2024 Jul 21–27; Vienna, Austria. p. 57217–43. [Google Scholar]
16. Li J, Zhou H, Qian Y, Dong Z, Wang SJ. Micro-expression recognition using dual-view self-supervised contrastive learning with intensity perception. Neurocomputing. 2025;619:129142. doi:10.2139/ssrn.4882306. [Google Scholar] [CrossRef]
17. Kim JH, Kim BG, Roy PP, Jeong DM. Efficient facial expression recognition algorithm based on hierarchical deep neural network structure. IEEE Access. 2019;7:41273–85. doi:10.1109/access.2019.2907327. [Google Scholar] [CrossRef]
18. Jain N, Kumar S, Kumar A, Shamsolmoali P, Zareapoor M. Hybrid deep neural networks for face emotion recognition. Pattern Recognit Lett. 2018;115(2):101–6. doi:10.1016/j.patrec.2018.04.010. [Google Scholar] [CrossRef]
19. Jain DK, Shamsolmoali P, Sehdev P. Extended deep neural network for facial emotion recognition. Pattern Recognit Lett. 2019;120:69–74. doi:10.1016/j.patrec.2019.01.008. [Google Scholar] [CrossRef]
20. Minaee S, Minaei M, Abdolrashidi A. Deep-emotion: facial expression recognition using attentional convolutional network. Sensors. 2021;21(9):3046. [Google Scholar]
21. Shao J, Qian Y. Three convolutional neural network models for facial expression recognition in the wild. Neurocomputing. 2019;355:82–92. doi:10.1016/j.neucom.2019.05.005. [Google Scholar] [CrossRef]
22. Umer S, Rout RK, Pero C, Nappi M. Facial expression recognition with trade-offs between data augmentation and deep learning features. J Ambient Intell Humaniz Comput. 2022;13(2):721–35. doi:10.1007/s12652-020-02845-8. [Google Scholar] [CrossRef]
23. Georgescu MI, Ionescu RT, Popescu M. Local learning with deep and handcrafted features for facial expression recognition. IEEE Access. 2019;7:64827–36. doi:10.1109/access.2019.2917266. [Google Scholar] [CrossRef]
24. Jeong M, Ko BC. Driver’s facial expression recognition in real-time for safe driving. Sensors. 2018;18(12):4270. doi:10.3390/s18124270. [Google Scholar] [CrossRef]
25. Khan M, Khan U, Awad M, Zaki N, Son G, Kwon S. Real-time emotion recognition system using adaptive distillation technique. Comput Model Eng Sci. 2026;147(1):1–10. doi:10.32604/cmes.2026.079697. [Google Scholar] [CrossRef]
26. Baghel N, Dubey SR, Singh SK. UpAttTrans: upscaled attention based transformer for facial image super-resolution. Image Vis Comput. 2025;163:105731. [Google Scholar]
27. Khan M, El Saddik A, Deriche M, Gueaieb W. STT-Net: simplified temporal transformer for emotion recognition. IEEE Access. 2024;12:86220–31. [Google Scholar]
28. Shu Y, Gu X, Yang GZ, Lo B. Revisiting self-supervised contrastive learning for facial expression recognition. arXiv:2210.03853. 2022. [Google Scholar]
29. Gao J, Chen M, Xiang L, Xu C. A comprehensive survey on evidential deep learning and its applications. IEEE Trans Pattern Anal Mach Intell. 2026;48(3):2118–38. doi:10.1109/tpami.2025.3625258. [Google Scholar] [CrossRef]
30. Goeleven E, De Raedt R, Leyman L, Verschuere B. The Karolinska directed emotional faces: a validation study. Cogn Emot. 2008;22(6):1094–118. doi:10.1080/02699930701626582. [Google Scholar] [CrossRef]
31. Lucey P, Cohn JF, Kanade T, Saragih J, Ambadar Z, Matthews I. The extended Cohn-Kanade dataset (ck+a complete dataset for action unit and emotion-specified expression. In: Proceedings of the 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition-Workshops; 2010 Jun 13–18; San Francisco, CA, USA. p. 94–101. [Google Scholar]
32. Kaur M, Kumar M, Dasi S. Fine-tuning MobileNetV2 for lightweight facial emotion recognition: FER-2013 and CK+ datasets. Natl Acad Sci Lett. 2026:1–8. [Google Scholar]
33. Ramirez-Quintana JA, Muñoz-Pacheco JJ, Ramirez-Alonso G, Medrano-Hermosillo JA, Corral-Saenz AD. Lightweight convolutional neural network with efficient channel attention mechanism for real-time facial emotion recognition in embedded systems. Sensors. 2025;25(23):7264. doi:10.3390/s25237264. [Google Scholar] [CrossRef]
34. Jayaraman S, Mahendran A. An interactive information based DCNN-BiLSTM model with dual attention mechanism for facial expression recognition. Sci Rep. 2025;15(1):26287. doi:10.1038/s41598-025-09709-1. [Google Scholar] [CrossRef]
35. Manal A, Alsulaiman FA. An ensemble learning approach for facial emotion recognition based on deep learning techniques. Electronics. 2025;14(17):3415. doi:10.3390/electronics14173415. [Google Scholar] [CrossRef]
Cite This Article
Copyright © 2026 The Author(s). Published by Tech Science Press.This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.


Submit a Paper
Propose a Special lssue
View Full Text
Download PDF
Downloads
Citation Tools