Open Access iconOpen Access

ARTICLE

Multimodal Emotion Recognition in Urdu through Late Fusion of Fine-Tuned Speech and Text Representations

Muhammad Sheraz1, Adil Majeed1, Shehzad Khalid2,3,*, Yazeed Alkhrijah4,*, Sulieman S. Alshuhri5, Hasan Mujtaba1

1 Department of Artificial Intelligence and Data Science, National University of Computer & Emerging Sciences, AK Brohi Rd., H-11/4, Islamabad, Pakistan
2 Computer and Information Sciences Research Center (CISRC), Imam Mohammad Ibn Saud Islamic University, Riyadh, Saudi Arabia
3 Department of Computer Engineering, Bahria School of Engineering and Applied Sciences (BSEAS), Bahria University, Islamabad, Pakistan
4 Department of Electrical Engineering, Imam Mohammad ibn Saud Islamic University (IMSIU), Riyadh, Saudi Arabia
5 Department of Information Technology, College of Computer and Information Sciences, Imam Mohammad ibn Saud Islamic University (IMSIU), Riyadh, Saudi Arabia

* Corresponding Authors: Shehzad Khalid. Email: email; Yazeed Alkhrijah. Email: email

(This article belongs to the Special Issue: Machine Learning and Deep Learning-Based Pattern Recognition, 2nd Edition)

Computer Modeling in Engineering & Sciences 2026, 148(2), 41 https://doi.org/10.32604/cmes.2026.086256

Abstract

Emotion recognition plays a crucial role in enabling intelligent human–computer interaction, yet research in low-resource languages such as Urdu remains limited, particularly in multimodal settings. This study proposes a multimodal deep learning framework for Urdu emotion recognition by integrating speech and text modalities. The approach leverages transformer-based models, namely wav2vec 2.0 for audio representation and MuRIL for text representation, combined using a late fusion strategy for classification. Experiments were conducted on the UMED dataset, consisting of 8269 multimodal instances across five emotion classes. The proposed multimodal model achieved an accuracy of 0.701 and an F1-score of 0.6915, outperforming unimodal baselines, where the audio-only and text-only models achieved accuracies of 0.6681 and 0.5085, respectively. Furthermore, the proposed approach surpasses the existing UMEDNet benchmark, demonstrating the effectiveness of transformer-based feature extraction and multimodal late fusion for Urdu emotion recognition. The results highlight the complementary nature of speech and text modalities and demonstrate that independently learned modality-specific classifiers combined through decision-level fusion can improve emotion recognition performance in low-resource languages. However, the performance improvement over alternative fusion strategies was relatively modest, indicating that more advanced multimodal interaction mechanisms may further enhance recognition performance.

Keywords

Multimodal; emotion detection; audio emotion detection; text emotion detection

Cite This Article

APA Style
Sheraz, M., Majeed, A., Khalid, S., Alkhrijah, Y., Alshuhri, S.S. et al. (2026). Multimodal Emotion Recognition in Urdu through Late Fusion of Fine-Tuned Speech and Text Representations. Computer Modeling in Engineering & Sciences, 148(2), 41. https://doi.org/10.32604/cmes.2026.086256
Vancouver Style
Sheraz M, Majeed A, Khalid S, Alkhrijah Y, Alshuhri SS, Mujtaba H. Multimodal Emotion Recognition in Urdu through Late Fusion of Fine-Tuned Speech and Text Representations. Comput Model Eng Sci. 2026;148(2):41. https://doi.org/10.32604/cmes.2026.086256
IEEE Style
M. Sheraz, A. Majeed, S. Khalid, Y. Alkhrijah, S. S. Alshuhri, and H. Mujtaba, “Multimodal Emotion Recognition in Urdu through Late Fusion of Fine-Tuned Speech and Text Representations,” Comput. Model. Eng. Sci., vol. 148, no. 2, pp. 41, 2026. https://doi.org/10.32604/cmes.2026.086256



cc Copyright © 2026 The Author(s). Published by Tech Science Press.
This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
  • 291

    View

  • 66

    Download

  • 0

    Like

Share Link