Open Access iconOpen Access

ARTICLE

A Comparative Study of Audio-Language Models for Speech Emotion Recognition in Spanish

Jorge Gómez-Navalón, Ronghao Pan, Tomas Bernal-Beltrán, José Antonio García-Díaz*, Rafael Valencia-García

Informatics and Systems, Universidad de Murcia, Murcia, Spain

* Corresponding Author: José Antonio García-Díaz. Email: email

(This article belongs to the Special Issue: Applied NLP with Large Language Models: AI Applications Across Domains)

Computer Modeling in Engineering & Sciences 2026, 148(1), 38 https://doi.org/10.32604/cmes.2026.085437

Abstract

Traditionally, speech emotion recognition has relied on supervised models that require task-specific training and annotated data. However, the recent emergence of audio-language models introduces a more flexible paradigm that enables multimodal reasoning through speech and natural language interaction. Nevertheless, their effectiveness for emotion recognition remains unclear. In this study, we evaluate audio-language models for speech emotion classification using the Spanish MEACorpus dataset and compare three approaches: prompt-based inference, embedding-based classification with lightweight classifiers, and instruction-tuned models with parameter-efficient fine-tuning plus a hybrid architecture based on class-specific confidence-driven routing. Our results show that the hybrid approach achieves the highest overall performance, reaching an 83.55% macro F1-score and an 84.37% weighted F1-score. Instruction tuning remains highly competitive, obtaining an 83.32% macro F1-score, which confirms the importance of supervised task adaptation for aligning ALMs with speech emotion recognition. We also include Whisper as a pretrained acoustic baseline to contextualize ALM-based representations against strong speech foundation models. Furthermore, the hybrid approach outperforms standalone embedding-based classification across all evaluated models, showing that class-specific confidence-driven routing can improve the use of embedding-based predictions. Although the best hybrid ALM configuration achieves competitive performance, it remains below the specialized MEACorpus baseline of 87.74% macro F1-score, indicating that general-purpose ALMs do not yet surpass highly specialized acoustic models for Spanish speech emotion recognition.

Keywords

Speech emotion recognition; audio-language models; multimodal learning; instruction tuning; prompting; embedding-based classification

Cite This Article

APA Style
Gómez-Navalón, J., Pan, R., Bernal-Beltrán, T., García-Díaz, J.A., Valencia-García, R. (2026). A Comparative Study of Audio-Language Models for Speech Emotion Recognition in Spanish. Computer Modeling in Engineering & Sciences, 148(1), 38. https://doi.org/10.32604/cmes.2026.085437
Vancouver Style
Gómez-Navalón J, Pan R, Bernal-Beltrán T, García-Díaz JA, Valencia-García R. A Comparative Study of Audio-Language Models for Speech Emotion Recognition in Spanish. Comput Model Eng Sci. 2026;148(1):38. https://doi.org/10.32604/cmes.2026.085437
IEEE Style
J. Gómez-Navalón, R. Pan, T. Bernal-Beltrán, J. A. García-Díaz, and R. Valencia-García, “A Comparative Study of Audio-Language Models for Speech Emotion Recognition in Spanish,” Comput. Model. Eng. Sci., vol. 148, no. 1, pp. 38, 2026. https://doi.org/10.32604/cmes.2026.085437



cc Copyright © 2026 The Author(s). Published by Tech Science Press.
This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
  • 0

    View

  • 0

    Download

  • 0

    Like

Share Link