Open Access
ARTICLE
DDE-SER: A Dual-Decomposition Ensemble Framework Fusing Adaptive Variational Modes and Harmonic-Percussive Spectrograms for Speech Emotion Recognition
1 School of Computer Science, The University of Technology Sydney, 15 Broadway, Ultimo, NSW, Australia
2 School of Science and Technology, University of New England, Elm Avenue, Armidale, NSW, Australia
3 Faculty of Science and Technology, Charles Darwin University, 54 Cavenagh St., Darwin, NT, Australia
* Corresponding Author: David Hason Rudd. Email:
(This article belongs to the Special Issue: Deep Learning for Emotion Recognition)
Computers, Materials & Continua 2026, 89(1), 17 https://doi.org/10.32604/cmc.2026.084015
Received 15 April 2026; Accepted 18 June 2026; Issue published 13 August 2026
Abstract
The accurate classification of human emotions from speech remains a formidable challenge due to the dynamic, non-stationary properties of audio signals and pervasive background noise. Traditional single-domain extraction methods frequently fail to capture overlapping acoustic phenomena, resulting in high misclassification rates among acoustically similar emotions. To overcome this, we propose the Dual-Decomposition Ensemble (DDE-SER), an architecture that synergizes 1D adaptive frequency filtering with 2D spatial spectrogram separation. The framework operates through two distinct pipelines: an adaptive time-domain branch that leverages VGG-optiVMD to autonomously extract Intrinsic Mode Functions (IMFs), and a structural spectrogram branch that applies orthogonal median filtering to decouple continuous harmonic formants from transient percussive noises. The primary novel contribution is a trainable Gated Attention mechanism that dynamically fuses these two previously established orthogonal decomposition pipelines based on the underlying emotional context, mitigating feature redundancy without the parameter overhead of large foundation models. Additionally, an acoustic perturbation strategy involving targeted pitch shifts and noise injection is applied to prevent overfitting on pristine studio datasets. Validated using a rigorous Speaker-Independent Cross-Validation (SICV) methodology, DDE-SER demonstrates competitive predictive performance, achieving an overall accuracy of 85.30% on the EMO-DB corpus and 62.17% on the seven-class RAVDESS corpus. Diagnostic results confirm the framework partially mitigates ambiguities between high-arousal states, such as Anger and Happiness, while maintaining the lightweight computational efficiency required for affective computing deployments.Keywords
Cite This Article
Copyright © 2026 The Author(s). Published by Tech Science Press.This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.


Submit a Paper
Propose a Special lssue
View Full Text
Download PDF
Downloads
Citation Tools