Home / Journals / CMC / Online First / doi:10.32604/cmc.2026.084015
Special Issues
Table of Content

Open Access

ARTICLE

DDE-SER: A Dual-Decomposition Ensemble Framework Fusing Adaptive Variational Modes and Harmonic-Percussive Spectrograms for Speech Emotion Recognition

David Hason Rudd1,*, Cesar Sanin2, Md Rafiqul Islam3, Xianzhi Wang1, Huan Huo1
1 School of Computer Science, The University of Technology Sydney, 15 Broadway, Ultimo, NSW, Australia
2 School of Science and Technology, University of New England, Elm Avenue, Armidale, NSW, Australia
3 Faculty of Science and Technology, Charles Darwin University, 54 Cavenagh St., Darwin, NT, Australia
* Corresponding Author: David Hason Rudd. Email: email
(This article belongs to the Special Issue: Deep Learning for Emotion Recognition)

Computers, Materials & Continua https://doi.org/10.32604/cmc.2026.084015

Received 15 April 2026; Accepted 18 June 2026; Published online 24 July 2026

Abstract

The accurate classification of human emotions from speech remains a formidable challenge due to the dynamic, non-stationary properties of audio signals and pervasive background noise. Traditional single-domain extraction methods frequently fail to capture overlapping acoustic phenomena, resulting in high misclassification rates among acoustically similar emotions. To overcome this, we propose the Dual-Decomposition Ensemble (DDE-SER), an architecture that synergizes 1D adaptive frequency filtering with 2D spatial spectrogram separation. The framework operates through two distinct pipelines: an adaptive time-domain branch that leverages VGG-optiVMD to autonomously extract Intrinsic Mode Functions (IMFs), and a structural spectrogram branch that applies orthogonal median filtering to decouple continuous harmonic formants from transient percussive noises. The primary novel contribution is a trainable Gated Attention mechanism that dynamically fuses these two previously established orthogonal decomposition pipelines based on the underlying emotional context, mitigating feature redundancy without the parameter overhead of large foundation models. Additionally, an acoustic perturbation strategy involving targeted pitch shifts and noise injection is applied to prevent overfitting on pristine studio datasets. Validated using a rigorous Speaker-Independent Cross-Validation (SICV) methodology, DDE-SER demonstrates competitive predictive performance, achieving an overall accuracy of 85.30% on the EMO-DB corpus and 62.17% on the seven-class RAVDESS corpus. Diagnostic results confirm the framework partially mitigates ambiguities between high-arousal states, such as Anger and Happiness, while maintaining the lightweight computational efficiency required for affective computing deployments.

Keywords

Acoustic emotion recognition; variational mode decomposition; harmonic-percussive separation; feature fusion; ensemble learning; convolutional neural networks
  • 250

    View

  • 31

    Download

  • 0

    Like

Share Link