Open Access
ARTICLE
DMHG-LEDS: Joint Differentiated Modality-Aware Heterogeneous Graph and Local Emotion Difference Supervision for Multimodal Emotion Recognition in Conversations
1 School of Computer Science and Artificial Intelligence, Hubei University of Technology, Wuhan, China
2 Hubei Provincial Key Laboratory of Green Intelligent Computing Power Network, Hubei University of Technology, Wuhan, China
3 Hubei Provincial Engineering Research Center for Digital & Intelligent Manufacturing Technologies and Applications, Hubei University of Technology, Wuhan, China
* Corresponding Author: Qun Zhang. Email:
Computers, Materials & Continua 2026, 89(2), 42 https://doi.org/10.32604/cmc.2026.087279
Received 15 June 2026; Accepted 24 July 2026; Issue published 15 September 2026
Abstract
Multimodal Emotion Recognition in Conversations (MERC) has garnered substantial research attention recently. Existing MERC methods face several challenges: (1) they apply shared or coarse-grained graph construction rules across modalities, overlooking their distinct dependency patterns; (2) they rely on fixed-activation MLPs for feature transformation, limiting nonlinear representation capacity in complex emotional scenarios; (3) they focus predominantly on contextual modeling while underexploring local emotion discrimination between related utterances. To address these issues, we propose Joint Differentiated Modality-Aware Heterogeneous Graph and Local Emotion Difference Supervision for Multimodal Emotion Recognition in Conversations (DMHG-LEDS), a novel MERC framework. Specifically, modality-aware heterogeneous graphs are constructed by assigning differentiated intra-modal connection strategies to text, visual, and audio modalities, enabling modality-dependent contextual relation modeling. Based on the resulting graph topology, graph convolution aggregates neighborhood information and Chebyshev-KAN subsequently performs adaptive nonlinear transformation within each propagation step. Finally, an Entropy-Gated Local Emotion Difference Supervision module is introduced as an auxiliary task. It constructs ordered utterance pairs within non-overlapping local segments and provides entropy-gated supervision based on emotion-label differences, thereby improving the discriminability of local utterance representations. Experiments on Interactive emotional dyadic motion capture database (IEMOCAP), A Multimodal Multi-Party Dataset for Emotion Recognition in Conversation (MELD), and Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph (CMU-MOSEI) demonstrate the effectiveness of the proposed method, which outperforms all baselines.Keywords
Cite This Article
Copyright © 2026 The Author(s). Published by Tech Science Press.This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.


Submit a Paper
Propose a Special lssue
View Full Text
Download PDF
Downloads
Citation Tools