Open Access iconOpen Access

ARTICLE

A Two-Stage Decoupled Matching Network for Multimodal Entity Linking

Huayu Li1, Xiang Wang1, Jia Luo2,3,4,*, Xiaotong He1, Peiying Zhang1

1 Qingdao Institute of Software, College of Computer Science and Technology, China University of Petroleum (East China), Qingdao, China
2 Interdisciplinary Faculty of Science and Engineering, Shimane University, Shimane, Japan
3 College of Economics and Management, Beijing University of Technology, Beijing, China
4 Chongqing Research Institute, Beijing University of Technology, Chongqing, China

* Corresponding Author: Jia Luo. Email: email

(This article belongs to the Special Issue: The Next-generation Deep Learning Approaches to Emerging Real-world Applications, 2nd Edition)

Computers, Materials & Continua 2026, 89(1), 61 https://doi.org/10.32604/cmc.2026.085456

Abstract

Multimodal Entity Linking (MEL) aims to map ambiguous mentions in multimodal contexts to their corresponding entities in a multimodal knowledge base. However, existing methods still face limitations in terms of feature extraction granularity, the depth of cross-modal interaction, and architectural coupling. To address these issues, we propose a Two-stage Decoupled Matching Network (TDMN) for multimodal entity linking. The matching process is divided into two stages: intra-modal matching and cross-modal interaction. In the intra-modal stage, textual and visual inputs are processed independently. The framework then proceeds to the cross-modal interaction stage, following the principle of “enhancement prior to interaction.” Specifically, unimodal features are first refined through a parallel dual-attention network consisting of Global Relational Attention and Adaptive Sharpening Attention, together with a multi-granularity calibration fusion module. Based on the refined representations, cross-modal alignment is subsequently performed within a symmetric bidirectional interaction architecture, in which a gated residual mechanism is introduced to facilitate information fusion. Experiments conducted on the public benchmark datasets WikiMEL and WikiDiverse demonstrate the effectiveness of TDMN. Compared with the M3EL baseline, TDMN achieves absolute improvements of 1.39% and 1.88% in MRR and Hits@1, respectively, on the WikiDiverse dataset. In addition, compared with MIMIC, TDMN improves MRR and Hits@1 by 0.8% and 1.21%, respectively, on the WikiMEL dataset. These results support the effectiveness of the proposed approach.

Keywords

Multimodal entity linking; multi-granularity feature fusion; attention mechanism; feature enhancement; multimodal representation learning

Cite This Article

APA Style
Li, H., Wang, X., Luo, J., He, X., Zhang, P. (2026). A Two-Stage Decoupled Matching Network for Multimodal Entity Linking. Computers, Materials & Continua, 89(1), 61. https://doi.org/10.32604/cmc.2026.085456
Vancouver Style
Li H, Wang X, Luo J, He X, Zhang P. A Two-Stage Decoupled Matching Network for Multimodal Entity Linking. Comput Mater Contin. 2026;89(1):61. https://doi.org/10.32604/cmc.2026.085456
IEEE Style
H. Li, X. Wang, J. Luo, X. He, and P. Zhang, “A Two-Stage Decoupled Matching Network for Multimodal Entity Linking,” Comput. Mater. Contin., vol. 89, no. 1, pp. 61, 2026. https://doi.org/10.32604/cmc.2026.085456



cc Copyright © 2026 The Author(s). Published by Tech Science Press.
This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
  • 172

    View

  • 49

    Download

  • 0

    Like

Share Link