Home / Journals / CMC / Online First / doi:10.32604/cmc.2026.085456
Special Issues
Table of Content

Open Access

ARTICLE

A Two-Stage Decoupled Matching Network for Multimodal Entity Linking

Huayu Li1, Xiang Wang1, Jia Luo2,3,4,*, Xiaotong He1, Peiying Zhang1
1 Qingdao Institute of Software, College of Computer Science and Technology, China University of Petroleum (East China), Qingdao, China
2 Interdisciplinary Faculty of Science and Engineering, Shimane University, Shimane, Japan
3 College of Economics and Management, Beijing University of Technology, Beijing, China
4 Chongqing Research Institute, Beijing University of Technology, Chongqing, China
* Corresponding Author: Jia Luo. Email: email
(This article belongs to the Special Issue: The Next-generation Deep Learning Approaches to Emerging Real-world Applications, 2nd Edition)

Computers, Materials & Continua https://doi.org/10.32604/cmc.2026.085456

Received 11 May 2026; Accepted 01 July 2026; Published online 27 July 2026

Abstract

Multimodal Entity Linking (MEL) aims to map ambiguous mentions in multimodal contexts to their corresponding entities in a multimodal knowledge base. However, existing methods still face limitations in terms of feature extraction granularity, the depth of cross-modal interaction, and architectural coupling. To address these issues, we propose a Two-stage Decoupled Matching Network (TDMN) for multimodal entity linking. The matching process is divided into two stages: intra-modal matching and cross-modal interaction. In the intra-modal stage, textual and visual inputs are processed independently. The framework then proceeds to the cross-modal interaction stage, following the principle of “enhancement prior to interaction.” Specifically, unimodal features are first refined through a parallel dual-attention network consisting of Global Relational Attention and Adaptive Sharpening Attention, together with a multi-granularity calibration fusion module. Based on the refined representations, cross-modal alignment is subsequently performed within a symmetric bidirectional interaction architecture, in which a gated residual mechanism is introduced to facilitate information fusion. Experiments conducted on the public benchmark datasets WikiMEL and WikiDiverse demonstrate the effectiveness of TDMN. Compared with the M3EL baseline, TDMN achieves absolute improvements of 1.39% and 1.88% in MRR and Hits@1, respectively, on the WikiDiverse dataset. In addition, compared with MIMIC, TDMN improves MRR and Hits@1 by 0.8% and 1.21%, respectively, on the WikiMEL dataset. These results support the effectiveness of the proposed approach.

Keywords

Multimodal entity linking; multi-granularity feature fusion; attention mechanism; feature enhancement; multimodal representation learning
  • 133

    View

  • 22

    Download

  • 0

    Like

Share Link