Muhammad Naeem Zafar1, Yunfei Yin1,*, Junaid Abbas2, Bayan Alabdullah3, Khaled Alnowaiser4, Yunyoung Nam5, Zepa Yang5,*
CMC-Computers, Materials & Continua, Vol.89, No.1, 2026, DOI:10.32604/cmc.2026.085136
- 13 August 2026
Abstract Accurate medical image classification increasingly relies on the joint modeling of global contextual semantics and fine-grained local structural cues, since many lesions are only reliably recognized when subtle local details are interpreted within their broader anatomical context. However, most recent hybrid CNN–Transformer and global–local frameworks still extract these features in separate streams and merge them only through late-stage static fusion, without explicit bidirectional interaction during representation learning. As a result, global context cannot effectively guide the refinement of subtle local structures, and local discriminative cues cannot recalibrate higher-level semantic reasoning before classification, which limits reciprocal… More >