BIAC-Net: Bidirectional Global-Local Communication for Feature Refinement in Medical Image Classification
Muhammad Naeem Zafar1, Yunfei Yin1,*, Junaid Abbas2, Bayan Alabdullah3, Khaled Alnowaiser4, Yunyoung Nam5, Zepa Yang5,*
1 College of Computer Science, Chongqing University, Shapingba, Chongqing, China
2 School of Big Data and Software Engineering, Chongqing University, Shapingba, Chongqing, China
3 Department of Information Systems, College of Computer and Information Sciences, Princess Nourah bint Abdulrahman University, Riyadh, Saudi Arabia
4 Department of Computer Engineering, College of Computer Engineering and Sciences, Prince Sattam Bin Abdulaziz University, Al-Kharj, Saudi Arabia
5 Department of Computer Science and Engineering, Soonchunhyang University, Asan, Republic of Korea
* Corresponding Author: Yunfei Yin. Email:
; Zepa Yang. Email:
Computers, Materials & Continua https://doi.org/10.32604/cmc.2026.085136
Received 06 May 2026; Accepted 29 June 2026; Published online 20 July 2026
Abstract
Accurate medical image classification increasingly relies on the joint modeling of global contextual semantics and fine-grained local structural cues, since many lesions are only reliably recognized when subtle local details are interpreted within their broader anatomical context. However, most recent hybrid CNN–Transformer and global–local frameworks still extract these features in separate streams and merge them only through late-stage static fusion, without explicit bidirectional interaction during representation learning. As a result, global context cannot effectively guide the refinement of subtle local structures, and local discriminative cues cannot recalibrate higher-level semantic reasoning before classification, which limits reciprocal feature refinement and undermines robust integration of contextual and structural information in challenging medical images. To address this limitation, we propose BIAC-Net, a bidirectional global–local feature communication network for medical image classification. BIAC-Net employs a Vision Transformer (ViT) backbone to capture contextual representations and introduces two complementary refinement branches that enhance global saliency and preserve local structural details. Unlike conventional hybrid designs that merge features directly, the proposed framework introduces a bidirectional cross-stream residual refinement mechanism between global and local representations, allowing contextual features to guide local refinement while local structural cues recalibrate global reasoning before fusion. The reciprocally refined representations are then integrated through an adaptive gated fusion mechanism that learns input-dependent feature weighting. Experimental results on the Kvasir and ISIC 2018 datasets demonstrate that BIAC-Net achieves 97.84% and 90.91% accuracy, respectively, outperforming several representative baseline models. Ablation studies further support the contribution of bidirectional communication, while Gradient-weighted Class Activation Mapping (Grad-CAM) visualizations provide preliminary qualitative evidence that the model attends to diagnostically relevant image regions.
Keywords
Medical image classification; vision transformer; global-local feature interaction; bidirectional attention; adaptive gated fusion; explainable AI