Open Access
ARTICLE
Attention-Guided Cross-Modal Transformer for Multimodal SAR-Optical Image Fusion and Flood Change Detection
1 Department of Information Systems, College of Computer and Information Sciences, Princess Nourah bint Abdulrahman University, Riyadh, Saudi Arabia
2 Department of Computer Games Development, Air University, E-9, Islamabad, Pakistan
3 Department of Software Engineering, College of Engineering and Advanced Computing, Alfaisal University, Riyadh, Saudi Arabia
4 Department of Computer Science, College of Computers and Information Technology, Taif University, Taif, Saudi Arabia
5 Department of Information Systems, College of Computer Engineering and Sciences, Prince Sattam bin Abdulaziz University, Al-Kharj, Saudi Arabia
6 Faculty of Computing and AI, Air University, E-9, Islamabad, Pakistan
7 Department of Computer Science and Engineering, College of Informatics, Korea University, Seoul, Republic of Korea
* Corresponding Authors: Mohammad Shorfuzzaman. Email: ; Ahmad Jalal. Email:
(This article belongs to the Special Issue: Multimodal Image Analysis, Data Fusion and Artificial Intelligence for Complex Visual and Material Data)
Computers, Materials & Continua 2026, 89(1), 88 https://doi.org/10.32604/cmc.2026.086985
Received 08 June 2026; Accepted 13 July 2026; Issue published 13 August 2026
Abstract
Multimodal data fusion and deep learning have opened new frontiers in the analysis of complex visual data acquired from heterogeneous sensing systems. Flood inundation mapping represents one of the most demanding applications in this domain, requiring robust interpretation of complementary but conflicting image modalities under severe real-world constraints. This paper presents CAG-Transformer, a novel multimodal AI architecture for bi-temporal flood change detection through intelligent fusion of Sentinel-1 SAR and Sentinel-2 multispectral imagery. Three tightly integrated contributions address the core challenges of heterogeneous multimodal image analysis. A Change Attention Gate (CAG) performs adaptive channel-wise representation learning, selectively amplifying flood-relevant spectral and backscatter variations while suppressing temporally static scene content. A Cross-Modal Transformer (CMT) bottleneck employs multi-head self-attention to model long-range spatial dependencies and enable context-aware reasoning across heterogeneous sensor representations capabilities fundamentally beyond convolutional fusion. A differentiable soft Dice loss ensures stable gradient flow under the severe class imbalance inherent to real-world flood datasets. Evaluated on the Ombria multimodal benchmark, CAG-Transformer achieves IoU = 0.7774, Dice = 0.8894, and AUC-ROC = 0.9768, outperforming single-modality and conventional fusion baselines. Cross-event validation on Albania (IoU = 0.7647) and Timor (IoU = 0.7313) confirms generalization across environmentally and spatially distinct flood. With 3.5 million parameters and sub-100 ms inference on freely available Copernicus data, the framework offers an efficient and deployable solution for near-real-time flood monitoring and multimodal image-based environmental assessment.Keywords
Cite This Article
Copyright © 2026 The Author(s). Published by Tech Science Press.This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.


Submit a Paper
Propose a Special lssue
View Full Text
Download PDF
Downloads
Citation Tools