Home / Journals / CMC / Online First / doi:10.32604/cmc.2026.085838
Special Issues
Table of Content

Open Access

ARTICLE

Semantic Context-Aware Multi-Scale Vision Transformer for UAV Disaster Scene Classification and Uncertainty-Aware Understanding

Hadeel Alsolai1, Muhammad Waqas Ahmed2, Bayan Alabdullah1, Fatimah Alhayan1, Mohammed Alonazi3, Ahmad Jalal4,5, Jeongmin Park6,*
1 Department of Information Systems, College of Computer and Information Sciences, Princess Nourah bint Abdulrahman University, Riyadh, Saudi Arabia
2 Department of Computer Games Development, Air University, Islamabad, Pakistan
3 Department of Information Systems, College of Computer Engineering and Sciences, Prince Sattam bin Abdulaziz University, Al-Kharj, Saudi Arabia
4 Department of Computer Science, Air University, Islamabad, Pakistan
5 Department of Computer Science and Engineering, College of Informatics, Korea University, Seoul, Republic of Korea
6 Department of Computer Engineering, Tech University of Korea, 237 Sangidaehak-Ro, Siheung-Si, Gyeonggi-Do, Republic of Korea
* Corresponding Author: Jeongmin Park. Email: email
(This article belongs to the Special Issue: Advances in Intelligent Video Object Tracking and Scene Understanding)

Computers, Materials & Continua https://doi.org/10.32604/cmc.2026.085838

Received 19 May 2026; Accepted 09 July 2026; Published online 17 August 2026

Abstract

Robust scene-level classification and semantic understanding from aerial and disaster-related imagery are essential for intelligent vision systems deployed in emergency response, UAV-based monitoring, and safety-critical environments. However, existing deep learning approaches, including convolutional neural networks and Vision Transformers (ViTs), often struggle to simultaneously capture fine-grained local object characteristics and global semantic scene context, while also lacking reliable uncertainty estimation mechanisms for trustworthy decision-making. To address these limitations, this paper proposes MS-SLCA-ViT, a novel multi-scale scene–local cross-attention Vision Transformer framework for robust and uncertainty-aware image scene understanding. The proposed architecture introduces three major contributions. First, a Multi-Scale Patch Tokenizer (MSPT) extracts hierarchical semantic representations at multiple spatial resolutions and adaptively integrates them through a learnable cross-scale attention mechanism to enhance contextual feature modeling. Second, a Scene–Local Cross-Attention (SLCA) module employs dual-stream transformer processing with bidirectional attention to model interactions between global scene semantics and discriminative local object regions, improving semantic context understanding in complex environments. Third, an uncertainty-aware inference strategy based on Monte Carlo Dropout generates calibrated confidence estimates alongside predictions, enabling more reliable and interpretable decision-making for intelligent visual analytics. Extensive experiments conducted on a public disaster image dataset demonstrate that the proposed framework achieves a top-1 classification accuracy of 93.7%, outperforming state-of-the-art approaches, including ViT, Swin Transformer, DeiT, ResNet-50, and EfficientNet-B0. Furthermore, the uncertainty modeling strategy significantly improves prediction reliability and confidence calibration under challenging environmental conditions. The experimental findings demonstrate that the proposed framework provides an effective and generalizable solution for robust scene understanding, semantic context modeling, and intelligent image analysis in UAV-assisted and safety-critical computer vision applications.

Keywords

UAV video scene understanding; semantic context modeling; vision transformer; multi-scale feature representation; transformer-based perception; intelligent aerial surveillance; safety-critical visual analytics
  • 85

    View

  • 24

    Download

  • 0

    Like

Share Link