Open Access
ARTICLE
Multiscale Long-Distance Feature Aggregation Network for Geospatial Semantic Segmentation in High-Resolution Remote Sensing Imagery
1 School of Automation, University of Electronic Science and Technology of China, Chengdu, China
2 School of the Environment, The University of Queensland, St Lucia, QLD, Australia
3 Future Tech Institute, Guangzhou Huashang University, Guangzhou, China
4 School of Biological and Environmental Engineering, Xi’an University, Xi’an, China
* Corresponding Author: Wenfeng Zheng. Email:
(This article belongs to the Special Issue: Recent Advances in Geospatial Artificial Intelligence (GeoAI) Models, Approaches, and Applications)
Computer Modeling in Engineering & Sciences 2026, 148(1), 32 https://doi.org/10.32604/cmes.2026.085484
Received 12 May 2026; Accepted 25 June 2026; Issue published 27 July 2026
Abstract
High-resolution remote sensing semantic segmentation is a fundamental task in Geospatial Artificial Intelligence (GeoAI). Existing CNN-based methods are effective for local and multiscale feature extraction but often lack progressive cross-scale semantic propagation, while attention- and Transformer-based methods improve global spatial modeling but generally ignore frequency-domain regularities. To address these limitations, this study proposes a Multiscale Long-Distance Feature Aggregation Network (MLFANet), a unified spatial-frequency segmentation framework for high-resolution remote sensing imagery. MLFANet introduces three key components: a Multiscale Global Dependency Extraction module for cascaded cross-scale contextual refinement, an FFT-based frequency-domain branch with learnable global filtering for capturing structural and texture regularities, and a bidirectional Spatial-Frequency Fusion module for adaptively aligning spatial details with frequency responses. Experiments on the ISPRS Potsdam and Vaihingen datasets demonstrate the effectiveness and feasibility of the proposed model. MLFANet achieves AF, MIoU, and OA values of 86.03%, 76.21%, and 88.70% on Potsdam, and 83.17%, 71.90%, and 86.33% on Vaihingen, respectively, outperforming representative CNN-based, attention-based, and hybrid models in overall metrics. In terms of computational complexity, MLFANet requires 17.49 G FLOPs under an input size of 256 × 256 pixels, indicating its practical feasibility for patch-based high-resolution remote sensing segmentation. Ablation studies further verify that multiscale dependency extraction, frequency-domain modeling, and adaptive spatial-frequency fusion each contribute to the final performance.Keywords
Cite This Article
Copyright © 2026 The Author(s). Published by Tech Science Press.This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.


Submit a Paper
Propose a Special lssue
View Full Text
Download PDF
Downloads
Citation Tools