Open Access
ARTICLE
Scale Ladder Consistency for Structure-Aware Multimodal Representation Learning in 3D Medical Image Segmentation
1 School of Automation Science and Engineering, South China University of Technology, Guangzhou, China
2 Department of Otolaryngology, Head and Neck Surgery, Guangdong Provincial People’s Hospital (Guangdong Academy of Medical Sciences), Southern Medical University, Guangzhou, China
* Corresponding Author: Bin Li. Email:
# These authors contributed equally to this work
(This article belongs to the Special Issue: Emerging Artificial Intelligence Technologies and Applications-II)
Computer Modeling in Engineering & Sciences 2026, 148(2), 39 https://doi.org/10.32604/cmes.2026.087647
Received 20 June 2026; Accepted 12 August 2026; Issue published 28 August 2026
Abstract
Self-supervised representation learning can reduce the dependence of three-dimensional (3D) medical image segmentation on dense voxel annotations. In multimodal 3D medical imaging, intensity-reconstruction pre-training provides dense appearance supervision but does not explicitly distinguish the structural regions that determine segmentation boundaries and small targets. A second mismatch arises in scale learning: encoder-decoder networks provide multi-scale feature maps, but they do not explicitly supervise how fine anatomical structures weaken or persist across neighboring scales. To address these mismatches, this study proposes Scale Ladder Consistency (SLC), a structure-aware self-supervised representation learning framework for multimodal 3D medical image segmentation. SLC combines Scale-Space Structural Reconstruction (SSR), Hybrid Mask, and Scale Ladder (SL) in its structural pre-training path. SSR replaces intensity recovery with structure prediction, Hybrid Mask increases supervision on fine structural regions, and SL learns neighboring-scale structural transitions through bidirectional prediction. Subset-to-Full Regularization (S2F) further stabilizes case-level representations during pre-training. Experimental results on the Brain Tumor Segmentation 2019 (BraTS19) and carotid artery datasets demonstrate that SLC consistently outperforms matched scratch fine-tuning and achieves competitive performance against recent segmentation and self-supervised methods. These results indicate that structure-aware and cross-scale self-supervised objectives can provide effective representations for multimodal 3D medical segmentation.Keywords
Cite This Article
Copyright © 2026 The Author(s). Published by Tech Science Press.This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.


Submit a Paper
Propose a Special lssue
View Full Text
Download PDF
Downloads
Citation Tools