Open Access iconOpen Access

ARTICLE

Scale Ladder Consistency for Structure-Aware Multimodal Representation Learning in 3D Medical Image Segmentation

Weiqing Liu1,#, Bin Li1,#,*, Lianfang Tian1, Qianhui Qiu2

1 School of Automation Science and Engineering, South China University of Technology, Guangzhou, China
2 Department of Otolaryngology, Head and Neck Surgery, Guangdong Provincial People’s Hospital (Guangdong Academy of Medical Sciences), Southern Medical University, Guangzhou, China

* Corresponding Author: Bin Li. Email: email
# These authors contributed equally to this work

(This article belongs to the Special Issue: Emerging Artificial Intelligence Technologies and Applications-II)

Computer Modeling in Engineering & Sciences 2026, 148(2), 39 https://doi.org/10.32604/cmes.2026.087647

Abstract

Self-supervised representation learning can reduce the dependence of three-dimensional (3D) medical image segmentation on dense voxel annotations. In multimodal 3D medical imaging, intensity-reconstruction pre-training provides dense appearance supervision but does not explicitly distinguish the structural regions that determine segmentation boundaries and small targets. A second mismatch arises in scale learning: encoder-decoder networks provide multi-scale feature maps, but they do not explicitly supervise how fine anatomical structures weaken or persist across neighboring scales. To address these mismatches, this study proposes Scale Ladder Consistency (SLC), a structure-aware self-supervised representation learning framework for multimodal 3D medical image segmentation. SLC combines Scale-Space Structural Reconstruction (SSR), Hybrid Mask, and Scale Ladder (SL) in its structural pre-training path. SSR replaces intensity recovery with structure prediction, Hybrid Mask increases supervision on fine structural regions, and SL learns neighboring-scale structural transitions through bidirectional prediction. Subset-to-Full Regularization (S2F) further stabilizes case-level representations during pre-training. Experimental results on the Brain Tumor Segmentation 2019 (BraTS19) and carotid artery datasets demonstrate that SLC consistently outperforms matched scratch fine-tuning and achieves competitive performance against recent segmentation and self-supervised methods. These results indicate that structure-aware and cross-scale self-supervised objectives can provide effective representations for multimodal 3D medical segmentation.

Keywords

Representation learning; self-supervised learning; multimodal imaging; 3D medical image segmentation; scale-space structural reconstruction; medical image analysis

Cite This Article

APA Style
Liu, W., Li, B., Tian, L., Qiu, Q. (2026). Scale Ladder Consistency for Structure-Aware Multimodal Representation Learning in 3D Medical Image Segmentation. Computer Modeling in Engineering & Sciences, 148(2), 39. https://doi.org/10.32604/cmes.2026.087647
Vancouver Style
Liu W, Li B, Tian L, Qiu Q. Scale Ladder Consistency for Structure-Aware Multimodal Representation Learning in 3D Medical Image Segmentation. Comput Model Eng Sci. 2026;148(2):39. https://doi.org/10.32604/cmes.2026.087647
IEEE Style
W. Liu, B. Li, L. Tian, and Q. Qiu, “Scale Ladder Consistency for Structure-Aware Multimodal Representation Learning in 3D Medical Image Segmentation,” Comput. Model. Eng. Sci., vol. 148, no. 2, pp. 39, 2026. https://doi.org/10.32604/cmes.2026.087647



cc Copyright © 2026 The Author(s). Published by Tech Science Press.
This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
  • 15

    View

  • 5

    Download

  • 0

    Like

Share Link