Learned Image Compression via Text-Semantic Guidance and Content-Aware Bitrate Control
Kaisen Li1, Yunwei Zhang1,*, Guoying Sun1, Bin Li2,*
1 Faculty of Information Engineering and Automation, Kunming University of Science and Technology, Kunming, China
2 Information Center, Yunnan Tobacco and Leaf Company, Kunming, China
* Corresponding Author: Yunwei Zhang. Email:
; Bin Li. Email:
Computers, Materials & Continua https://doi.org/10.32604/cmc.2026.084755
Received 28 April 2026; Accepted 24 June 2026; Published online 13 July 2026
Abstract
With the development of vision-language pre-trained models, effectively exploiting high-level semantics and precisely controlling bitrate in learned image compression remains a challenging problem. Existing methods mainly rely on image feature modeling alone, making it difficult to jointly preserve fine-grained details and semantic consistency under a given bitrate budget. To address this issue, this paper proposes a learned image compression framework that integrates text-semantic guidance with content-aware bitrate control. The framework combines Bootstrapping Language-Image Pre-training (BLIP) and Contrastive Language-Image Pre-training (CLIP) to extract image semantic information, and performs conditional modulation on multi-scale visual features through feature-wise linear modulation to construct semantically enhanced latent representations, thereby improving structural preservation and visual reconstruction quality under low-bitrate conditions. Meanwhile, we design a bitrate control module that integrates content awareness with historical feedback. Regional importance is estimated by channel attention, and the quantization strength is adaptively adjusted according to the target and actual bitrates, enabling the model to approach the preset bitrate without manual tuning. Experimental results on different datasets show that, compared with representative learned image compression methods, including Cheng2020, Efficient Learned Image Compression (ELIC), and Frequency-aware Transformer for Learned Image Compression (FTIC), the proposed method achieves a more competitive rate-distortion trade-off. In particular, it consistently exhibits superior rate-distortion (RD) performance, delivers higher objective reconstruction quality at the same bitrate, and maintains favorable subjective perceptual consistency.
Keywords
Learned image compression; feature-wise linear modulation; semantic guidance; semantically enhanced latent representation; content-aware bitrate control