Home / Journals / CMC / Online First / doi:10.32604/cmc.2026.081249
Special Issues
Table of Content

Open Access

REVIEW

A Review of Vision Language Models for Architectures, Training Methods, Datasets, Evaluation Metrics, Results, and Fine-Tuning Techniques for Vietnamese

Van-Thuan Nguyen1,2, Van-Nui Nguyen2, Van-Hung Le3,*
1 Faculty of Engineering Technology, Hung Vuong University, Nong Trang, Phu Tho, Vietnam
2 Thai Nguyen University of Information and Communication Technology, Thai Nguyen, Vietnam
3 Information Technology Department, Tan Trao University, Minh Xuan, Tuyen Quang, Vietnam
* Corresponding Author: Van-Hung Le. Email: email

Computers, Materials & Continua https://doi.org/10.32604/cmc.2026.081249

Received 27 February 2026; Accepted 14 May 2026; Published online 03 August 2026

Abstract

The vision-language models (VLM) combine the image and text to solve practical applications. Specifically, VLM leverages the results of computer vision in conjunction with natural language processing (NLP), like a large language model (LLM), to address real-world problems such as automating and improving the quality of medical examinations and treatments in healthcare, building autonomous driving systems, image captioning, and generating automated chatbots. To understand the development and application of VLM, we surveyed VLM, classifying it according to model architecture, learning methods, evaluation measures, datasets, challenges, and future development directions of VLM based on the model architecture. Simultaneously, to experiment with the VLM model, focusing on LLM, we collected the VQA-TQU1 dataset with multimodal information streams: images, text descriptions, and audio data of Tan Trao University from 2024 to 2026. The VQA-TQU1 dataset was fine-tuned on the Vintern-1B model for generating visual question answering (VQA) (the results on the BLEU, ROUGE-L, METEOR, F1, Char-F1 measures were 0.5677, 0.739, 0.7532, 0.7435, 0.7461, respectively and compare it with state-of-the-art methods), generating automated responses about Tan Trao University’s admissions, and adjusting the data based on the LLM fine-tuned in the Vintern-1B model. We also tested the online video captioning problem on the VQA-TQU1 dataset with two videos based on the SmolVLM2 and Wav2vec 2.0 models, with results showing RTF = 0.24 and RTF = 0.36, respectively.

Keywords

Vision-language models (VLMs); VLMs classification; embedding-based VLM; generative VLM; cross-modal transformer; VLM + LLM; embodied VLM; fine-tuning VLM for Vietnamese
  • 98

    View

  • 18

    Download

  • 0

    Like

Share Link