Open Access
REVIEW
A Review of Vision Language Models for Architectures, Training Methods, Datasets, Evaluation Metrics, Results, and Fine-Tuning Techniques for Vietnamese
1 Faculty of Engineering Technology, Hung Vuong University, Nong Trang, Phu Tho, Vietnam
2 Thai Nguyen University of Information and Communication Technology, Thai Nguyen, Vietnam
3 Information Technology Department, Tan Trao University, Minh Xuan, Tuyen Quang, Vietnam
* Corresponding Author: Van-Hung Le. Email:
Computers, Materials & Continua 2026, 89(1), 8 https://doi.org/10.32604/cmc.2026.081249
Received 27 February 2026; Accepted 14 May 2026; Issue published 13 August 2026
Abstract
The vision-language models (VLM) combine the image and text to solve practical applications. Specifically, VLM leverages the results of computer vision in conjunction with natural language processing (NLP), like a large language model (LLM), to address real-world problems such as automating and improving the quality of medical examinations and treatments in healthcare, building autonomous driving systems, image captioning, and generating automated chatbots. To understand the development and application of VLM, we surveyed VLM, classifying it according to model architecture, learning methods, evaluation measures, datasets, challenges, and future development directions of VLM based on the model architecture. Simultaneously, to experiment with the VLM model, focusing on LLM, we collected the VQA-TQU1 dataset with multimodal information streams: images, text descriptions, and audio data of Tan Trao University from 2024 to 2026. The VQA-TQU1 dataset was fine-tuned on the Vintern-1B model for generating visual question answering (VQA) (the results on the BLEU, ROUGE-L, METEOR, F1, Char-F1 measures were 0.5677, 0.739, 0.7532, 0.7435, 0.7461, respectively and compare it with state-of-the-art methods), generating automated responses about Tan Trao University’s admissions, and adjusting the data based on the LLM fine-tuned in the Vintern-1B model. We also tested the online video captioning problem on the VQA-TQU1 dataset with two videos based on the SmolVLM2 and Wav2vec 2.0 models, with results showing RTF = 0.24 and RTF = 0.36, respectively.Keywords
Cite This Article
Copyright © 2026 The Author(s). Published by Tech Science Press.This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.


Submit a Paper
Propose a Special lssue
View Full Text
Download PDF
Downloads
Citation Tools