TY - EJOU AU - Nguyen, Van-Thuan AU - Nguyen, Van-Nui AU - Le, Van-Hung TI - A Review of Vision Language Models for Architectures, Training Methods, Datasets, Evaluation Metrics, Results, and Fine-Tuning Techniques for Vietnamese T2 - Computers, Materials \& Continua PY - VL - IS - SN - 1546-2226 AB - The vision-language models (VLM) combine the image and text to solve practical applications. Specifically, VLM leverages the results of computer vision in conjunction with natural language processing (NLP), like a large language model (LLM), to address real-world problems such as automating and improving the quality of medical examinations and treatments in healthcare, building autonomous driving systems, image captioning, and generating automated chatbots. To understand the development and application of VLM, we surveyed VLM, classifying it according to model architecture, learning methods, evaluation measures, datasets, challenges, and future development directions of VLM based on the model architecture. Simultaneously, to experiment with the VLM model, focusing on LLM, we collected the VQA-TQU1 dataset with multimodal information streams: images, text descriptions, and audio data of Tan Trao University from 2024 to 2026. The VQA-TQU1 dataset was fine-tuned on the Vintern-1B model for generating visual question answering (VQA) (the results on the BLEU, ROUGE-L, METEOR, F1, Char-F1 measures were 0.5677, 0.739, 0.7532, 0.7435, 0.7461, respectively and compare it with state-of-the-art methods), generating automated responses about Tan Trao University’s admissions, and adjusting the data based on the LLM fine-tuned in the Vintern-1B model. We also tested the online video captioning problem on the VQA-TQU1 dataset with two videos based on the SmolVLM2 and Wav2vec 2.0 models, with results showing RTF = 0.24 and RTF = 0.36, respectively. KW - Vision-language models (VLMs); VLMs classification; embedding-based VLM; generative VLM; cross-modal transformer; VLM + LLM; embodied VLM; fine-tuning VLM for Vietnamese DO - 10.32604/cmc.2026.081249