Open Access iconOpen Access

REVIEW

A Comprehensive Review of Complex Logical Reasoning in Large Vision-Language Models

Weiqiang Jin1,2,#, Yang Liu2,#, Yang Gao1,#, Shixiang Tang2, Yanghao Zhou3, Jinhu Qi4, Wentao Zhang4, Junli Wang5, Jing Gao2, Yue Ma4, Ziwei Zhang1,*, Biao Zhao2,*

1 Institute of SRIICL, Xi’an Jiaotong University, Xi’an, China
2 School of Information and Communications Engineering, Xi’an Jiaotong University, Xi’an, China
3 Department of Electrical and Computer Engineering, National University of Singapore, Kent Ridge, Singapore
4 Department of Computer Science and Engineering, The Chinese University of Hong Kong, Hong Kong SAR, China
5 School of Computer Science and Technology, University of Science and Technology of China, Hefei, China

* Corresponding Authors: Ziwei Zhang. Email: email; Biao Zhao. Email: email
# These authors contributed equally to this work

(This article belongs to the Special Issue: Emerging Artificial Intelligence Technologies and Applications-II)

Computer Modeling in Engineering & Sciences 2026, 148(1), 1 https://doi.org/10.32604/cmes.2026.083586

Abstract

Large Vision-Language Models (LVLMs) have achieved strong performance in multimodal perception, understanding, and generation, but their ability to perform complex logical reasoning remains insufficiently understood. In particular, it is still unclear whether current LVLMs can reliably conduct explicit logical operations, multi-step inference, abstract relational reasoning, and cross-modal evidence integration. Reasoning abilities such as deductive, inductive, abductive, multi-hop, and causal inference are fundamental to robust decision making, trustworthy interaction, and real-world deployment, yet they have not been systematically examined in the LVLM literature. Existing surveys mainly discuss mathematical reasoning, general multimodal intelligence, or benchmark progress, but they do not provide a unified account of complex logical reasoning in LVLMs, including its definition, reasoning types, modeling paradigms, evaluation protocols, and unresolved limitations. To address this gap, this survey develops a unified analytical framework for complex logical reasoning in LVLMs. This survey provides a structured review of this emerging area. We first formalize complex logical reasoning in multimodal settings and organize the literature into five recurrent reasoning families: deductive, inductive, abductive, multi-hop, and causal reasoning. We then review reasoning-oriented LVLM architectures, including unified, modular, and tool-augmented paradigms, and summarize major reasoning mechanisms such as chain-of-thought, program-based reasoning, self-correction, and interpretability-oriented analysis. We further examine representative benchmarks and evaluation protocols, with particular attention to the mismatch between final-answer accuracy and genuine reasoning validity. Based on empirical evidence from representative LVLMs and datasets, we identify common capability trends, recurring failure modes, and key open challenges. Our analysis shows that current LVLMs still struggle with reasoning faithfulness, long-horizon inference, cross-modal grounding, hallucination control, and process-aware evaluation. Finally, we outline future directions in reasoning-oriented data construction, model design, training strategies, evaluation methodology, and deployment. Overall, this survey offers a unified conceptual framework and technical roadmap for advancing LVLMs from strong perceptual systems toward reliable multimodal reasoning agents.

Keywords

Large vision-language model; complex logical reasoning; multimodal reasoning; chain-of-thought; evaluation benchmark; reasoning faithfulness

Cite This Article

APA Style
Jin, W., Liu, Y., Gao, Y., Tang, S., Zhou, Y. et al. (2026). A Comprehensive Review of Complex Logical Reasoning in Large Vision-Language Models. Computer Modeling in Engineering & Sciences, 148(1), 1. https://doi.org/10.32604/cmes.2026.083586
Vancouver Style
Jin W, Liu Y, Gao Y, Tang S, Zhou Y, Qi J, et al. A Comprehensive Review of Complex Logical Reasoning in Large Vision-Language Models. Comput Model Eng Sci. 2026;148(1):1. https://doi.org/10.32604/cmes.2026.083586
IEEE Style
W. Jin et al., “A Comprehensive Review of Complex Logical Reasoning in Large Vision-Language Models,” Comput. Model. Eng. Sci., vol. 148, no. 1, pp. 1, 2026. https://doi.org/10.32604/cmes.2026.083586



cc Copyright © 2026 The Author(s). Published by Tech Science Press.
This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
  • 333

    View

  • 56

    Download

  • 0

    Like

Share Link