Home / Journals / CMES / Online First / doi:10.32604/cmes.2026.083586
Special Issues
Table of Content

Open Access

REVIEW

A Comprehensive Review of Complex Logical Reasoning in Large Vision-Language Models

Weiqiang Jin1,2,#, Yang Liu2,#, Yang Gao1,#, Shixiang Tang2, Yanghao Zhou3, Jinhu Qi4, Wentao Zhang4, Junli Wang5, Jing Gao2, Yue Ma4, Ziwei Zhang1,*, Biao Zhao2,*
1 Institute of SRIICL, Xi’an Jiaotong University, Xi’an, China
2 School of Information and Communications Engineering, Xi’an Jiaotong University, Xi’an, China
3 Department of Electrical and Computer Engineering, National University of Singapore, Kent Ridge, Singapore
4 Department of Computer Science and Engineering, The Chinese University of Hong Kong, Hong Kong SAR, China
5 School of Computer Science and Technology, University of Science and Technology of China, Hefei, China
* Corresponding Author: Ziwei Zhang. Email: email; Biao Zhao. Email: email
# These authors contributed equally to this work
(This article belongs to the Special Issue: Emerging Artificial Intelligence Technologies and Applications-II)

Computer Modeling in Engineering & Sciences https://doi.org/10.32604/cmes.2026.083586

Received 10 April 2026; Accepted 15 June 2026; Published online 06 July 2026

Abstract

Large Vision-Language Models (LVLMs) have achieved strong performance in multimodal perception, understanding, and generation, but their ability to perform complex logical reasoning remains insufficiently understood. In particular, it is still unclear whether current LVLMs can reliably conduct explicit logical operations, multi-step inference, abstract relational reasoning, and cross-modal evidence integration. Reasoning abilities such as deductive, inductive, abductive, multi-hop, and causal inference are fundamental to robust decision making, trustworthy interaction, and real-world deployment, yet they have not been systematically examined in the LVLM literature. Existing surveys mainly discuss mathematical reasoning, general multimodal intelligence, or benchmark progress, but they do not provide a unified account of complex logical reasoning in LVLMs, including its definition, reasoning types, modeling paradigms, evaluation protocols, and unresolved limitations. To address this gap, this survey develops a unified analytical framework for complex logical reasoning in LVLMs. This survey provides a structured review of this emerging area. We first formalize complex logical reasoning in multimodal settings and organize the literature into five recurrent reasoning families: deductive, inductive, abductive, multi-hop, and causal reasoning. We then review reasoning-oriented LVLM architectures, including unified, modular, and tool-augmented paradigms, and summarize major reasoning mechanisms such as chain-of-thought, program-based reasoning, self-correction, and interpretability-oriented analysis. We further examine representative benchmarks and evaluation protocols, with particular attention to the mismatch between final-answer accuracy and genuine reasoning validity. Based on empirical evidence from representative LVLMs and datasets, we identify common capability trends, recurring failure modes, and key open challenges. Our analysis shows that current LVLMs still struggle with reasoning faithfulness, long-horizon inference, cross-modal grounding, hallucination control, and process-aware evaluation. Finally, we outline future directions in reasoning-oriented data construction, model design, training strategies, evaluation methodology, and deployment. Overall, this survey offers a unified conceptual framework and technical roadmap for advancing LVLMs from strong perceptual systems toward reliable multimodal reasoning agents.

Keywords

Large vision-language model; complex logical reasoning; multimodal reasoning; chain-of-thought; evaluation benchmark; reasoning faithfulness
  • 249

    View

  • 43

    Download

  • 0

    Like

Share Link