A Comprehensive Review of Complex Logical Reasoning in Large Vision-Language Models
Weiqiang Jin1,2,#, Yang Liu2,#, Yang Gao1,#, Shixiang Tang2, Yanghao Zhou3, Jinhu Qi4, Wentao Zhang4, Junli Wang5, Jing Gao2, Yue Ma4, Ziwei Zhang1,*, Biao Zhao2,*
CMES-Computer Modeling in Engineering & Sciences, Vol.148, No.1, 2026, DOI:10.32604/cmes.2026.083586
- 27 July 2026
(This article belongs to the Special Issue: Emerging Artificial Intelligence Technologies and Applications-II)
Abstract Large Vision-Language Models (LVLMs) have achieved strong performance in multimodal perception, understanding, and generation, but their ability to perform complex logical reasoning remains insufficiently understood. In particular, it is still unclear whether current LVLMs can reliably conduct explicit logical operations, multi-step inference, abstract relational reasoning, and cross-modal evidence integration. Reasoning abilities such as deductive, inductive, abductive, multi-hop, and causal inference are fundamental to robust decision making, trustworthy interaction, and real-world deployment, yet they have not been systematically examined in the LVLM literature. Existing surveys mainly discuss mathematical reasoning, general multimodal intelligence, or benchmark progress, but… More >