When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs

Hongcheng Liu; Pingjie Wang; Yuhao Wang; Siqu Ou; Yanfeng Wang; Yu Wang

When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs

Hongcheng Liu, Pingjie Wang, Yuhao Wang, Siqu Ou, Yanfeng Wang, Yu Wang

Abstract

Multimodal large language models (MLLMs) have shown strong capabilities across a broad range of benchmarks. However, most existing evaluations focus on passive inference, where models perform step-by-step reasoning under complete information. This setup is misaligned with real-world use, where seeing is not enough. This raises a fundamental question: Can MLLMs actively acquire missing evidence under incomplete information? To bridge this gap, we require the MLLMs to actively acquire missing evidence and iteratively refine decisions under incomplete information, by selecting a target image from a candidate pool without task-specific priors. To support systematic study, we propose GuessBench, a benchmark with both perception-oriented and knowledge-oriented images for evaluating active reasoning in MLLMs. We evaluate 20 superior MLLMs and find that performance on active reasoning lags far behind it on passive settings, indicating substantial room for improvement. Further analysis identifies fine-grained perception and timely decision-making as key challenges. Ablation studies show that perceptual enhancements benefit smaller models, whereas thinking-oriented methods provide consistent gains across model sizes. These results suggest promising directions for future research on multimodal active reasoning.

Anthology ID:: 2026.acl-long.1264
Volume:: Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Month:: July
Year:: 2026
Address:: San Diego, California, United States
Editors:: Maria Liakata, Viviane P. Moreira, Jiajun Zhang, David Jurgens
Venue:: ACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 27395–27422
Language:
URL:: https://aclanthology.org/2026.acl-long.1264/
DOI:
Bibkey:
Cite (ACL):: Hongcheng Liu, Pingjie Wang, Yuhao Wang, Siqu Ou, Yanfeng Wang, and Yu Wang. 2026. When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 27395–27422, San Diego, California, United States. Association for Computational Linguistics.
Cite (Informal):: When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs (Liu et al., ACL 2026)
Copy Citation:
PDF:: https://aclanthology.org/2026.acl-long.1264.pdf
Checklist:: 2026.acl-long.1264.checklist.pdf

PDF Cite Search Checklist Fix data