GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?

Yiyang Zhou; Linjie Li; Shi Qiu; Zhengyuan Yang; Yuyang Zhao; Siwei Han; Yangfan He; Kangqi Li; Haonian Ji; Zihao Zhao; Haibo Tong; Lijuan Wang; Huaxiu Yao

doi:10.18653/v1/2025.emnlp-main.1415

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?

Yiyang Zhou, Linjie Li, Shi Qiu, Zhengyuan Yang, Yuyang Zhao, Siwei Han, Yangfan He, Kangqi Li, Haonian Ji, Zihao Zhao, Haibo Tong, Lijuan Wang, Huaxiu Yao

Abstract

Existing video benchmarks often resemble image-based benchmarks, with question types like “What actions does the person perform throughout the video?” or “What color is the woman’s dress in the video?” For these, models can often answer by scanning just a few key frames, without deep temporal reasoning. This limits our ability to assess whether large vision-language models (LVLMs) can truly think with videos rather than perform superficial frame-level analysis. To address this, we introduce , a benchmark specifically designed to evaluate whether LVLMs can genuinely think with videos. Unlike prior benchmarks, emphasizes comprehensive video understanding beyond static image cues. It consists of 3,269 videos and over 4,342 highly visual-centric questions across 11 categories, including Trajectory Analysis, Temporal Reasoning, and Forensics Detection. All questions are carefully crafted by human annotators and require watching the entire video and reasoning over full video context—this is what we mean by thinking with video. These questions cannot be answered by scanning selected frames or relying on text alone. In human evaluations, achieves 94.82% accuracy, but current LVLMs face significant challenges. Even the best-performing model, GPT-o3, reaches only 66.43%, highlighting that LVLMs still struggle to move beyond surface-level reasoning to truly think with videos. We publicly release our benchmark and code at https://github.com/aiming-lab/GLIMPSE.

Anthology ID:: 2025.emnlp-main.1415
Volume:: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Month:: November
Year:: 2025
Address:: Suzhou, China
Editors:: Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng
Venue:: EMNLP
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 27842–27856
Language:
URL:: https://aclanthology.org/2025.emnlp-main.1415/
DOI:: 10.18653/v1/2025.emnlp-main.1415
Bibkey:
Cite (ACL):: Yiyang Zhou, Linjie Li, Shi Qiu, Zhengyuan Yang, Yuyang Zhao, Siwei Han, Yangfan He, Kangqi Li, Haonian Ji, Zihao Zhao, Haibo Tong, Lijuan Wang, and Huaxiu Yao. 2025. GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 27842–27856, Suzhou, China. Association for Computational Linguistics.
Cite (Informal):: GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? (Zhou et al., EMNLP 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.emnlp-main.1415.pdf
Checklist:: 2025.emnlp-main.1415.checklist.pdf

PDF Cite Search Checklist Fix data