Beyond the Last Frame: Process-aware Evaluation for Generative Video Reasoning

Yifan Li; YuKai Gu; Yingqian Min; Zikang Liu; Yifan Du; Kun Zhou; Min Yang; Wayne Xin Zhao; Minghui Qiu

Beyond the Last Frame: Process-aware Evaluation for Generative Video Reasoning

Yifan Li, YuKai Gu, Yingqian Min, Zikang Liu, Yifan Du, Kun Zhou, Min Yang, Xin Zhao, Minghui Qiu

Abstract

Recent breakthroughs in video generation have demonstrated an emerging capability termed Chain-of-Frames (CoF) reasoning, where models resolve complex tasks through the generation of continuous frames. While these models show promise for Generative Video Reasoning (GVR), existing evaluation frameworks often rely on single-frame assessments, which can lead to outcome-hacking, where a model reaches a correct conclusion through an erroneous process. To address this, we propose a process-aware evaluation paradigm. We introduce VIPER, a comprehensive benchmark spanning 16 tasks across temporal, structural, symbolic, spatial, physics, and planning reasoning. Furthermore, we propose Process-outcome Consistency (POC@r), a new metric that utilizes VLM-as-Judge with a hierarchical rubric to evaluate both the validity of the intermediate steps and the final result. Our experiments reveal that state-of-the-art video models achieve POC@1.0 only about 20% and exhibit a significant outcome-hacking. We further explore the impact of test-time scaling and sampling robustness, highlighting a substantial gap between current video generation and true generalized visual reasoning. Our benchmark are released at https://github.com/RUCAIBox/VIPER.

Anthology ID:: 2026.acl-long.934
Volume:: Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Month:: July
Year:: 2026
Address:: San Diego, California, United States
Editors:: Maria Liakata, Viviane P. Moreira, Jiajun Zhang, David Jurgens
Venue:: ACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 20393–20409
Language:
URL:: https://aclanthology.org/2026.acl-long.934/
DOI:
Bibkey:
Cite (ACL):: Yifan Li, YuKai Gu, Yingqian Min, Zikang Liu, Yifan Du, Kun Zhou, Min Yang, Xin Zhao, and Minghui Qiu. 2026. Beyond the Last Frame: Process-aware Evaluation for Generative Video Reasoning. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 20393–20409, San Diego, California, United States. Association for Computational Linguistics.
Cite (Informal):: Beyond the Last Frame: Process-aware Evaluation for Generative Video Reasoning (Li et al., ACL 2026)
Copy Citation:
PDF:: https://aclanthology.org/2026.acl-long.934.pdf
Checklist:: 2026.acl-long.934.checklist.pdf

PDF Cite Search Checklist Fix data