FIBER: Fill-in-the-Blanks as a Challenging Video Understanding Evaluation Framework

Santiago Castro, Ruoyao Wang, Pingxuan Huang, Ian Stewart, Oana Ignat, Nan Liu, Jonathan Stroud, Rada Mihalcea


Abstract
We propose fill-in-the-blanks as a video understanding evaluation framework and introduce FIBER – a novel dataset consisting of 28,000 videos and descriptions in support of this evaluation framework. The fill-in-the-blanks setting tests a model’s understanding of a video by requiring it to predict a masked noun phrase in the caption of the video, given the video and the surrounding text. The FIBER benchmark does not share the weaknesses of the current state-of-the-art language-informed video understanding tasks, namely: (1) video question answering using multiple-choice questions, where models perform relatively well because they exploit linguistic biases in the task formulation, thus making our framework challenging for the current state-of-the-art systems to solve; and (2) video captioning, which relies on an open-ended evaluation framework that is often inaccurate because system answers may be perceived as incorrect if they differ in form from the ground truth. The FIBER dataset and our code are available at https://lit.eecs.umich.edu/fiber/.
Anthology ID:
2022.acl-long.209
Volume:
Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Month:
May
Year:
2022
Address:
Dublin, Ireland
Editors:
Smaranda Muresan, Preslav Nakov, Aline Villavicencio
Venue:
ACL
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
2925–2940
Language:
URL:
https://aclanthology.org/2022.acl-long.209
DOI:
10.18653/v1/2022.acl-long.209
Bibkey:
Cite (ACL):
Santiago Castro, Ruoyao Wang, Pingxuan Huang, Ian Stewart, Oana Ignat, Nan Liu, Jonathan Stroud, and Rada Mihalcea. 2022. FIBER: Fill-in-the-Blanks as a Challenging Video Understanding Evaluation Framework. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2925–2940, Dublin, Ireland. Association for Computational Linguistics.
Cite (Informal):
FIBER: Fill-in-the-Blanks as a Challenging Video Understanding Evaluation Framework (Castro et al., ACL 2022)
Copy Citation:
PDF:
https://aclanthology.org/2022.acl-long.209.pdf
Video:
 https://aclanthology.org/2022.acl-long.209.mp4
Code
 MichiganNLP/video-fill-in-the-blank
Data
ActivityNet CaptionsVATEXVisual Question Answering