VDAct 2.0: Scaling Video-Grounded Dialogue for Event-driven Activity Understanding with LLM-Assisted Filtering

Wiradee Imrattanatrai, Masaki Asada, Kimihiro Hasegawa, Ken Fukuda, Teruko Mitamura


Abstract
We present VDAct 2.0, an enhanced benchmark for video-grounded dialogue that builds upon the original VDAct by expanding dialogue coverage and introducing a scalable LLM-assisted filtering pipeline to ensure high-quality, grounded QA pairs. VDAct 2.0 comprises 6,356 human-annotated dialogues with a total of 63,958 turns, grounded in 2,975 household activity videos, with undesirable dialogue turns systematically identified and removed. To achieve this, we design a trigger-based quality framework and calibrate a panel of high-agreement LLMs through human-in-the-loop calibration, allowing scalable QA-turn-level filtering. We benchmark a wide range of pretrained and fine-tuned models, both open-source and proprietary, across standard text generation metrics and LLM-based evaluations. The results highlight both recent advances and remaining challenges in video-grounded dialogue modeling, positioning VDAct 2.0 as a high-fidelity testbed for evaluating and advancing multimodal reasoning in interactive settings.
Anthology ID:
2026.lrec-1.208
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
2653–2666
Language:
External URL:
https://lrec.elra.info/lrec2026-main-208
DOI:
10.63317/4vcv4ncvs6xx
Bibkey:
Cite (ACL):
Wiradee Imrattanatrai, Masaki Asada, Kimihiro Hasegawa, Ken Fukuda, and Teruko Mitamura. 2026. VDAct 2.0: Scaling Video-Grounded Dialogue for Event-driven Activity Understanding with LLM-Assisted Filtering. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 2653–2666, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
VDAct 2.0: Scaling Video-Grounded Dialogue for Event-driven Activity Understanding with LLM-Assisted Filtering (Imrattanatrai et al., LREC 2026)
Copy Citation: