Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models

Masanari Oi, Masahiro Kaneko, Naoaki Okazaki, Nakamasa Inoue


Abstract
Vision-language models (VLMs) have shown impressive abilities across a range of multi-modal tasks. However, existing metrics for evaluating the quality of text generated by VLMs typically focus on an overall evaluation for a specific task, such as image captioning. While the overall evaluation is essential for any task, the criteria prioritized can differ depending on the task, making it challenging for current metrics to adapt to multi-task scenarios. To address this limitation, we propose HarmonicEval, a reference-free comprehensive evaluation metric that aggregates criterion-wise scores to produce the overall score in a bottom-up manner. Furthermore, to assess the generalizability of automatic evaluation metrics in multi-task scenarios, we construct the Multi-task Multi-criteria Human Evaluation (MMHE) benchmark, which comprises 18,000 expert human judgments across four multi-modal tasks. Our experiments demonstrate that HarmonicEval achieves higher correlations with human judgments than conventional metrics while providing numerical scores for each criterion. Our code and data will be available publicly.
Anthology ID:
2026.lrec-1.738
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
9396–9408
Language:
External URL:
https://lrec.elra.info/lrec2026-main-738
DOI:
10.63317/575aknbvd9hs
Bibkey:
Cite (ACL):
Masanari Oi, Masahiro Kaneko, Naoaki Okazaki, and Nakamasa Inoue. 2026. Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 9396–9408, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models (Oi et al., LREC 2026)
Copy Citation: