Unveiling NLG Human-Evaluation Reproducibility: Lessons Learned and Key Insights from Participating in the ReproNLP Challenge

Lewis Watson; Dimitra Gkatzia

Unveiling NLG Human-Evaluation Reproducibility: Lessons Learned and Key Insights from Participating in the ReproNLP Challenge

Abstract

Human evaluation is crucial for NLG systems as it provides a reliable assessment of the quality, effectiveness, and utility of generated language outputs. However, concerns about the reproducibility of such evaluations have emerged, casting doubt on the reliability and generalisability of reported results. In this paper, we present the findings of a reproducibility study on a data-to-text system, conducted under two conditions: (1) replicating the original setup as closely as possible with evaluators from AMT, and (2) replicating the original human evaluation but this time, utilising evaluators with a background in academia. Our experiments show that there is a loss of statistical significance between the original and reproduction studies, i.e. the human evaluation results are not reproducible. In addition, we found that employing local participants led to more robust results. We finally discuss lessons learned, addressing the challenges and best practices for ensuring reproducibility in NLG human evaluations.

Anthology ID:: 2023.humeval-1.6
Volume:: Proceedings of the 3rd Workshop on Human Evaluation of NLP Systems
Month:: September
Year:: 2023
Address:: Varna, Bulgaria
Editors:: Anya Belz, Maja Popović, Ehud Reiter, Craig Thomson, João Sedoc
Venues:: HumEval | WS
SIG:
Publisher:: INCOMA Ltd., Shoumen, Bulgaria
Note:
Pages:: 69–74
Language:
URL:: https://aclanthology.org/2023.humeval-1.6
DOI:
Bibkey:
Cite (ACL):: Lewis Watson and Dimitra Gkatzia. 2023. Unveiling NLG Human-Evaluation Reproducibility: Lessons Learned and Key Insights from Participating in the ReproNLP Challenge. In Proceedings of the 3rd Workshop on Human Evaluation of NLP Systems, pages 69–74, Varna, Bulgaria. INCOMA Ltd., Shoumen, Bulgaria.
Cite (Informal):: Unveiling NLG Human-Evaluation Reproducibility: Lessons Learned and Key Insights from Participating in the ReproNLP Challenge (Watson & Gkatzia, HumEval-WS 2023)
Copy Citation:
PDF:: https://aclanthology.org/2023.humeval-1.6.pdf

PDF Cite Search