Improving Statistical Significance in Human Evaluation of Automatic Metrics via Soft Pairwise Accuracy

Brian Thompson; Nitika Mathur; Daniel Deutsch; Huda Khayrallah

doi:10.18653/v1/2024.wmt-1.118

Improving Statistical Significance in Human Evaluation of Automatic Metrics via Soft Pairwise Accuracy

Brian Thompson, Nitika Mathur, Daniel Deutsch, Huda Khayrallah

Abstract

Selecting an automatic metric that best emulates human annotators is often non-trivial, because there is no clear definition of “best emulates.” A meta-metric is required to compare the human judgments to the automatic metric scores, and metric rankings depend on the choice of meta-metric. We propose Soft Pairwise Accuracy (SPA), a new meta-metric that builds on Pairwise Accuracy (PA) but incorporates the statistical significance of both the human judgments and the metric scores. We show that SPA is more stable than PA with respect to changes in the number of systems/segments used for evaluation. We also show that PA can only assign a small set of distinct output values to metrics, and this results in many metrics being artificially assigned the exact same PA score. We demonstrate that SPA fixes this issue. Finally, we show that SPA is more discriminative than PA, producing more statistically significant comparisons between metrics. SPA was selected as the official system-level metric for the 2024 WMT Metrics Shared Task.

Anthology ID:: 2024.wmt-1.118
Volume:: Proceedings of the Ninth Conference on Machine Translation
Month:: November
Year:: 2024
Address:: Miami, Florida, USA
Editors:: Barry Haddow, Tom Kocmi, Philipp Koehn, Christof Monz
Venues:: WMT | WS
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 1222–1234
Language:
URL:: https://aclanthology.org/2024.wmt-1.118/
DOI:: 10.18653/v1/2024.wmt-1.118
Bibkey:
Cite (ACL):: Brian Thompson, Nitika Mathur, Daniel Deutsch, and Huda Khayrallah. 2024. Improving Statistical Significance in Human Evaluation of Automatic Metrics via Soft Pairwise Accuracy. In Proceedings of the Ninth Conference on Machine Translation, pages 1222–1234, Miami, Florida, USA. Association for Computational Linguistics.
Cite (Informal):: Improving Statistical Significance in Human Evaluation of Automatic Metrics via Soft Pairwise Accuracy (Thompson et al., WMT 2024)
Copy Citation:
PDF:: https://aclanthology.org/2024.wmt-1.118.pdf

PDF Cite Search Fix data