Can Automatic Metrics Assess High-Quality Translations?

Sweta Agrawal, António Farinhas, Ricardo Rei, Andre Martins


Abstract
Automatic metrics for evaluating translation quality are typically validated by measuring how well they correlate with human assessments. However, correlation methods tend to capture only the ability of metrics to differentiate between good and bad source-translation pairs, overlooking their reliability in distinguishing alternative translations for the same source. In this paper, we confirm that this is indeed the case by showing that current metrics are insensitive to nuanced differences in translation quality. This effect is most pronounced when the quality is high and the variance among alternatives is low. Given this finding, we shift towards detecting high-quality correct translations, an important problem in practical decision-making scenarios where a binary check of correctness is prioritized over a nuanced evaluation of quality. Using the MQM framework as the gold standard, we systematically stress-test the ability of current metrics to identify translations with no errors as marked by humans. Our findings reveal that current metrics often over or underestimate translation quality, indicating significant room for improvement in machine translation evaluation.
Anthology ID:
2024.emnlp-main.802
Volume:
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
Month:
November
Year:
2024
Address:
Miami, Florida, USA
Editors:
Yaser Al-Onaizan, Mohit Bansal, Yun-Nung Chen
Venue:
EMNLP
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
14491–14502
Language:
URL:
https://aclanthology.org/2024.emnlp-main.802
DOI:
Bibkey:
Cite (ACL):
Sweta Agrawal, António Farinhas, Ricardo Rei, and Andre Martins. 2024. Can Automatic Metrics Assess High-Quality Translations?. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 14491–14502, Miami, Florida, USA. Association for Computational Linguistics.
Cite (Informal):
Can Automatic Metrics Assess High-Quality Translations? (Agrawal et al., EMNLP 2024)
Copy Citation:
PDF:
https://aclanthology.org/2024.emnlp-main.802.pdf