Correcting FLORES Evaluation Dataset for Four African Languages

Idris Abdulmumin; Sthembiso Mkhwanazi; Mahlatse Mbooi; Shamsuddeen Hassan Muhammad; Ibrahim Said Ahmad; Neo Putini; Miehleketo Mathebula; Matimba Shingange; Tajuddeen Gwadabe; Vukosi Marivate

doi:10.18653/v1/2024.wmt-1.44

Correcting FLORES Evaluation Dataset for Four African Languages

Idris Abdulmumin, Sthembiso Mkhwanazi, Mahlatse Mbooi, Shamsuddeen Hassan Muhammad, Ibrahim Said Ahmad, Neo Putini, Miehleketo Mathebula, Matimba Shingange, Tajuddeen Gwadabe, Vukosi Marivate

Abstract

This paper describes the corrections made to the FLORES evaluation (dev and devtest) dataset for four African languages, namely Hausa, Northern Sotho (Sepedi), Xitsonga, and isiZulu. The original dataset, though groundbreaking in its coverage of low-resource languages, exhibited various inconsistencies and inaccuracies in the reviewed languages that could potentially hinder the integrity of the evaluation of downstream tasks in natural language processing (NLP), especially machine translation. Through a meticulous review process by native speakers, several corrections were identified and implemented, improving the dataset’s overall quality and reliability. For each language, we provide a concise summary of the errors encountered and corrected and also present some statistical analysis that measures the difference between the existing and corrected datasets. We believe that our corrections enhance the linguistic accuracy and reliability of the data and, thereby, contribute to a more effective evaluation of NLP tasks involving the four African languages. Finally, we recommend that future translation efforts, particularly in low-resource languages, prioritize the active involvement of native speakers at every stage of the process to ensure linguistic accuracy and cultural relevance.

Anthology ID:: 2024.wmt-1.44
Volume:: Proceedings of the Ninth Conference on Machine Translation
Month:: November
Year:: 2024
Address:: Miami, Florida, USA
Editors:: Barry Haddow, Tom Kocmi, Philipp Koehn, Christof Monz
Venues:: WMT | WS
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 570–578
Language:
URL:: https://aclanthology.org/2024.wmt-1.44/
DOI:: 10.18653/v1/2024.wmt-1.44
Bibkey:
Cite (ACL):: Idris Abdulmumin, Sthembiso Mkhwanazi, Mahlatse Mbooi, Shamsuddeen Hassan Muhammad, Ibrahim Said Ahmad, Neo Putini, Miehleketo Mathebula, Matimba Shingange, Tajuddeen Gwadabe, and Vukosi Marivate. 2024. Correcting FLORES Evaluation Dataset for Four African Languages. In Proceedings of the Ninth Conference on Machine Translation, pages 570–578, Miami, Florida, USA. Association for Computational Linguistics.
Cite (Informal):: Correcting FLORES Evaluation Dataset for Four African Languages (Abdulmumin et al., WMT 2024)
Copy Citation:
PDF:: https://aclanthology.org/2024.wmt-1.44.pdf

PDF Cite Search Fix data