When the Gold Standard Isn’t Necessarily Standard: Challenges of Evaluating the Translation of User-Generated Content

Lydia Nishimwe, Benoît Sagot, Rachel Bawden


Abstract
User-generated content (UGC) is characterised by frequent use of non-standard language, from spelling errors to expressive choices such as slang, character repetitions, and emojis. This makes evaluating UGC translation challenging: what counts as a "good" translation depends on the desired standardness level of the output. To explore this, we examine the human translation guidelines of four UGC datasets, and derive a taxonomy of twelve non-standard phenomena and five translation actions (NORMALISE, COPY, TRANSFER, OMIT, CENSOR). Our analysis reveals notable differences in how UGC is treated, resulting in a spectrum of standardness in reference translations. We show that translation scores of large language models are highly sensitive to prompts with explicit UGC translation instructions, and that they improve when they align with the dataset guidelines. We argue that fair evaluation requires both models and metrics to be aware of translation guidelines. Finally, we call for clear guidelines during dataset creation and for the development of controllable, guideline-aware evaluation frameworks for UGC translation.
Anthology ID:
2026.eamt-1.30
Volume:
Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 1)
Month:
June
Year:
2026
Address:
Tilburg, The Netherlands
Editors:
Dimitar Shterionov, Eva Vanmassenhove, Mirella De Sisto, Fred Blain, Javad Pourmostafa Roshan Sharami, Lisa Lepp, Chiara Manna, Argentina Anna Rescigno, Alina Karakanta, Ayla Rigouts Terryn, Manuel Lardelli, Natalia Resende, Elena Murgolo, Janiça Hackenbuchner, Anna Zaretskaya, Miquel Esplà-Gomis, Thierry Etchegoyhen, Dagmar Gromann, Rachel Bawden, Barry Haddow, Sara Szoc, Mikel Forcada, Helena Moniz
Venue:
EAMT
SIG:
Publisher:
European Association for Machine Translation
Note:
Pages:
473–495
Language:
URL:
https://aclanthology.org/2026.eamt-1.30/
DOI:
Bibkey:
Cite (ACL):
Lydia Nishimwe, Benoît Sagot, and Rachel Bawden. 2026. When the Gold Standard Isn’t Necessarily Standard: Challenges of Evaluating the Translation of User-Generated Content. In Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 1), pages 473–495, Tilburg, The Netherlands. European Association for Machine Translation.
Cite (Informal):
When the Gold Standard Isn’t Necessarily Standard: Challenges of Evaluating the Translation of User-Generated Content (Nishimwe et al., EAMT 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.eamt-1.30.pdf