Using Model Disagreement to Identify Unstable Regions in MT Evaluation

Vitalii Iakivchuk


Abstract
Human evaluation of MT is essential but exhibits substantial annotator variability that limits evaluation reliability and super- vised learning. Rather than treating dis- agreement as noise or correcting it through protocol changes, we analyze its structure via learned severity classifiers. Across training regimes defined by base- line model reproducibility, we observe in- ternally coherent but mutually incompati- ble severity mappings: models trained on one regime produce confident predictions within that regime but reduced separability on the other. Margin–correctness analysis shows that instability is not uniformly low confidence; separability depends on align- ment between model-internalized and hu- man annotation regimes. These results indicate that unstable MT evaluation regions arise primarily from competing severity interpretations rather than intrinsic example difficulty. Model– annotator disagreement therefore provides a practical signal for identifying unstable evaluation regions during MT evaluation.
Anthology ID:
2026.eamt-1.22
Volume:
Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 1)
Month:
June
Year:
2026
Address:
Tilburg, The Netherlands
Editors:
Dimitar Shterionov, Eva Vanmassenhove, Mirella De Sisto, Fred Blain, Javad Pourmostafa Roshan Sharami, Lisa Lepp, Chiara Manna, Argentina Anna Rescigno, Alina Karakanta, Ayla Rigouts Terryn, Manuel Lardelli, Natalia Resende, Elena Murgolo, Janiça Hackenbuchner, Anna Zaretskaya, Miquel Esplà-Gomis, Thierry Etchegoyhen, Dagmar Gromann, Rachel Bawden, Barry Haddow, Sara Szoc, Mikel Forcada, Helena Moniz
Venue:
EAMT
SIG:
Publisher:
European Association for Machine Translation
Note:
Pages:
339–347
Language:
URL:
https://aclanthology.org/2026.eamt-1.22/
DOI:
Bibkey:
Cite (ACL):
Vitalii Iakivchuk. 2026. Using Model Disagreement to Identify Unstable Regions in MT Evaluation. In Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 1), pages 339–347, Tilburg, The Netherlands. European Association for Machine Translation.
Cite (Informal):
Using Model Disagreement to Identify Unstable Regions in MT Evaluation (Iakivchuk, EAMT 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.eamt-1.22.pdf