The Challenge of Finding Robust and Efficient Strategies for Training Machine Translation Models with Noisy Data

Mikko Aulamo, Sami Virpioja, Yves Scherrer, Jörg Tiedemann


Abstract
Most machine translation datasets come with a certain level of noise, and strategies for handling such data need to be robust and efficient. Data selection and filtering are challenging and may depend on expensive language-specific tools that are not necessarily available, especially for low-resource languages. This paper looks at training strategies that combine cheap heuristic filters with curriculum learning to implement iterative procedures that robustly operate on raw noisy data without expensive prior preprocessing and data selection. The intuition is that we can cluster data into buckets with varying noise levels and use different sets of buckets at different stages of MT model training. We test various strategies and compare them to pre-filtering approaches for a diverse set of low-resource languages and conclude that curriculum learning can improve robustness but does not necessarily lead to improved translation performance. Overall, the experiments demonstrate the importance of proper experimental workflows, which cannot easily generalize from one language pair and scenario to another.
Anthology ID:
2026.eamt-1.13
Volume:
Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 1)
Month:
June
Year:
2026
Address:
Tilburg, The Netherlands
Editors:
Dimitar Shterionov, Eva Vanmassenhove, Mirella De Sisto, Fred Blain, Javad Pourmostafa Roshan Sharami, Lisa Lepp, Chiara Manna, Argentina Anna Rescigno, Alina Karakanta, Ayla Rigouts Terryn, Manuel Lardelli, Natalia Resende, Elena Murgolo, Janiça Hackenbuchner, Anna Zaretskaya, Miquel Esplà-Gomis, Thierry Etchegoyhen, Dagmar Gromann, Rachel Bawden, Barry Haddow, Sara Szoc, Mikel Forcada, Helena Moniz
Venue:
EAMT
SIG:
Publisher:
European Association for Machine Translation
Note:
Pages:
158–172
Language:
URL:
https://aclanthology.org/2026.eamt-1.13/
DOI:
Bibkey:
Cite (ACL):
Mikko Aulamo, Sami Virpioja, Yves Scherrer, and Jörg Tiedemann. 2026. The Challenge of Finding Robust and Efficient Strategies for Training Machine Translation Models with Noisy Data. In Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 1), pages 158–172, Tilburg, The Netherlands. European Association for Machine Translation.
Cite (Informal):
The Challenge of Finding Robust and Efficient Strategies for Training Machine Translation Models with Noisy Data (Aulamo et al., EAMT 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.eamt-1.13.pdf