SinMix2Mono: A Dataset for Code-mixed Romanized Sinhala Translation and Transliteration

Rukshan Dias, Deshan Sumanathilaka, Archchana Sindhujan, Minidu Nimna


Abstract
Code-mixed and Romanized texts are widely used in digital content, yet they remain largely underexplored for many low-resource languages, including Sinhala. The scarcity of high-quality parallel data has limited progress on downstream tasks, such as machine translation and transliteration. We introduce SinMix2Mono, the largest manually annotated parallel training dataset, followed by the first gold standard benchmark and code-mixed transliteration ambiguity corpora for code-mixed romanized Sinhala to Sinhala conversion. The dataset comprises approximately 25,000 real-world sentences collected from social media, covering diverse domains and authentic code-mixing patterns. To ensure high-quality translations, we used an annotation pipeline that combined rule-based transliteration, LLM-assisted translation, and human validation. The golden test dataset, which includes 2549 sentences, and the code-mixed transliteration ambiguity test were validated by three annotators, yielding Gwet’s AC1 scores of 0.7465 and 0.7068, respectively. We benchmarked nine systems, including statistical, neural and commercial LLMs. SinMix2Mono provides a robust training and evaluation resource, establishing a strong benchmark for future research on Sinhala code-mixed translation and transliteration.
Anthology ID:
2026.eamt-1.10
Volume:
Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 1)
Month:
June
Year:
2026
Address:
Tilburg, The Netherlands
Editors:
Dimitar Shterionov, Eva Vanmassenhove, Mirella De Sisto, Fred Blain, Javad Pourmostafa Roshan Sharami, Lisa Lepp, Chiara Manna, Argentina Anna Rescigno, Alina Karakanta, Ayla Rigouts Terryn, Manuel Lardelli, Natalia Resende, Elena Murgolo, Janiça Hackenbuchner, Anna Zaretskaya, Miquel Esplà-Gomis, Thierry Etchegoyhen, Dagmar Gromann, Rachel Bawden, Barry Haddow, Sara Szoc, Mikel Forcada, Helena Moniz
Venue:
EAMT
SIG:
Publisher:
European Association for Machine Translation
Note:
Pages:
114–129
Language:
URL:
https://aclanthology.org/2026.eamt-1.10/
DOI:
Bibkey:
Cite (ACL):
Rukshan Dias, Deshan Sumanathilaka, Archchana Sindhujan, and Minidu Nimna. 2026. SinMix2Mono: A Dataset for Code-mixed Romanized Sinhala Translation and Transliteration. In Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 1), pages 114–129, Tilburg, The Netherlands. European Association for Machine Translation.
Cite (Informal):
SinMix2Mono: A Dataset for Code-mixed Romanized Sinhala Translation and Transliteration (Dias et al., EAMT 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.eamt-1.10.pdf