Nesciun Lengaz Lascià Endò: Machine Translation for Fassa Ladin

Giovanni Valer; Nicolò Penzo; Jacopo Staiano

Nesciun Lengaz Lascià Endò: Machine Translation for Fassa Ladin

Giovanni Valer, Nicolò Penzo, Jacopo Staiano

Abstract

Despite the remarkable success recently obtained by Large Language Models, a significant gap in performance still exists when dealing with low-resource languages which are often poorly supported by off-the-shelf models. In this work we focus on Fassa Ladin, a Rhaeto-Romance linguistic variety spoken by less than ten thousand people in the Dolomitic regions, and set to build the first bidirectional Machine Translation system supporting Italian, English, and Fassa Ladin. To this end, we collected a small though representative corpus compounding 1135 parallel sentences in these three languages, and spanning five domains. We evaluated several models including the open (Meta AI’s No Language Left Behind, NLLB-200) and commercial (OpenAI’s gpt-4o) state-of-the-art, and indeed found that both obtain unsatisfactory performance. We therefore proceeded to finetune the NLLB-200 model on the data collected, using different approaches. We report a comparative analysis of the results obtained, showing that 1) jointly training for multilingual translation (Ladin-Italian and Ladin-English) significantly improves the performance, and 2) knowledge-transfer is highly effective (e.g., leveraging similarities between Ladin and Friulian), highlighting the importance of targeted data collection and model adaptation in the context of low-resource/endangered languages for which little textual data is available.

Anthology ID:: 2024.clicit-1.104
Volume:: Proceedings of the Tenth Italian Conference on Computational Linguistics (CLiC-it 2024)
Month:: December
Year:: 2024
Address:: Pisa, Italy
Editors:: Felice Dell'Orletta, Alessandro Lenci, Simonetta Montemagni, Rachele Sprugnoli
Venue:: CLiC-it
SIG:
Publisher:: CEUR Workshop Proceedings
Note:
Pages:: 967–975
Language:
URL:: https://aclanthology.org/2024.clicit-1.104/
DOI:
Bibkey:
Cite (ACL):: Giovanni Valer, Nicolò Penzo, and Jacopo Staiano. 2024. Nesciun Lengaz Lascià Endò: Machine Translation for Fassa Ladin. In Proceedings of the Tenth Italian Conference on Computational Linguistics (CLiC-it 2024), pages 967–975, Pisa, Italy. CEUR Workshop Proceedings.
Cite (Informal):: Nesciun Lengaz Lascià Endò: Machine Translation for Fassa Ladin (Valer et al., CLiC-it 2024)
Copy Citation:
PDF:: https://aclanthology.org/2024.clicit-1.104.pdf

PDF Cite Search Fix data