Marta Ba Nón
2023
MaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages
Marta Ba Nón
|
Mălina Chichirău
|
Miquel Esplà-Gomis
|
Mikel Forcada
|
Aarón Galiano-Jiménez
|
Taja Kuzman
|
Nikola Ljubešić
|
Rik van Noord
|
Leopoldo Pla Sempere
|
Gema Ramírez-Sánchez
|
Peter Rupnik
|
Vit Suchomel
|
Antonio Toral
|
Jaume Zaragoza-Bernabeu
Proceedings of the 24th Annual Conference of the European Association for Machine Translation
We present the most relevant results of the project MaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages in its second year. To date, parallel and monolingual corpora have been produced for seven low-resourced European languages by crawling large amounts of textual data from selected top-level domains of the Internet; both human and automatic evaluation show its usefulness. In addition, several large language models pretrained on MaCoCu data have been published, as well as the code used to collect and curate the data.