Towards Automatic Finnish Text Simplification

Anna Dmitrieva, Jörg Tiedemann


Abstract
Automatic text simplification (ATS/TS) models typically require substantial parallel training data. This paper describes our work on expanding the Finnish-Easy Finnish parallel corpus and making baseline simplification models. We discuss different approaches to document and sentence alignment. After finding the optimal alignment methodologies, we increase the amount of document-aligned data 6.5 times and add a sentence-aligned version of the dataset consisting of more than twelve thousand sentence pairs. Using sentence-aligned data, we fine-tune two models for text simplification. The first is mBART, a sequence-to-sequence translation architecture proven to show good results for monolingual translation tasks. The second is the Finnish GPT model, for which we utilize instruction fine-tuning. This work is the first attempt to create simplification models for Finnish using monolingual parallel data in this language. The data has been deposited in the Finnish Language Bank (Kielipankki) and is available for non-commercial use, and the models will be made accessible through either Kielipankki or public repositories such as Huggingface or GitHub.
Anthology ID:
2024.determit-1.4
Volume:
Proceedings of the Workshop on DeTermIt! Evaluating Text Difficulty in a Multilingual Context @ LREC-COLING 2024
Month:
May
Year:
2024
Address:
Torino, Italia
Editors:
Giorgio Maria Di Nunzio, Federica Vezzani, Liana Ermakova, Hosein Azarbonyad, Jaap Kamps
Venues:
DeTermIt | WS
SIG:
Publisher:
ELRA and ICCL
Note:
Pages:
39–50
Language:
URL:
https://aclanthology.org/2024.determit-1.4
DOI:
Bibkey:
Cite (ACL):
Anna Dmitrieva and Jörg Tiedemann. 2024. Towards Automatic Finnish Text Simplification. In Proceedings of the Workshop on DeTermIt! Evaluating Text Difficulty in a Multilingual Context @ LREC-COLING 2024, pages 39–50, Torino, Italia. ELRA and ICCL.
Cite (Informal):
Towards Automatic Finnish Text Simplification (Dmitrieva & Tiedemann, DeTermIt-WS 2024)
Copy Citation:
PDF:
https://aclanthology.org/2024.determit-1.4.pdf