CoDeRooMor: A new dataset for non-inflectional morphology studies of Swedish

Elena Volodina, Yousuf Ali Mohammed, Therese Lindström Tiedemann


Abstract
The paper introduces a new resource, CoDeRooMor, for studying the morphology of modern Swedish word formation. The approximately 16.000 lexical items in the resource have been manually segmented into word-formation morphemes, and labeled for their categories, such as prefixes, suffixes, roots, etc. Word-formation mechanisms, such as derivation and compounding have been associated with each item on the list. The article describes the selection of items for manual annotation and the principles of annotation, reports on the reliability of the manual annotation, and presents tools, resources and some first statistics. Given the”gold” nature of the resource, it is possible to use it for empirical studies as well as to develop linguistically-aware algorithms for morpheme segmentation and labeling (cf statistical subword approach). The resource will be made freely available.
Anthology ID:
2021.nodalida-main.18
Volume:
Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa)
Month:
May 31--2 June
Year:
2021
Address:
Reykjavik, Iceland (Online)
Venue:
NoDaLiDa
SIG:
Publisher:
Linköping University Electronic Press, Sweden
Note:
Pages:
178–189
Language:
URL:
https://aclanthology.org/2021.nodalida-main.18
DOI:
Bibkey:
Cite (ACL):
Elena Volodina, Yousuf Ali Mohammed, and Therese Lindström Tiedemann. 2021. CoDeRooMor: A new dataset for non-inflectional morphology studies of Swedish. In Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa), pages 178–189, Reykjavik, Iceland (Online). Linköping University Electronic Press, Sweden.
Cite (Informal):
CoDeRooMor: A new dataset for non-inflectional morphology studies of Swedish (Volodina et al., NoDaLiDa 2021)
Copy Citation:
PDF:
https://aclanthology.org/2021.nodalida-main.18.pdf