Lemmatising Serbian as Category Tagging with Bidirectional Sequence Classification

Andrea Gesmundo, Tanja Samardžić


Abstract
We present a novel tool for morphological analysis of Serbian, which is a low-resource language with rich morphology. Our tool produces lemmatisation and morphological analysis reaching accuracy that is considerably higher compared to the existing alternative tools: 83.6% relative error reduction on lemmatisation and 8.1% relative error reduction on morphological analysis. The system is trained on a small manually annotated corpus with an approach based on Bidirectional Sequence Classification and Guided Learning techniques, which have recently been adapted with success to a broad set of NLP tagging tasks. In the system presented in this paper, this general approach to tagging is applied to the lemmatisation task for the first time thanks to our novel formulation of lemmatisation as a category tagging task. We show that learning lemmatisation rules from annotated corpus and integrating the context information in the process of morphological analysis provides a state-of-the-art performance despite the lack of resources. The proposed system can be used via a web GUI that deploys its best scoring configuration
Anthology ID:
L12-1411
Volume:
Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12)
Month:
May
Year:
2012
Address:
Istanbul, Turkey
Editors:
Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Mehmet Uğur Doğan, Bente Maegaard, Joseph Mariani, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association (ELRA)
Note:
Pages:
2103–2106
Language:
URL:
http://www.lrec-conf.org/proceedings/lrec2012/pdf/708_Paper.pdf
DOI:
Bibkey:
Cite (ACL):
Andrea Gesmundo and Tanja Samardžić. 2012. Lemmatising Serbian as Category Tagging with Bidirectional Sequence Classification. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12), pages 2103–2106, Istanbul, Turkey. European Language Resources Association (ELRA).
Cite (Informal):
Lemmatising Serbian as Category Tagging with Bidirectional Sequence Classification (Gesmundo & Samardžić, LREC 2012)
Copy Citation:
PDF:
http://www.lrec-conf.org/proceedings/lrec2012/pdf/708_Paper.pdf