Comprehensive Part-Of-Speech Tag Set and SVM based POS Tagger for Sinhala

Sandareka Fernando, Surangika Ranathunga, Sanath Jayasena, Gihan Dias


Abstract
This paper presents a new comprehensive multi-level Part-Of-Speech tag set and a Support Vector Machine based Part-Of-Speech tagger for the Sinhala language. The currently available tag set for Sinhala has two limitations: the unavailability of tags to represent some word classes and the lack of tags to capture inflection based grammatical variations of words. The new tag set, presented in this paper overcomes both of these limitations. The accuracy of available Sinhala Part-Of-Speech taggers, which are based on Hidden Markov Models, still falls far behind state of the art. Our Support Vector Machine based tagger achieved an overall accuracy of 84.68% with 59.86% accuracy for unknown words and 87.12% for known words, when the test set contains 10% of unknown words.
Anthology ID:
W16-3718
Volume:
Proceedings of the 6th Workshop on South and Southeast Asian Natural Language Processing (WSSANLP2016)
Month:
December
Year:
2016
Address:
Osaka, Japan
Editors:
Dekai Wu, Pushpak Bhattacharyya
Venue:
WSSANLP
SIG:
Publisher:
The COLING 2016 Organizing Committee
Note:
Pages:
173–182
Language:
URL:
https://aclanthology.org/W16-3718/
DOI:
Bibkey:
Cite (ACL):
Sandareka Fernando, Surangika Ranathunga, Sanath Jayasena, and Gihan Dias. 2016. Comprehensive Part-Of-Speech Tag Set and SVM based POS Tagger for Sinhala. In Proceedings of the 6th Workshop on South and Southeast Asian Natural Language Processing (WSSANLP2016), pages 173–182, Osaka, Japan. The COLING 2016 Organizing Committee.
Cite (Informal):
Comprehensive Part-Of-Speech Tag Set and SVM based POS Tagger for Sinhala (Fernando et al., WSSANLP 2016)
Copy Citation:
PDF:
https://aclanthology.org/W16-3718.pdf