Evaluation of Off-the-Shelf Language Identification Tools on Bulgarian Social Media Posts

Silvia Gargova, Irina Temnikova, Ivo Dzhumerov, Hristiana Nikolaeva


Abstract
Automatic Language Identification (LI) is a widely addressed task, but not all users (for example linguists) have the means or interest to develop their own tool or to train the existing ones with their own data. There are several off-the-shelf LI tools, but for some languages, it is unclear which tool is the best for specific types of text. This article presents a comparison of the performance of several off-the-shelf language identification tools on Bulgarian social media data. The LI tools are tested on a multilingual Twitter dataset (composed of 2966 tweets) and an existing Bulgarian Twitter dataset on the topic of fake content detection of 3350 tweets. The article presents the manual annotation procedure of the first dataset, a dis- cussion of the decisions of the two annotators, and the results from testing the 7 off-the-shelf LI tools on both datasets. Our findings show that the tool, which is the easiest for users with no programming skills, achieves the highest F1-Score on Bulgarian social media data, while other tools have very useful functionalities for Bulgarian social media texts.
Anthology ID:
2022.clib-1.18
Volume:
Proceedings of the 5th International Conference on Computational Linguistics in Bulgaria (CLIB 2022)
Month:
September
Year:
2022
Address:
Sofia, Bulgaria
Venue:
CLIB
SIG:
Publisher:
Department of Computational Linguistics, IBL -- BAS
Note:
Pages:
152–161
Language:
URL:
https://aclanthology.org/2022.clib-1.18
DOI:
Bibkey:
Cite (ACL):
Silvia Gargova, Irina Temnikova, Ivo Dzhumerov, and Hristiana Nikolaeva. 2022. Evaluation of Off-the-Shelf Language Identification Tools on Bulgarian Social Media Posts. In Proceedings of the 5th International Conference on Computational Linguistics in Bulgaria (CLIB 2022), pages 152–161, Sofia, Bulgaria. Department of Computational Linguistics, IBL -- BAS.
Cite (Informal):
Evaluation of Off-the-Shelf Language Identification Tools on Bulgarian Social Media Posts (Gargova et al., CLIB 2022)
Copy Citation:
PDF:
https://aclanthology.org/2022.clib-1.18.pdf