Hate Speech Classification in Bulgarian

Radoslav Ralev; Jürgen Pfeffer

Hate Speech Classification in Bulgarian

Abstract

In recent years, we have seen a surge in the propagation of online hate speech on social media platforms. According to a multitude of sources such as the European Council, hate speech can lead to acts of violence and conflict on a broader scale. That has led to in- creased awareness by governments, companies, and the scientific community, and although the field is relatively new, there have been considerable advancements in the field as a result of the collective effort. Despite the increasingly better results, most of the research focuses on the more popular languages (i.e., English, German, or Arabic), whereas less popular languages such as Bulgarian and other Balkan languages have been neglected. We have aggregated a real-world dataset from Bulgarian online forums and manually annotated 108,142 sentences. About 1.74% of which can be described with the categories racism, sexism, rudeness, and profanity. We then developed and evaluated various classifiers on the dataset and found that a support vector machine with a linear kernel trained on character-level TF-IDF features is the best model. Our work can be seen as another piece in the puzzle to building a strong foundation for future work on hate speech classification in Bulgarian.

Anthology ID:: 2022.clib-1.5
Volume:: Proceedings of the Fifth International Conference on Computational Linguistics in Bulgaria (CLIB 2022)
Month:: September
Year:: 2022
Address:: Sofia, Bulgaria
Venue:: CLIB
SIG:
Publisher:: Department of Computational Linguistics, IBL -- BAS
Note:
Pages:: 49–58
Language:
URL:: https://aclanthology.org/2022.clib-1.5/
DOI:
Bibkey:
Cite (ACL):: Radoslav Ralev and Jürgen Pfeffer. 2022. Hate Speech Classification in Bulgarian. In Proceedings of the Fifth International Conference on Computational Linguistics in Bulgaria (CLIB 2022), pages 49–58, Sofia, Bulgaria. Department of Computational Linguistics, IBL -- BAS.
Cite (Informal):: Hate Speech Classification in Bulgarian (Ralev & Pfeffer, CLIB 2022)
Copy Citation:
PDF:: https://aclanthology.org/2022.clib-1.5.pdf

PDF Cite Search Fix data