An Annotated Corpus of Arabic Tweets for Hate Speech Analysis

Wajdi Zaghouani, Md. Rafiul Biswas


Abstract
Identifying hate speech content in the Arabic language is challenging due to the rich quality of dialectal variations. This study introduces a multilabel hate speech dataset in the Arabic language. We have collected 10,000 Arabic tweets and annotated each tweet, whether it contains offensive content or not. If a text contains offensive content, we further classify it into different hate speech targets such as religion, gender, politics, ethnicity, origin, and others. A text can contain either single or multiple targets. Multiple annotators are involved in the data annotation task. We calculated the inter-annotator agreement, which was reported to be 0.86 for offensive content and 0.71 for multiple hate speech targets. Finally, we evaluated the data annotation task by employing a different transformers-based model in which AraBERTv2 outperformed with a micro-F1 score of 0.7865 and an accuracy of 0.786.
Anthology ID:
2025.ranlp-1.163
Volume:
Proceedings of the 15th International Conference on Recent Advances in Natural Language Processing - Natural Language Processing in the Generative AI Era
Month:
September
Year:
2025
Address:
Varna, Bulgaria
Editors:
Galia Angelova, Maria Kunilovskaya, Marie Escribe, Ruslan Mitkov
Venue:
RANLP
SIG:
Publisher:
INCOMA Ltd., Shoumen, Bulgaria
Note:
Pages:
1413–1419
Language:
URL:
https://aclanthology.org/2025.ranlp-1.163/
DOI:
Bibkey:
Cite (ACL):
Wajdi Zaghouani and Md. Rafiul Biswas. 2025. An Annotated Corpus of Arabic Tweets for Hate Speech Analysis. In Proceedings of the 15th International Conference on Recent Advances in Natural Language Processing - Natural Language Processing in the Generative AI Era, pages 1413–1419, Varna, Bulgaria. INCOMA Ltd., Shoumen, Bulgaria.
Cite (Informal):
An Annotated Corpus of Arabic Tweets for Hate Speech Analysis (Zaghouani & Biswas, RANLP 2025)
Copy Citation:
PDF:
https://aclanthology.org/2025.ranlp-1.163.pdf