AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages

Shamsuddeen Muhammad; Idris Abdulmumin; Abinew Ayele; Nedjma Ousidhoum; David Adelani; Seid Yimam; Ibrahim Ahmad; Meriem Beloucif; Saif Mohammad; Sebastian Ruder; Oumaima Hourrane; Alipio Jorge; Pavel Brazdil; Felermino Ali; Davis David; Salomey Osei; Bello Shehu Bello; Falalu Lawan; Tajuddeen Gwadabe; Samuel Rutunda; Tadesse Belay; Wendimu Messelle; Hailu Balcha; Sisay Chala; Hagos Gebremichael; Bernard Opoku; Stephen Arthur

doi:10.18653/v1/2023.emnlp-main.862

AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages

Shamsuddeen Muhammad, Idris Abdulmumin, Abinew Ayele, Nedjma Ousidhoum, David Adelani, Seid Yimam, Ibrahim Ahmad, Meriem Beloucif, Saif Mohammad, Sebastian Ruder, Oumaima Hourrane, Alipio Jorge, Pavel Brazdil, Felermino Ali, Davis David, Salomey Osei, Bello Shehu-Bello, Falalu Lawan, Tajuddeen Gwadabe, Samuel Rutunda, Tadesse Belay, Wendimu Messelle, Hailu Balcha, Sisay Chala, Hagos Gebremichael, Bernard Opoku, Stephen Arthur

Abstract

Africa is home to over 2,000 languages from over six language families and has the highest linguistic diversity among all continents. This includes 75 languages with at least one million speakers each. Yet, there is little NLP research conducted on African languages. Crucial in enabling such research is the availability of high-quality annotated datasets. In this paper, we introduce AfriSenti, a sentiment analysis benchmark that contains a total of >110,000 tweets in 14 African languages (Amharic, Algerian Arabic, Hausa, Igbo, Kinyarwanda, Moroccan Arabic, Mozambican Portuguese, Nigerian Pidgin, Oromo, Swahili, Tigrinya, Twi, Xitsonga, and Yoruba) from four language families. The tweets were annotated by native speakers and used in the AfriSenti-SemEval shared task (with over 200 participants, see website: https://afrisenti-semeval.github.io). We describe the data collection methodology, annotation process, and the challenges we dealt with when curating each dataset. We further report baseline experiments conducted on the AfriSenti datasets and discuss their usefulness.

Anthology ID:: 2023.emnlp-main.862
Volume:: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing
Month:: December
Year:: 2023
Address:: Singapore
Editors:: Houda Bouamor, Juan Pino, Kalika Bali
Venue:: EMNLP
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 13968–13981
Language:
URL:: https://aclanthology.org/2023.emnlp-main.862
DOI:: 10.18653/v1/2023.emnlp-main.862
Bibkey:
Cite (ACL):: Shamsuddeen Muhammad, Idris Abdulmumin, Abinew Ayele, Nedjma Ousidhoum, David Adelani, Seid Yimam, Ibrahim Ahmad, Meriem Beloucif, Saif Mohammad, Sebastian Ruder, Oumaima Hourrane, Alipio Jorge, Pavel Brazdil, Felermino Ali, Davis David, Salomey Osei, Bello Shehu-Bello, Falalu Lawan, Tajuddeen Gwadabe, et al.. 2023. AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 13968–13981, Singapore. Association for Computational Linguistics.
Cite (Informal):: AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages (Muhammad et al., EMNLP 2023)
Copy Citation:
PDF:: https://aclanthology.org/2023.emnlp-main.862.pdf
Video:: https://aclanthology.org/2023.emnlp-main.862.mp4

PDF Cite Search Video