Sentiment Analysis of Arabic Tweets Using Large Language Models

Pankaj Dadure; Ananya Dixit; Kunal Tewatia; Nandini Paliwal; Anshika Malla

Sentiment Analysis of Arabic Tweets Using Large Language Models

Pankaj Dadure, Ananya Dixit, Kunal Tewatia, Nandini Paliwal, Anshika Malla

Abstract

In the digital era, sentiment analysis has become an indispensable tool for understanding public sentiments, optimizing market strategies, and enhancing customer engagement across diverse sectors. While significant advancements have been made in sentiment analysis for high-resource languages such as English, French, etc. This study focuses on Arabic, a low-resource language, to address its unique challenges like morphological complexity, diverse dialects, and limited linguistic resources. Existing works in Arabic sentiment analysis have utilized deep learning architectures like LSTM, BiLSTM, and CNN-LSTM, alongside embedding techniques such as Word2Vec and contextualized models like ARABERT. Building on this foundation, our research investigates sentiment classification of Arabic tweets, categorizing them as positive or negative, using embeddings derived from three large language models (LLMs): Universal Sentence Encoder (USE), XLM-RoBERTa base (XLM-R base), and MiniLM-L12-v2. Experimental results demonstrate that incorporating emojis in the dataset and using the MiniLM embeddings yield an accuracy of 85.98%. In contrast, excluding emojis and using embeddings from the XLM-R base resulted in a lower accuracy of 78.98%. These findings highlight the impact of both dataset composition and embedding techniques on Arabic sentiment analysis performance.

Anthology ID:: 2025.abjadnlp-1.10
Volume:: Proceedings of the 1st Workshop on NLP for Languages Using Arabic Script
Month:: January
Year:: 2025
Address:: Abu Dhabi, UAE
Editor:: Mo El-Haj
Venues:: AbjadNLP | WS
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 88–94
Language:
URL:: https://aclanthology.org/2025.abjadnlp-1.10/
DOI:
Bibkey:
Cite (ACL):: Pankaj Dadure, Ananya Dixit, Kunal Tewatia, Nandini Paliwal, and Anshika Malla. 2025. Sentiment Analysis of Arabic Tweets Using Large Language Models. In Proceedings of the 1st Workshop on NLP for Languages Using Arabic Script, pages 88–94, Abu Dhabi, UAE. Association for Computational Linguistics.
Cite (Informal):: Sentiment Analysis of Arabic Tweets Using Large Language Models (Dadure et al., AbjadNLP 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.abjadnlp-1.10.pdf

PDF Cite Search Fix data