When Bigger Isn’t Better: Evaluating LLMs for Arabic Sentiment Analysis

Mohamed Ibrahim, Abdullah Makki, Youssef Barakat, Nour Samy, Sarah AlHumoud


Abstract
This study evaluates the performance of a fine-tuned Arabic sentiment transformer (CAMeL-MSA) against eight large language models (LLMs). Using zero-shot prompting across six Arabic sentiment datasets, we compare a specialized, task-specific approach against generalized model capabilities. Results show that the fine-tuned baseline substantially outperformed all LLMs on five of the six datasets in both accuracy and Macro F1-score. While LLMs offer versatility, this comparison highlights the continued practical superiority of task-specific fine-tuning over zero-shot prompting.
Anthology ID:
2026.osact-1.4
Volume:
The 7th Workshop on Open-Source Arabic Corpora and Processing Tools (OSACT7) with 5 Shared Tasks
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Hend Al-Khalifa, Mo El-Haj, Saad Ezzini
Venues:
OSACT | WS
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
35–39
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-osact-04
DOI:
10.63317/2kimof4u6y8x
Bibkey:
Cite (ACL):
Mohamed Ibrahim, Abdullah Makki, Youssef Barakat, Nour Samy, and Sarah AlHumoud. 2026. When Bigger Isn’t Better: Evaluating LLMs for Arabic Sentiment Analysis. In The 7th Workshop on Open-Source Arabic Corpora and Processing Tools (OSACT7) with 5 Shared Tasks, pages 35–39, Palma, Mallorca (Spain). Association for Computational Linguistics.
Cite (Informal):
When Bigger Isn’t Better: Evaluating LLMs for Arabic Sentiment Analysis (Ibrahim et al., OSACT 2026)
Copy Citation: