How Foundation Models Behave for Arabic Image Captioning?

Khaoula Dahimi, Amel Belabbaci, Hadda Cherroun, Abdelhamid Haouhat


Abstract
Image captioning plays a crucial role in numerous applications, including educational systems. However, ensuring caption quality remains a significant challenge, particularly for morphologically rich, low-resource languages such as Arabic. We investigate an evaluation of Arabic image captioning using state-of-the-art multimodal foundation models. We systematically assess the performance of leading models—Gemini, Gemma, LLaMA, and Fanar. Our evaluation framework employs a diverse set of metrics spanning rule-based, learnable, visually-grounded, and LLM-based approaches to capture semantic accuracy, linguistic fluency, and hallucination detection. Experiments are conducted on two benchmark datasets: Flickr8k-Arabic and JEEM. Our findings reveal significant performance variations across models and evaluation metrics, highlighting the need for Arabic-specific optimization in multimodal architectures.
Anthology ID:
2026.osact-1.6
Volume:
The 7th Workshop on Open-Source Arabic Corpora and Processing Tools (OSACT7) with 5 Shared Tasks
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Hend Al-Khalifa, Mo El-Haj, Saad Ezzini
Venues:
OSACT | WS
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
49–58
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-osact-06
DOI:
10.63317/3bhwcpon3fv5
Bibkey:
Cite (ACL):
Khaoula Dahimi, Amel Belabbaci, Hadda Cherroun, and Abdelhamid Haouhat. 2026. How Foundation Models Behave for Arabic Image Captioning?. In The 7th Workshop on Open-Source Arabic Corpora and Processing Tools (OSACT7) with 5 Shared Tasks, pages 49–58, Palma, Mallorca (Spain). Association for Computational Linguistics.
Cite (Informal):
How Foundation Models Behave for Arabic Image Captioning? (Dahimi et al., OSACT 2026)
Copy Citation: