Parsing Arabic Dialects Revisited: New Benchmarks, Models, and Insights

Ahmed Farouk Zakaria Elshabrawy, Go Inoue, Muhammed AbuOdeh, Nizar Habash


Abstract
Parsing dialectal Arabic remains underexplored, with limited progress over the past two decades. Existing Modern Standard Arabic (MSA) parsers perform poorly on dialectal data, motivating the need for dialect-specific approaches. We revisit this task using modern neural models and present new results on Egyptian and Gulf Arabic dependency parsing. We demonstrate that even small amounts of dialectal training data yield substantial improvements in parsing accuracy. Our contributions include: (1) introducing a new annotated dataset for Gulf Arabic, (2) releasing a state-of-the-art multi-variety Arabic parser, and (3) employing dialect identification as a diagnostic tool to better understand how training data affects parsing performance across dialects and test sets.
Anthology ID:
2026.osact-1.12
Volume:
The 7th Workshop on Open-Source Arabic Corpora and Processing Tools (OSACT7) with 5 Shared Tasks
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Hend Al-Khalifa, Mo El-Haj, Saad Ezzini
Venues:
OSACT | WS
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
94–105
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-osact-12
DOI:
10.63317/3eyeu3k726ab
Bibkey:
Cite (ACL):
Ahmed Farouk Zakaria Elshabrawy, Go Inoue, Muhammed AbuOdeh, and Nizar Habash. 2026. Parsing Arabic Dialects Revisited: New Benchmarks, Models, and Insights. In The 7th Workshop on Open-Source Arabic Corpora and Processing Tools (OSACT7) with 5 Shared Tasks, pages 94–105, Palma, Mallorca (Spain). Association for Computational Linguistics.
Cite (Informal):
Parsing Arabic Dialects Revisited: New Benchmarks, Models, and Insights (Farouk Zakaria Elshabrawy et al., OSACT 2026)
Copy Citation: