A Large and Balanced Multi-Domain Arabic Corpus Annotated for Morphology, Syntax, and Readability

Khalid N. Elmadani, Adel Mahmoud Wizani, Hanada Taha Thomure, Nizar Habash


Abstract
We present BAREC-10M, an expanded version of the Balanced Arabic Readability Evaluation Corpus (BAREC). This new release extends the original 1M-word corpus to 10 million words and broadens its scope to include balanced multi-domain coverage annotated for morphology, syntax, and readability. The corpus integrates 45 sub-corpora drawn from diverse sources, including news, educational materials, literature, children’s texts, and religious discourse. Each text is labeled for domain, readership level, and genre, and automatically analyzed using state-of-the-art morphological and syntactic tools. To enhance coverage of underrepresented varieties, we manually digitized and included children’s materials, magazines, and curriculum-based content. The resulting dataset provides a balanced resource for studying Arabic linguistic variation across styles, audiences, and levels of complexity.
Anthology ID:
2026.lrec-1.921
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
11761–11775
Language:
External URL:
https://lrec.elra.info/lrec2026-main-921
DOI:
10.63317/45f2o6t8piyi
Bibkey:
Cite (ACL):
Khalid N. Elmadani, Adel Mahmoud Wizani, Hanada Taha Thomure, and Nizar Habash. 2026. A Large and Balanced Multi-Domain Arabic Corpus Annotated for Morphology, Syntax, and Readability. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 11761–11775, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
A Large and Balanced Multi-Domain Arabic Corpus Annotated for Morphology, Syntax, and Readability (Elmadani et al., LREC 2026)
Copy Citation: