DIA2 - a Comprehensive and Diverse Diacritized Arabic Corpus for NLP Research

Fatima Dekmak, Shady Elbassuoni, Khaled Shaban, Hazem Hajj, Wassim El-Hajj, Yasmine Abu Adla, Buthaina Alabrash


Abstract
The development of Arabic natural language processing (NLP) applications and large language models (LLMs) faces substantial challenges, primarily due to the scarcity of high-quality native Arabic datasets. To address this critical gap, we present DIA2 (a Comprehensive and Diverse Diacritized Modern Standard Arabic Corpus), a novel dataset curated from 28 diverse, carefully selected Arabic sources. DIA2 emphasizes the use of original Arabic text and explicitly avoids machine-translated content. The corpus incorporates substantial amounts of text from books, news articles, and poetry, and employs extensive data preprocessing to support NLP research and LLM development. Our preprocessing pipeline includes rigorous text cleaning, URL- and document-level deduplication, and automatic diacritization, while preserving a gold diacritized subset derived from manually annotated sources. The resulting corpus comprises over 140 GB of high-quality text, containing more than 26 million unique words and 41.9 billion tokens. To evaluate the proposed pipeline, we conducted controlled continued pretraining experiments using Llama3.1-8B on both raw and processed subsets of DIA2. The model trained on processed data consistently outperformed its counterpart across multiple Arabic evaluation benchmarks. These results highlight the positive impact of systematic preprocessing and the utility of DIA2 in empowering native Arabic LLMs and downstream NLP tasks.
Anthology ID:
2026.osact-1.14
Volume:
The 7th Workshop on Open-Source Arabic Corpora and Processing Tools (OSACT7) with 5 Shared Tasks
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Hend Al-Khalifa, Mo El-Haj, Saad Ezzini
Venues:
OSACT | WS
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
115–130
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-osact-14
DOI:
10.63317/3k2m7vtzunuk
Bibkey:
Cite (ACL):
Fatima Dekmak, Shady Elbassuoni, Khaled Shaban, Hazem Hajj, Wassim El-Hajj, Yasmine Abu Adla, and Buthaina Alabrash. 2026. DIA2 - a Comprehensive and Diverse Diacritized Arabic Corpus for NLP Research. In The 7th Workshop on Open-Source Arabic Corpora and Processing Tools (OSACT7) with 5 Shared Tasks, pages 115–130, Palma, Mallorca (Spain). Association for Computational Linguistics.
Cite (Informal):
DIA2 - a Comprehensive and Diverse Diacritized Arabic Corpus for NLP Research (Dekmak et al., OSACT 2026)
Copy Citation: