A Comprehensive Full-Form Lexicon for Arabic NLP and Speech Technology

Yannis Haralambous, Jack Halpern


Abstract
Natural Language Processing (NLP) applications require morphological data with precise grammatical attributes, while speech technology requires abundant phonemic and phonetic data. This presents a challenge for Arabic due to its abundant morphological, orthographic, and phonemic ambiguity in both MSA and its various dialects. Existing systems struggle with incomplete and unstructured web data, leading to suboptimal performance in both morphological analysis and speech applications. This paper presents ArabLEX, a full-form lexicon (includes all wordforms, i.e., fully inflected/cliticized members of a lexeme class) that addresses these issues by providing a large-scale database designed to enhance NLP accuracy. It comprises approximately 570 million entries with fully inflected forms and detailed morphological, phonetic, and orthographic attributes. ArabLEX serves as a foundational framework for developing comprehensive Arabic lexical resources for NLP, particularly for speech technology, as well as dialect databases.
Anthology ID:
2026.lrec-1.108
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
1382–1393
Language:
External URL:
https://lrec.elra.info/lrec2026-main-108
DOI:
10.63317/2gbvmmu4ix5e
Bibkey:
Cite (ACL):
Yannis Haralambous and Jack Halpern. 2026. A Comprehensive Full-Form Lexicon for Arabic NLP and Speech Technology. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 1382–1393, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
A Comprehensive Full-Form Lexicon for Arabic NLP and Speech Technology (Haralambous & Halpern, LREC 2026)
Copy Citation: