Building Arabic NLP from the Ground Up: Twenty Years of Lessons, Failures, and Open Problems

Wajdi Zaghouani

doi:10.18653/v1/2026.bigpicture-main.8

Building Arabic NLP from the Ground Up: Twenty Years of Lessons, Failures, and Open Problems

Abstract

This paper reflects on twenty years of building NLP resources and research infrastructure for Arabic, a language spoken by hundreds of millions yet historically underserved relative to languages such as English or Chinese. The first decade focused on foundational linguistic infrastructure; the second shifted toward computational social science, social media analysis, and socially oriented applications. Rather than cataloguing outputs, the paper examines what the experience of building them revealed. Three counterintuitive lessons emerge: building datasets is as much a social process as a technical one; communities formed around shared tasks often matter more than the tasks themselves; and moving from language resources to computational social science exposes challenges that traditional NLP training does not address. We discuss three failures: a depression detection corpus that never reached clinical practice, a period of spreading across too many shared tasks without sufficient depth, and a long-standing assumption that Modern Standard Arabic infrastructure would transfer cleanly to dialectal tasks. These experiences suggest that the hardest problems in developing NLP for underserved communities are not linguistic but social, institutional, and epistemic, and require competencies the field rarely teaches.

Anthology ID:: 2026.bigpicture-main.8
Volume:: Proceedings of The Big Picture v2: Crafting a Research Narrative
Month:: July
Year:: 2026
Address:: San Diego, CA, USA
Editors:: Yanai Elazar, Allyson Ettinger, Nora Kassner, Sebastian Ruder
Venues:: BigPicture | WS
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 94–106
Language:
URL:: https://aclanthology.org/2026.bigpicture-main.8/
DOI:: 10.18653/v1/2026.bigpicture-main.8
Bibkey:
Cite (ACL):: Wajdi Zaghouani. 2026. Building Arabic NLP from the Ground Up: Twenty Years of Lessons, Failures, and Open Problems. In Proceedings of The Big Picture v2: Crafting a Research Narrative, pages 94–106, San Diego, CA, USA. Association for Computational Linguistics.
Cite (Informal):: Building Arabic NLP from the Ground Up: Twenty Years of Lessons, Failures, and Open Problems (Zaghouani, BigPicture 2026)
Copy Citation:
PDF:: https://aclanthology.org/2026.bigpicture-main.8.pdf

PDF Cite Search Fix data