Cross-Dialectal Transfer for Low-Resource Arabic: The Tunisian Arabic Dependency Treebank

Amal Aissaoui


Abstract
This paper presents a small-scale dependency treebank for Tunisian Arabic (TADT) developed within the Universal Dependencies framework, addressing the scarcity of linguistic resources for the Arabic varieties. The approach employs domain adaptation, leveraging a machine learning model (UDPipe 1.0) trained on Algerian Arabic data to annotate 100 Tunisian Arabic social media comments, followed by manual correction. This pilot study evaluates the feasibility of using machine learning-assisted annotation to scale resource development for spoken Arabic and identifies key challenges in cross-dialectal transfer for improving annotation quality and efficiency. This work contributes to more inclusive and fair representation of Arabic linguistic varieties in academic research and NLP applications.
Anthology ID:
2026.udw-1.23
Volume:
Proceedings of the Ninth Workshop on Universal Dependencies (UDW 2026)
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Çağrı Çöltekin, Kaja Dobrovoljc
Venues:
UDW | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
258–267
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-udw-23
DOI:
10.63317/4vghkox8ptiq
Bibkey:
Cite (ACL):
Amal Aissaoui. 2026. Cross-Dialectal Transfer for Low-Resource Arabic: The Tunisian Arabic Dependency Treebank. In Proceedings of the Ninth Workshop on Universal Dependencies (UDW 2026), pages 258–267, Palma de Mallorca, Spain. ELRA Language Resources Association (ELRA).
Cite (Informal):
Cross-Dialectal Transfer for Low-Resource Arabic: The Tunisian Arabic Dependency Treebank (Aissaoui, UDW 2026)
Copy Citation: