A Simple Approach to Unifying Ambiguously Encoded Kurdish Characters

Sardar Jaf


Abstract
In this study we outline a potential problem in the normalisation stage of processing texts that are based on a modified version of the Arabic alphabet. The main source of resources available for processing resource-scarce languages is raw text. We have identified an interesting challenge that must be addressed when normalising certain natural language texts. Many less-resourced languages, such as Kurdish, Farsi, Urdu, Pashtu, etc., use a modified version of the Arabic writing system. Many characters in harvested data from the Internet may have exactly the same form but encoded with different Unicode values (ambiguous characters). It is important to identify ambiguous characters during the normalisation stage of most text processing tasks. We will demonstrate cases related to ambiguous Kurdish and Farsi characters and propose a semi-automatic approach to identifying and unifying ambiguously encoded characters.
Anthology ID:
2016.clib-1.11
Volume:
Proceedings of the Second International Conference on Computational Linguistics in Bulgaria (CLIB 2016)
Month:
September
Year:
2016
Address:
Sofia, Bulgaria
Venue:
CLIB
SIG:
Publisher:
Department of Computational Linguistics, Institute for Bulgarian Language, Bulgarian Academy of Sciences
Note:
Pages:
86–94
Language:
URL:
https://aclanthology.org/2016.clib-1.11
DOI:
Bibkey:
Cite (ACL):
Sardar Jaf. 2016. A Simple Approach to Unifying Ambiguously Encoded Kurdish Characters. In Proceedings of the Second International Conference on Computational Linguistics in Bulgaria (CLIB 2016), pages 86–94, Sofia, Bulgaria. Department of Computational Linguistics, Institute for Bulgarian Language, Bulgarian Academy of Sciences.
Cite (Informal):
A Simple Approach to Unifying Ambiguously Encoded Kurdish Characters (Jaf, CLIB 2016)
Copy Citation:
PDF:
https://aclanthology.org/2016.clib-1.11.pdf