Sardar Jaf


2016

pdf bib
A Simple Approach to Unifying Ambiguously Encoded Kurdish Characters
Sardar Jaf
Proceedings of the Second International Conference on Computational Linguistics in Bulgaria (CLIB 2016)

In this study we outline a potential problem in the normalisation stage of processing texts that are based on a modified version of the Arabic alphabet. The main source of resources available for processing resource-scarce languages is raw text. We have identified an interesting challenge that must be addressed when normalising certain natural language texts. Many less-resourced languages, such as Kurdish, Farsi, Urdu, Pashtu, etc., use a modified version of the Arabic writing system. Many characters in harvested data from the Internet may have exactly the same form but encoded with different Unicode values (ambiguous characters). It is important to identify ambiguous characters during the normalisation stage of most text processing tasks. We will demonstrate cases related to ambiguous Kurdish and Farsi characters and propose a semi-automatic approach to identifying and unifying ambiguously encoded characters.

2015

pdf bib
The Application of Constraint Rules to Data-driven Parsing
Sardar Jaf | Allan Ramsay
Proceedings of the International Conference Recent Advances in Natural Language Processing

Search
Co-authors
Venues