Towards a Standardized Dataset on Indonesian Named Entity Recognition

Siti Oryza Khairunnisa; Aizhan Imankulova; Mamoru Komachi

Towards a Standardized Dataset on Indonesian Named Entity Recognition

Siti Oryza Khairunnisa, Aizhan Imankulova, Mamoru Komachi

Abstract

In recent years, named entity recognition (NER) tasks in the Indonesian language have undergone extensive development. There are only a few corpora for Indonesian NER; hence, recent Indonesian NER studies have used diverse datasets. Although an open dataset is available, it includes only approximately 2,000 sentences and contains inconsistent annotations, thereby preventing accurate training of NER models without reliance on pre-trained models. Therefore, we re-annotated the dataset and compared the two annotations’ performance using the Bidirectional Long Short-Term Memory and Conditional Random Field (BiLSTM-CRF) approach. Fixing the annotation yielded a more consistent result for the organization tag and improved the prediction score by a large margin. Moreover, to take full advantage of pre-trained models, we compared different feature embeddings to determine their impact on the NER task for the Indonesian language.

Anthology ID:: 2020.aacl-srw.10
Volume:: Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing: Student Research Workshop
Month:: December
Year:: 2020
Address:: Suzhou, China
Editors:: Boaz Shmueli, Yin Jou Huang
Venue:: AACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 64–71
Language:
URL:: https://aclanthology.org/2020.aacl-srw.10
DOI:
Bibkey:
Cite (ACL):: Siti Oryza Khairunnisa, Aizhan Imankulova, and Mamoru Komachi. 2020. Towards a Standardized Dataset on Indonesian Named Entity Recognition. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing: Student Research Workshop, pages 64–71, Suzhou, China. Association for Computational Linguistics.
Cite (Informal):: Towards a Standardized Dataset on Indonesian Named Entity Recognition (Khairunnisa et al., AACL 2020)
Copy Citation:
PDF:: https://aclanthology.org/2020.aacl-srw.10.pdf
Code: khairunnisaor/idner-news-2k

PDF Cite Search Code