Bengali and Magahi PUD Treebank and Parser

Pritha Majumdar, Deepak Alok, Akanksha Bansal, Atul Kr. Ojha, John P. McCrae


Abstract
This paper presents the development of the Parallel Universal Dependency (PUD) Treebank for two Indo-Aryan languages: Bengali and Magahi. A treebank of 1,000 sentences has been created using a parallel corpus of English and the UD framework. A preliminary set of sentences was annotated manually - 600 for Bengali and 200 for Magahi. The rest of the sentences were built using the Bengali and Magahi parser. The sentences have been translated and annotated manually by the authors, some of whom are also native speakers of the languages. The objective behind this work is to build a syntactically-annotated linguistic repository for the aforementioned languages, that can prove to be a useful resource for building further NLP tools. Additionally, Bengali and Magahi parsers were also created which is built on machine learning approach. The accuracy of the Bengali parser is 78.13% in the case of UPOS; 76.99% in the case of XPOS, 56.12% in the case of UAS; and 47.19% in the case of LAS. The accuracy of Magahi parser is 71.53% in the case of UPOS; 66.44% in the case of XPOS, 58.05% in the case of UAS; and 33.07% in the case of LAS. This paper also includes an illustration of the annotation schema followed, the findings of the Parallel Universal Dependency (PUD) treebank, and it’s resulting linguistic analysis
Anthology ID:
2022.wildre-1.11
Volume:
Proceedings of the WILDRE-6 Workshop within the 13th Language Resources and Evaluation Conference
Month:
June
Year:
2022
Address:
Marseille, France
Editors:
Girish Nath Jha, Sobha L., Kalika Bali, Atul Kr. Ojha
Venue:
WILDRE
SIG:
Publisher:
European Language Resources Association
Note:
Pages:
60–67
Language:
URL:
https://aclanthology.org/2022.wildre-1.11
DOI:
Bibkey:
Cite (ACL):
Pritha Majumdar, Deepak Alok, Akanksha Bansal, Atul Kr. Ojha, and John P. McCrae. 2022. Bengali and Magahi PUD Treebank and Parser. In Proceedings of the WILDRE-6 Workshop within the 13th Language Resources and Evaluation Conference, pages 60–67, Marseille, France. European Language Resources Association.
Cite (Informal):
Bengali and Magahi PUD Treebank and Parser (Majumdar et al., WILDRE 2022)
Copy Citation:
PDF:
https://aclanthology.org/2022.wildre-1.11.pdf