Towards Building Semantic Role Labeler for Indian Languages

Maaz Anwar; Dipti Misra Sharma

Towards Building Semantic Role Labeler for Indian Languages

Abstract

We present a statistical system for identifying the semantic relationships or semantic roles for two major Indian Languages, Hindi and Urdu. Given an input sentence and a predicate/verb, the system first identifies the arguments pertaining to that verb and then classifies it into one of the semantic labels which can either be a DOER, THEME, LOCATIVE, CAUSE, PURPOSE etc. The system is based on 2 statistical classifiers trained on roughly 130,000 words for Urdu and 100,000 words for Hindi that were hand-annotated with semantic roles under the PropBank project for these two languages. Our system achieves an accuracy of 86% in identifying the arguments of a verb for Hindi and 75% for Urdu. At the subsequent task of classifying the constituents into their semantic roles, the Hindi system achieved 58% precision and 42% recall whereas Urdu system performed better and achieved 83% precision and 80% recall. Our study also allowed us to compare the usefulness of different linguistic features and feature combinations in the semantic role labeling task. We also examine the use of statistical syntactic parsing as feature in the role labeling task.

Anthology ID:: L16-1727
Volume:: Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16)
Month:: May
Year:: 2016
Address:: Portorož, Slovenia
Editors:: Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Sara Goggi, Marko Grobelnik, Bente Maegaard, Joseph Mariani, Helene Mazo, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:: LREC
SIG:
Publisher:: European Language Resources Association (ELRA)
Note:
Pages:: 4588–4595
Language:
URL:: https://aclanthology.org/L16-1727/
DOI:
Bibkey:
Cite (ACL):: Maaz Anwar and Dipti Sharma. 2016. Towards Building Semantic Role Labeler for Indian Languages. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16), pages 4588–4595, Portorož, Slovenia. European Language Resources Association (ELRA).
Cite (Informal):: Towards Building Semantic Role Labeler for Indian Languages (Anwar & Sharma, LREC 2016)
Copy Citation:
PDF:: https://aclanthology.org/L16-1727.pdf

PDF Cite Search Fix data