A Dataset for Semantic Role Labelling of Hindi-English Code-Mixed Tweets

Riya Pal, Dipti Sharma


Abstract
We present a data set of 1460 Hindi-English code-mixed tweets consisting of 20,949 tokens labelled with Proposition Bank labels marking their semantic roles. We created verb frames for complex predicates present in the corpus and formulated mappings from Paninian dependency labels to Proposition Bank labels. With the help of these mappings and the dependency tree, we propose a baseline rule based system for Semantic Role Labelling of Hindi-English code-mixed data. We obtain an accuracy of 96.74% for Argument Identification and are able to further classify 73.93% of the labels correctly. While there is relevant ongoing research on Semantic Role Labelling and on building tools for code-mixed social media data, this is the first attempt at labelling semantic roles in code-mixed data, to the best of our knowledge.
Anthology ID:
W19-4020
Volume:
Proceedings of the 13th Linguistic Annotation Workshop
Month:
August
Year:
2019
Address:
Florence, Italy
Editors:
Annemarie Friedrich, Deniz Zeyrek, Jet Hoek
Venue:
LAW
SIG:
SIGANN
Publisher:
Association for Computational Linguistics
Note:
Pages:
178–188
Language:
URL:
https://aclanthology.org/W19-4020
DOI:
10.18653/v1/W19-4020
Bibkey:
Cite (ACL):
Riya Pal and Dipti Sharma. 2019. A Dataset for Semantic Role Labelling of Hindi-English Code-Mixed Tweets. In Proceedings of the 13th Linguistic Annotation Workshop, pages 178–188, Florence, Italy. Association for Computational Linguistics.
Cite (Informal):
A Dataset for Semantic Role Labelling of Hindi-English Code-Mixed Tweets (Pal & Sharma, LAW 2019)
Copy Citation:
PDF:
https://aclanthology.org/W19-4020.pdf