POS Tagging in Low-Resource Maithili Language: Specific Challenges and Nuances

Shivani Priya, Shruti Jha, Urmila Jha, Girish Nath Jha, Deepali Tiwari, Jyoti Raj


Abstract
Abstract Part-of-Speech (POS) tagging is a key step in Natural Language Processing (NLP), laying the groundwork for more advanced syntactic and semantic tasks. Despite Maithili’s status as an Indo-Aryan language with a rich literary tradition and official recognition in India, computational resources for it are still very limited. In this paper, the creation of an annotated corpus of 25,000 sentences drawn from the fields of health, tourism, and administration is described with the hierarchical tagset currently used for Maithili. This paper also indicates that standard tagsets, typically adapted from English or Hindi, fail to capture the linguistic nuances of Maithili. This underestimates the need for a dedicated tagging framework that considers characteristics like vocative particles, verbal nuances, honorific complexities. Keywords: Parts of Speech, Natural Language Processing, Maithili, annotation
Anthology ID:
2026.wildre-1.9
Volume:
Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation
Month:
May
Year:
2026
Address:
Palma, Mallorca, Spain
Editors:
Girish Nath Jha, Kalika Bali, Sobha L, Devendr Kumar
Venues:
WILDRE | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
67–74
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-wildre-09
DOI:
10.63317/27ugx7nj4vvs
Bibkey:
Cite (ACL):
Shivani Priya, Shruti Jha, Urmila Jha, Girish Nath Jha, Deepali Tiwari, and Jyoti Raj. 2026. POS Tagging in Low-Resource Maithili Language: Specific Challenges and Nuances. In Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation, pages 67–74, Palma, Mallorca, Spain. ELRA Language Resources Association (ELRA).
Cite (Informal):
POS Tagging in Low-Resource Maithili Language: Specific Challenges and Nuances (Priya et al., WILDRE 2026)
Copy Citation: