The Lock, Stock, and Barrel of Marathi Multiwords

Aakanksha Padhye, Ashwini Vaidya


Abstract
Multiword expressions are an important area of study in linguistics and natural language processing as they represent combination of words that function as a single unit, and display properties that cannot be predicated fully from their individual components. This paper describes annotated corpora of about 3000 multiword expressions across syntactic categories in Marathi. This is the first exhaustive resource for Marathi which includes both verbal and non-verbal multiwords. In order to develop the guidelines for annotation, we have used the existing literature on the identification and classification of these expressions. Following the PARSEME 2.0 guidelines, we discuss the categories of multiwords and their behaviour in the corpus. Throughout the annotation process, we encounter variability in compositionality and syntactic realization and discuss our design decisions during annotation. Such a dataset will further our understanding of how grammatical structure can be integrated with lexically stored multiword units in Marathi.
Anthology ID:
2026.mwe-1.11
Volume:
Proceedings of the 22nd Workshop on Multiword Expressions (MWE 2026)
Month:
March
Year:
2026
Address:
Rabat, Marocco
Editors:
Atul Kr. Ojha, Verginica Barbu Mititelu, Mathieu Constant, Ivelina Stoyanova, A. Seza Doğruöz, Alexandre Rademaker
Venues:
MWE | WS
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
96–102
Language:
URL:
https://aclanthology.org/2026.mwe-1.11/
DOI:
Bibkey:
Cite (ACL):
Aakanksha Padhye and Ashwini Vaidya. 2026. The Lock, Stock, and Barrel of Marathi Multiwords. In Proceedings of the 22nd Workshop on Multiword Expressions (MWE 2026), pages 96–102, Rabat, Marocco. Association for Computational Linguistics.
Cite (Informal):
The Lock, Stock, and Barrel of Marathi Multiwords (Padhye & Vaidya, MWE 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.mwe-1.11.pdf