Aakanksha Padhye
2026
The Lock, Stock, and Barrel of Marathi Multiwords
Aakanksha Padhye | Ashwini Vaidya
Proceedings of the 22nd Workshop on Multiword Expressions (MWE 2026)
Aakanksha Padhye | Ashwini Vaidya
Proceedings of the 22nd Workshop on Multiword Expressions (MWE 2026)
Multiword expressions are an important area of study in linguistics and natural language processing as they represent combination of words that function as a single unit, and display properties that cannot be predicated fully from their individual components. This paper describes annotated corpora of about 3000 multiword expressions across syntactic categories in Marathi. This is the first exhaustive resource for Marathi which includes both verbal and non-verbal multiwords. In order to develop the guidelines for annotation, we have used the existing literature on the identification and classification of these expressions. Following the PARSEME 2.0 guidelines, we discuss the categories of multiwords and their behaviour in the corpus. Throughout the annotation process, we encounter variability in compositionality and syntactic realization and discuss our design decisions during annotation. Such a dataset will further our understanding of how grammatical structure can be integrated with lexically stored multiword units in Marathi.