Lilian Teixeira de Sousa

Author directory

2026

While Part-of-speech tagging is considered a well-understood task in the literature, most work has focused on indo-european languages with fusional morphology. For low-resource agglutinative languages, such as brazillian indigenous languages, the task is commonly constrained by small annotated corpora, high lexical sparsity, and morphological patterns that are poorly represented by word-level models. This paper investigates the impact of unsupervised subword segmentation techniques on POS tagging for brazillian indigenous languages. We compare a word-level baseline with Byte Pair Encoding, Morfessor, and FlatCat on Bororo, Nheengatu, and Tupinamba corpora, including a controlled experiment on the size of the training corpus. Our findings suggest that unsupervised segmentation can reduce sparsity in low-resource POS tagging, although its benefit depends on the language, corpus size, and segmentation method.