Dominick Maia Alexandre
Author directory2026
Extending Enhanced Universal Dependencies to Nheengatu: Semi-Automated Control and Raising Annotation
Dominick Maia Alexandre | Leonel Figueiredo de Alencar
Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology
Dominick Maia Alexandre | Leonel Figueiredo de Alencar
Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology
This paper presents the first Enhanced Universal Dependencies annotation effort for a Brazilian Indigenous language, focusing on XCOMP constructions in Nheengatu, an endangered Tupian language. We describe a semiautomatic pipeline for the UD Nheengatu-CompLin treebank to annotate control and raising predicates and propagated subjects. The method combines theoretical discussions on XCOMP constructions with corpus-based analysis of Nheengatu syntax to create rule-based enhancement heuristics. The work also provides annotation guidelines and expands EUD resources for Indigenous and low-resource languages within the Universal Dependencies framework.
Parsing Nheengatu: Performance Gains for a Brazilian Indigenous Universal Dependencies Treebank
Dominick Maia Alexandre | Leonel Figueiredo de Alencar
Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 2
Dominick Maia Alexandre | Leonel Figueiredo de Alencar
Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 2
This paper evaluates the impact of expanding the UD_Nheengatu-CompLin treebank on parsing performance for Nheengatu, a Brazilian endangered Indigenous language. We hypothesized that the inclusion of annotated data would result in a 10% improvement in the Labeled Attachment Score (LAS). To test this hypothesis, we conducted a 10-fold cross-validation experiment using UDPipe 1.4 under two conditions: parsing with gold tokenization and gold tags, and automatic parsing from raw text. Statistical significance was determined using the Mann-Whitney U test. Although the expected gain was not achieved, the results show improvements in parsing accuracy and reduced variance across folds. The findings highlight the importance of corpus expansion and standardized annotation workflows for improving parsing performance in low-resource language scenarios and for supporting reproducible evaluation methods in the computational modeling of minority languages.