Towards Uniform Meaning Representation for Brazilian Portuguese: Building a First UMR-Annotated Dataset

Adriana S. Pagano, Magali Duran, Federica Gamba, Daniel Zeman


Abstract
This paper presents the first Brazilian Portuguese dataset annotated for Uniform Meaning Representation (UMR). It contains 96 sentences from the Brazilian Portuguese portion of the Parallel Universal Dependencies treebank, parallel to existing UMR annotations in English, Czech, and Italian. The sentences were parsed with PortParser, revised in Arborator-Grew, converted from CoNLL-U into preliminary sentence-level UMR graphs, and manually revised in PENMAN format. The results show that CoNLL-U is a useful starting point, but semantic graph construction requires manual interpretation and language-specific lexical resources. The dataset expands multilingual UMR coverage and supports future Portuguese semantic annotation.
Anthology ID:
2026.stil-1.24
Volume:
Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology
Month:
October
Year:
2026
Address:
Cuiabá, Mato Grosso, Brazil
Editors:
Bryan Khelven da Silva Barbosa, Aline Paes, Ariani Di Felippo
Venue:
STIL
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
285–296
Language:
URL:
https://aclanthology.org/2026.stil-1.24/
DOI:
10.5753/stil.2026.26607
Bibkey:
Cite (ACL):
Adriana S. Pagano, Magali Duran, Federica Gamba, and Daniel Zeman. 2026. Towards Uniform Meaning Representation for Brazilian Portuguese: Building a First UMR-Annotated Dataset. In Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology, pages 285–296, Cuiabá, Mato Grosso, Brazil. Association for Computational Linguistics.
Cite (Informal):
Towards Uniform Meaning Representation for Brazilian Portuguese: Building a First UMR-Annotated Dataset (Pagano et al., STIL 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.stil-1.24.pdf