MesoTree: Annotated Linguistic Resources for Quantitative Comparative Linguistic Analysis and NLP in Mesoamerica

Robert Pugh, Francis Tyers, Robert Henderson


Abstract
One aspect of descriptive and documentary linguistic materials that is becoming increasingly important in the information age is that they be searchable, quantifiable, and comparable. In this paper, we describe an effort to create morphosyntactically-annotated corpora for a number of under-served Mesoamerican languages using Universal Dependencies. We describe the Mesoamerican linguistic area and languages involved in the project, the training and annotation process, and give a status report on the current state of the corpora. Finally, we describe a comparitive syntax experiment and train UD parsing models on the data, demonstrating the usefulness of UD for facilitating quantitative, comparative linguistic research.
Anthology ID:
2026.udw-1.17
Volume:
Proceedings of the Ninth Workshop on Universal Dependencies (UDW 2026)
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Çağrı Çöltekin, Kaja Dobrovoljc
Venues:
UDW | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
197–207
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-udw-17
DOI:
10.63317/2xvtti733shi
Bibkey:
Cite (ACL):
Robert Pugh, Francis Tyers, and Robert Henderson. 2026. MesoTree: Annotated Linguistic Resources for Quantitative Comparative Linguistic Analysis and NLP in Mesoamerica. In Proceedings of the Ninth Workshop on Universal Dependencies (UDW 2026), pages 197–207, Palma de Mallorca, Spain. ELRA Language Resources Association (ELRA).
Cite (Informal):
MesoTree: Annotated Linguistic Resources for Quantitative Comparative Linguistic Analysis and NLP in Mesoamerica (Pugh et al., UDW 2026)
Copy Citation: