Evaluating Portuguese Tokenizers as Morpheme Sequence in Relation to LLM Downstream Performance

Guilherme L. Mello, Marcelo Finger


Abstract
Sub-word tokenizers allow us to handle open vocabulary problems using a relatively small set of tokens. Despite its ease of use, it usually relies on a data-driven approach that does not directly employ linguistic or morphologic features for text tokenization. [Bostrom and Durrett 2020] and [Hofmann et al. 2021] demonstrate that morphemes improve LLM performance on English texts. In this work, we explore the hypothesis that tokenizers that are capable of producing a token sequence better aligned with a morpheme sequence can improve LLMs performance on Brazilian Portuguese. To evaluate how the presence of morphemes can impact LLMs for Brazilian Portuguese, we propose MorphEval-PT, a new evaluation procedure based on the psycholinguistic concept of morphological models of word processing. We build new BPE and Unigram vocabularies that are evaluated on MorphEval-PT and, in order to validate the impact of morphemes on LLMs performance, train new LLMs from scratch and evaluate its performance on downstream tasks. Consistently, BPE demonstrates a higher precision score than Unigram in its ability to represent morphemes, as well as better performance on every downstream task. These promising results indicate that accessing the ability of tokenizers to represent morphemes is an important feature in the development of LLMs for Brazilian Portuguese and that MorphEval-PT is a good and lightweight method to improve LLM performance before any pre-training.
Anthology ID:
2026.stil-1.21
Volume:
Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology
Month:
October
Year:
2026
Address:
Cuiabá, Mato Grosso, Brazil
Editors:
Bryan Khelven da Silva Barbosa, Aline Paes, Ariani Di Felippo
Venue:
STIL
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
245–257
Language:
URL:
https://aclanthology.org/2026.stil-1.21/
DOI:
10.5753/stil.2026.26562
Bibkey:
Cite (ACL):
Guilherme L. Mello and Marcelo Finger. 2026. Evaluating Portuguese Tokenizers as Morpheme Sequence in Relation to LLM Downstream Performance. In Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology, pages 245–257, Cuiabá, Mato Grosso, Brazil. Association for Computational Linguistics.
Cite (Informal):
Evaluating Portuguese Tokenizers as Morpheme Sequence in Relation to LLM Downstream Performance (Mello & Finger, STIL 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.stil-1.21.pdf