Beyond Aggregate Scores: A Diagnostic Analysis of Portuguese LLM Evaluation Suites

Gustavo Hochgraf, Gabriel Assis, Aline Paes, Fabio G. Cozman


Abstract
Benchmarks play a central role in evaluating large language models (LLMs), providing standardized comparisons across models, tasks, and adaptation strategies. However, aggregate benchmark scores often compress performance into a single number, obscuring important variation across tasks. This issue is especially relevant for less-resourced languages like Portuguese, where benchmarks may combine native tasks with translated datasets spanning heterogeneous categories and subareas. In this work, we examine this issue using PoETa v2, a broad Portuguese evaluation benchmark, as a diagnostic setting to evaluate three Qwen3 1.7B variants under different Portuguese adaptation settings. Our results show that, although overall scores differ by less than one percentage point across models, disaggregated analyses reveal substantially different patterns across task origin, category, and subarea. In particular, gains on native Portuguese tasks may coexist with losses on translated tasks, while different adaptation corpora redistribute performance in distinct ways. Taken together, our findings highlight the importance of complementing aggregate scores with multi-level analyses when evaluating Portuguese LLMs, particularly in the context of language-specific adaptation.
Anthology ID:
2026.stil-1.18
Volume:
Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology
Month:
October
Year:
2026
Address:
Cuiabá, Mato Grosso, Brazil
Editors:
Bryan Khelven da Silva Barbosa, Aline Paes, Ariani Di Felippo
Venue:
STIL
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
209–217
Language:
URL:
https://aclanthology.org/2026.stil-1.18/
DOI:
10.5753/stil.2026.26539
Bibkey:
Cite (ACL):
Gustavo Hochgraf, Gabriel Assis, Aline Paes, and Fabio G. Cozman. 2026. Beyond Aggregate Scores: A Diagnostic Analysis of Portuguese LLM Evaluation Suites. In Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology, pages 209–217, Cuiabá, Mato Grosso, Brazil. Association for Computational Linguistics.
Cite (Informal):
Beyond Aggregate Scores: A Diagnostic Analysis of Portuguese LLM Evaluation Suites (Hochgraf et al., STIL 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.stil-1.18.pdf