Gustavo Hochgraf
Author directory2026
Beyond Aggregate Scores: A Diagnostic Analysis of Portuguese LLM Evaluation Suites
Gustavo Hochgraf | Gabriel Assis | Aline Paes | Fabio G. Cozman
Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology
Gustavo Hochgraf | Gabriel Assis | Aline Paes | Fabio G. Cozman
Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology
Benchmarks play a central role in evaluating large language models (LLMs), providing standardized comparisons across models, tasks, and adaptation strategies. However, aggregate benchmark scores often compress performance into a single number, obscuring important variation across tasks. This issue is especially relevant for less-resourced languages like Portuguese, where benchmarks may combine native tasks with translated datasets spanning heterogeneous categories and subareas. In this work, we examine this issue using PoETa v2, a broad Portuguese evaluation benchmark, as a diagnostic setting to evaluate three Qwen3 1.7B variants under different Portuguese adaptation settings. Our results show that, although overall scores differ by less than one percentage point across models, disaggregated analyses reveal substantially different patterns across task origin, category, and subarea. In particular, gains on native Portuguese tasks may coexist with losses on translated tasks, while different adaptation corpora redistribute performance in distinct ways. Taken together, our findings highlight the importance of complementing aggregate scores with multi-level analyses when evaluating Portuguese LLMs, particularly in the context of language-specific adaptation.