A Comparison of Commonly Used Automatic Evaluation Metrics for Open-Ended Tasks in Portuguese

Eduardo D. Faé, Dennis Giovani Balreira, Viviane Pereira Moreira


Abstract
Open-ended tasks in Natural Language processing are characterized by the absence of a fixed set of outputs or short, predetermined responses. Typical examples include open-domain question answering and summarization. The automatic evaluation of these tasks is challenging, and existing metrics suffer from important limitations. Thus, many works have already been proposed to evaluate the performance of such metrics. However, the great majority of these works were conducted in English, leaving a gap for analyses in other languages, such as Portuguese. This work investigates the performance of some of the most popular automated intrinsic evaluation metrics for open-ended tasks by analyzing their correlation with human judgment. Using both the Summarization and Question Answering tasks, this study compares traditional n-gram-based metrics with metrics based on contextual embeddings generated by Pre-trained Language Models (PLMs), performing all analyses using models and datasets available for Portuguese. Results indicate that while PLM-based metrics generally show slightly higher correlation with human perception, their performance is highly reliant on the quality of the models used. Still, none of the evaluated metrics displayed a strong correlation with human judgments, suggesting the need for more robust metrics.
Anthology ID:
2026.stil-1.11
Volume:
Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology
Month:
October
Year:
2026
Address:
Cuiabá, Mato Grosso, Brazil
Editors:
Bryan Khelven da Silva Barbosa, Aline Paes, Ariani Di Felippo
Venue:
STIL
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
126–138
Language:
URL:
https://aclanthology.org/2026.stil-1.11/
DOI:
10.5753/stil.2026.26578
Bibkey:
Cite (ACL):
Eduardo D. Faé, Dennis Giovani Balreira, and Viviane Pereira Moreira. 2026. A Comparison of Commonly Used Automatic Evaluation Metrics for Open-Ended Tasks in Portuguese. In Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology, pages 126–138, Cuiabá, Mato Grosso, Brazil. Association for Computational Linguistics.
Cite (Informal):
A Comparison of Commonly Used Automatic Evaluation Metrics for Open-Ended Tasks in Portuguese (Faé et al., STIL 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.stil-1.11.pdf