Benchmark Data Contamination in Underrepresented Languages: A Comprehensive Analysis Using Brazilian Data

Iriedson Souto Maior de Moraes Vilar, David Candeia Maia, João Brunet, Fabio Morais, Leandro Balby Marinho


Abstract
Large Language Models (LLMs) are typically evaluated using standardized benchmarks to enable consistent performance measurement and model comparison. However, the reliability of these benchmarks can be undermined by data contamination, which occurs when evaluation items are inadvertently included in training corpora. While this issue has been investigated primarily in high-resource languages such as English and Chinese, its impact on underrepresented languages — such as Brazilian Portuguese — remains understudied. In this paper, we present one of the first systematic investigations of benchmark data contamination (BDC) in an underrepresented language setting, using Brazilian Portuguese as a case study. Using validated methodologies from the literature, we evaluate specialized and multilingual models across four benchmarks: BLUEX, ENEM Challenge, OAB Exams, and HealthQA-BR. Our approach applyes TS-Guessing to detect contamination via memorized knowledge, alongside a 50-character n-gram similarity strategy to identify benchmark items leaked into training data. Our results provide consistent evidence of contamination, revealing that models with stronger memorization and retrieval abilities tend to achieve artificially inflated benchmark scores. Our contributions include: (i) classifying models according to their contamination risk, (ii) identifying the benchmarks most affected by data leakage, and (iii) reporting contaminated training corpora.
Anthology ID:
2026.lrec-1.374
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
4765–4777
Language:
External URL:
https://lrec.elra.info/lrec2026-main-374
DOI:
10.63317/39wbjvajnh7t
Bibkey:
Cite (ACL):
Iriedson Souto Maior de Moraes Vilar, David Candeia Maia, João Brunet, Fabio Morais, and Leandro Balby Marinho. 2026. Benchmark Data Contamination in Underrepresented Languages: A Comprehensive Analysis Using Brazilian Data. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 4765–4777, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Benchmark Data Contamination in Underrepresented Languages: A Comprehensive Analysis Using Brazilian Data (Vilar et al., LREC 2026)
Copy Citation: