Breaking the Benchmark: Revealing LLM Bias via Minimal Contextual Augmentation

Kaveh Eskandari Miandoab, Mahammed Kamruzzaman, Arshia Gharooni, Gene Louis Kim, Vasanth Sarathy, Ninareh Mehrabi


Abstract
Large Language Models have been shown to demonstrate stereotypical biases in their representations and behavior due to the discriminative nature of the data that they have been trained on. Despite significant progress in the development of methods and models that refrain from using stereotypical information in their decision-making, recent work has shown that approaches used for bias alignment are brittle. In this work, we introduce a novel and general augmentation framework that involves three plug-and-play steps and is applicable to a number of fairness evaluation benchmarks. Through application of augmentation to a fairness evaluation dataset (Bias Benchmark for Question Answering (BBQ)), we find that Large Language Models (LLMs), including state-of-the-art open and closed weight models, are susceptible to perturbations to their inputs, showcasing a higher likelihood to behave stereotypically. Furthermore, we find that such models are more likely to have biased behavior in cases where the target demographic belongs to a community less studied by the literature, underlining the need to expand the fairness and safety research to include more diverse communities.
Anthology ID:
2026.lrec-1.322
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
4070–4092
Language:
External URL:
https://lrec.elra.info/lrec2026-main-322
DOI:
10.63317/5a6nbh2tnoeb
Bibkey:
Cite (ACL):
Kaveh Eskandari Miandoab, Mahammed Kamruzzaman, Arshia Gharooni, Gene Louis Kim, Vasanth Sarathy, and Ninareh Mehrabi. 2026. Breaking the Benchmark: Revealing LLM Bias via Minimal Contextual Augmentation. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 4070–4092, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Breaking the Benchmark: Revealing LLM Bias via Minimal Contextual Augmentation (Eskandari Miandoab et al., LREC 2026)
Copy Citation:
Optionalsupplementarymaterial:
 2026.lrec-1.322.OptionalSupplementaryMaterial.zip