Are Social Biases in LLMs Consistent across Generative Tasks? A Case Study for Basque

Muitze Zulaika, Xabier Saralegi, Julia Shershneva, Lia Gonzalez, Arkaitz Fullaondo


Abstract
Most bias benchmarks for Large Language Models (LLMs) rely on multiple-choice formats, overlooking subtler biases that emerge in open-ended text generation. This gap is particularly relevant for low-resource languages like Basque, where culturally grounded evaluation resources are limited. We introduce BasqBBG (Basque Bias Benchmark for Generation), the first systematic benchmark for social bias in Basque Natural Language Generation (NLG), covering eight bias categories—including a newly added feminism dimension—adapted from the BasqBBQ dataset. We validate an LLM-as-a-Judge framework against expert human evaluations on two NLG tasks (story continuation and generative QA), achieving strong agreement (agreement of 0.78 in bias presence and 0.92 in bias directionality). We scale this approach to ten additional tasks and five models. Results show that bias levels vary markedly across tasks and depend more on model family than size: Llama-based models exhibit higher and less consistent bias (45–50%), whereas GPT-4o and the Gemma-based Kimu-9B remain substantially fairer (≤20%). Our findings highlight the need for task-aware, language-specific frameworks to assess social bias in generative LLMs. Keywords: Large Language Models, Social Bias, Basque, Natural Language Generation, Benchmarking, Manual Evaluation, LLM-as-a-judge.
Anthology ID:
2026.lrec-1.664
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
8415–8430
Language:
External URL:
https://lrec.elra.info/lrec2026-main-664
DOI:
10.63317/52zk8uyjrw5k
Bibkey:
Cite (ACL):
Muitze Zulaika, Xabier Saralegi, Julia Shershneva, Lia Gonzalez, and Arkaitz Fullaondo. 2026. Are Social Biases in LLMs Consistent across Generative Tasks? A Case Study for Basque. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 8415–8430, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Are Social Biases in LLMs Consistent across Generative Tasks? A Case Study for Basque (Zulaika et al., LREC 2026)
Copy Citation: