How Far Can Bias Go? Tracing Bias from Pre-Training Data to Alignment

Marion Thaler, Abdullatif Köksal, Alina Leidinger, Anna Anna Korhonen, Hinrich Schütze


Abstract
As LLMs are increasingly integrated into user-facing applications, addressing biases that perpetuate societal inequalities is crucial. While much work has gone into measuring and mitigating biases, fewer studies have investigated their origins. Therefore, this study examines the propagation of representational gender-occupation bias from pre-training data to LLM generations. Using zero-shot prompting and token co-occurrence analyses, we explore how biases in the pre-training data influence model generations. Our findings reveal that representational biases present in the pre-training data are amplified in the model generations, regardless of hyperparameters and prompting type. By comparing gender representation in the pre-training data with real-world distributions, our research highlights discrepancies between the data and the model, underscoring the importance of further work in mitigating bias at the data level.
Anthology ID:
2026.lrec-1.315
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
3975–3995
Language:
External URL:
https://lrec.elra.info/lrec2026-main-315
DOI:
10.63317/4zeoky6waeng
Bibkey:
Cite (ACL):
Marion Thaler, Abdullatif Köksal, Alina Leidinger, Anna Anna Korhonen, and Hinrich Schütze. 2026. How Far Can Bias Go? Tracing Bias from Pre-Training Data to Alignment. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 3975–3995, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
How Far Can Bias Go? Tracing Bias from Pre-Training Data to Alignment (Thaler et al., LREC 2026)
Copy Citation: