A Wikidata-Based Framework to Measure Cross-Lingual Bias in Multilingual Large Language Models

Mouloud Iferroudjene, Lisa Poggel, Andrea Schimmenti, Duo Yang, Kanchan Shivashankar, Jan-Christoph Kalo, Marta Boscariol


Abstract
Multilingual large language models (LLMs) are increasingly used for factual question answering, yet their accuracy varies across languages in ways that are difficult to interpret. A central challenge is that many multilingual probing benchmarks conflate multiple factors: the language used to ask the question, the cultural-linguistic context of the entities being queried, and the popularity skew of entities. In our paper, we disentangle these factors by asking: (i) how strongly does the Language of the Question (LoQ) affect factual recall, (ii) does matching LoQ to an entity-associated Language of the Entity (LoE) improve performance, and (iii) do these effects persist when entity popularity is controlled. To this end, we introduce WILA-PopQA, a new Wikidata-grounded benchmark spanning 9 languages with matched popularity profiles, and probe 12 open-weight models of varying sizes and architectures under aligned and misaligned LoQ–LoE conditions. We evaluate models’ answers to 4 types of questions about entity biographical properties in all selected languages. Results show that LoQ is the dominant source of variation. LoQ–LoE alignment does not consistently yield the highest accuracy, and performance depends on the property being asked. These results suggest that prompt language is an actionable experimental factor for multilingual factual evaluation.
Anthology ID:
2026.kallm-1.19
Volume:
Proceedings of the Knowledge Graphs and Large Language Models Workshop (KG-LLM) @ LREC26
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Gilles Sérasset, Katerina Gkirtzou, Michael Cochez, Jan-Christoph Kalo
Venues:
KaLLM | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
190–209
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-kgllm-19
DOI:
10.63317/29b354pejyrt
Bibkey:
Cite (ACL):
Mouloud Iferroudjene, Lisa Poggel, Andrea Schimmenti, Duo Yang, Kanchan Shivashankar, Jan-Christoph Kalo, and Marta Boscariol. 2026. A Wikidata-Based Framework to Measure Cross-Lingual Bias in Multilingual Large Language Models. In Proceedings of the Knowledge Graphs and Large Language Models Workshop (KG-LLM) @ LREC26, pages 190–209, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
A Wikidata-Based Framework to Measure Cross-Lingual Bias in Multilingual Large Language Models (Iferroudjene et al., KaLLM 2026)
Copy Citation: