Investigating Language and Retrieval Bias in Multilingual Previously Fact-Checked Claim Detection

Ivan Vykopal; Antonia Karamolegkou; Jaroslav Kopčan; Qiwei Peng; Tomáš Javůrek; Michal Gregor; Marian Simko

Investigating Language and Retrieval Bias in Multilingual Previously Fact-Checked Claim Detection

Ivan Vykopal, Antonia Karamolegkou, Jaroslav Kopčan, Qiwei Peng, Tomáš Javůrek, Michal Gregor, Marian Simko

Abstract

Multilingual Large Language Models (LLMs) offer powerful capabilities for cross-lingual fact-checking. However, these models often exhibit language bias, performing disproportionately better on high-resource languages such as English than on low-resource counterparts. We also present and inspect a novel concept - retrieval bias, when information retrieval systems tend to favor certain information over others, leaving the retrieval process skewed. In this paper, we study language and retrieval bias in the context of Previously Fact-Checked Claim Detection (PFCD). We evaluate six open-source multilingual LLMs across 20 languages using a fully multilingual prompting strategy, leveraging the AMC-16K dataset. By translating task prompts into each language, we uncover disparities in monolingual and cross-lingual performance and identify key trends based on model family, size, and prompting strategy. Our findings highlight persistent bias in LLM behavior and offer recommendations for improving equity in multilingual fact-checking. To investigate retrieval bias, we employed multilingual embedding models and look into the frequency of retrieved claims. Our analysis reveals that certain claims are retrieved disproportionately across different posts, leading to inflated retrieval performance for popular claims while under-representing less common ones.

Anthology ID:: 2026.eacl-long.240
Volume:: Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers)
Month:: March
Year:: 2026
Address:: Rabat, Morocco
Editors:: Vera Demberg, Kentaro Inui, Lluís Marquez
Venue:: EACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 5195–5221
Language:
URL:: https://aclanthology.org/2026.eacl-long.240/
DOI:
Bibkey:
Cite (ACL):: Ivan Vykopal, Antonia Karamolegkou, Jaroslav Kopčan, Qiwei Peng, Tomáš Javůrek, Michal Gregor, and Marian Simko. 2026. Investigating Language and Retrieval Bias in Multilingual Previously Fact-Checked Claim Detection. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5195–5221, Rabat, Morocco. Association for Computational Linguistics.
Cite (Informal):: Investigating Language and Retrieval Bias in Multilingual Previously Fact-Checked Claim Detection (Vykopal et al., EACL 2026)
Copy Citation:
PDF:: https://aclanthology.org/2026.eacl-long.240.pdf
Checklist:: 2026.eacl-long.240.checklist.pdf

PDF Cite Search Checklist Fix data