LongTailQA: Benchmarking LLMs and RAG Models on Disambiguated Long-Tail Entities

William Xion, Uwe Hadler, Tim Cofala, Maximilian Idahl, Soumyadeep Roy, Wolfgang Nejdl


Abstract
Large Language Models (LLMs) struggle with memorizing long-tail facts. Retrieval-Augmented Generation (RAG) models show better performance on long-tail Question Answering (QA) by offloading memory to external knowledge sources. We demonstrate that popular QA benchmarks such as PopQA, WITQA, and EntityQA contain significant entity ambiguity, with 8-30% of long-tail questions referencing entities with non-unique names. This ambiguity confounds evaluation, obscuring true model capabilities. To perform robust benchmarking, we disambiguate these questions with the Wikipedia knowledge graph to develop LongTailQA, an improved QA benchmark that mitigates entity ambiguity in long-tail entity questions. We evaluate various recent LLMs and RAG models, such as Self-RAG and InstructRAG, investigating retriever quality and retrieval depth impacts on QA performance. We observe that: (i) disambiguation improves model accuracy up to 24.7%, (ii) RAG models benefit significantly more than vanilla LLMs, (iii) simply increasing retrieval depth does not improve RAG performance, and (iv) RAG models achieve high accuracy with perfect information, highlighting the need to filter noisy documents during retrieval. The LongTailQA benchmark facilitates robust evaluation of long-tail knowledge recall and RAG system effectiveness. We make the codebase and datasets publicly available at https://github.com/williamx854/LongTailQA-Benchmark
Anthology ID:
2026.lrec-1.405
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
5182–5191
Language:
External URL:
https://lrec.elra.info/lrec2026-main-405
DOI:
10.63317/4tdekxzqph7x
Bibkey:
Cite (ACL):
William Xion, Uwe Hadler, Tim Cofala, Maximilian Idahl, Soumyadeep Roy, and Wolfgang Nejdl. 2026. LongTailQA: Benchmarking LLMs and RAG Models on Disambiguated Long-Tail Entities. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 5182–5191, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
LongTailQA: Benchmarking LLMs and RAG Models on Disambiguated Long-Tail Entities (Xion et al., LREC 2026)
Copy Citation: