HEAD-QA v2: Expanding a Healthcare Benchmark for Reasoning

Alexis Correa, Carlos Gómez-Rodríguez, David Vilares


Abstract
We introduce HEAD-QA v2, an expanded and updated version of a Spanish/English healthcare multiple-choice reasoning dataset originally released by Vilares and Gómez-Rodríguez (2019). The update responds to the growing need for high-quality datasets that capture the linguistic and conceptual complexity of healthcare reasoning. We extend the dataset to over 12,000 questions from ten years of Spanish professional exams, benchmark several open-source LLMs using prompting, RAG, and probability-based answer selection, and provide additional multilingual versions to support future work. Results indicate that performance is mainly driven by model scale and intrinsic reasoning ability, with complex inference strategies obtaining limited gains. Together, these results establish HEAD-QA v2 as a reliable resource for advancing research on biomedical reasoning and model improvement.
Anthology ID:
2026.lrec-1.407
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
5203–5214
Language:
External URL:
https://lrec.elra.info/lrec2026-main-407
DOI:
10.63317/2dvxxrgarr9d
Bibkey:
Cite (ACL):
Alexis Correa, Carlos Gómez-Rodríguez, and David Vilares. 2026. HEAD-QA v2: Expanding a Healthcare Benchmark for Reasoning. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 5203–5214, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
HEAD-QA v2: Expanding a Healthcare Benchmark for Reasoning (Correa et al., LREC 2026)
Copy Citation: