Charting the European LLM Benchmarking Landscape: A New Taxonomy and Registry

Spela Vintar, Mojca Brglez, Taja Kuzman Pungeršek, Nikola Ljubešić


Abstract
While new benchmarks for large language models (LLMs) are being developed continuously to catch up with the growing capabilities of new models and AI in general, using and evaluating LLMs in non-English languages remains a poorly-charted landscape. We give a concise overview of recent developments in LLM benchmarking, and then propose a new taxonomy for the categorization of benchmarks that is tailored to multilingual or non-English use scenarios. We further propose a registry of benchmarks implementing the new categorization and documenting benchmarks with a rich set of metadescriptors. While still at a pilot stage, such a registry can lead to a more coordinated development of benchmarks for European languages. We conclude with a review of current trends and advocate for a higher language and culture sensitivity of evaluation methods.
Anthology ID:
2026.llms4ssh-1.22
Volume:
Proceedings of Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities (LLMs4SSH) @ LREC 2026
Month:
May
Year:
2026
Address:
Palma de Mallorca (Spain)
Editors:
Arturo Montejo-Raez, Cristina Grisot, Joanna Blochowiak, Nikola Ljubešić, Elena Battaner, German Rigau
Venues:
LLMs4SSH | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
205–217
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-llms4ssh-22
DOI:
10.63317/4kixe3c9zmde
Bibkey:
Cite (ACL):
Spela Vintar, Mojca Brglez, Taja Kuzman Pungeršek, and Nikola Ljubešić. 2026. Charting the European LLM Benchmarking Landscape: A New Taxonomy and Registry. In Proceedings of Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities (LLMs4SSH) @ LREC 2026, pages 205–217, Palma de Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
Charting the European LLM Benchmarking Landscape: A New Taxonomy and Registry (Vintar et al., LLMs4SSH 2026)
Copy Citation: