MELD: Melding Diverse Multilingual and Multi-Domain Datasets for Named Entity Recognition Evaluation

Kevin Glocker, Marco Kuhlmann


Abstract
Zero-shot Named Entity Recognition (NER) has gained prominence for information extraction across diverse domains without being limited to a single, fixed tag set. However, existing NER resources vary widely in data format, licensing terms, annotation schemes, and availability, making it difficult to systematically evaluate the generalization capabilities of zero-shot NER models. Prior attempts to aggregate datasets with broad coverage across domains have largely focused on a small subset of languages, and it is often not transparent how datasets were processed from their sources. This paper introduces MELD, a comprehensive multilingual and multi-domain data collection designed to address these gaps. MELD integrates 60 NER datasets spanning 194 languages, 14 domains, and 601 normalized entity types. While previously introduced multilingual NER datasets are mainly silver-standard, MELD contains gold-standard annotations for 60 languages. All data processing steps are fully open-source and reproducible, facilitating future extensions and ensuring long-term accessibility. While MELD is primarily designed for zero-shot evaluation, it also provides training and development splits in a single, consistent format to support future research in few-shot and supervised NER settings.
Anthology ID:
2026.lrec-1.148
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
1889–1903
Language:
External URL:
https://lrec.elra.info/lrec2026-main-148
DOI:
10.63317/32qrd24xac2e
Bibkey:
Cite (ACL):
Kevin Glocker and Marco Kuhlmann. 2026. MELD: Melding Diverse Multilingual and Multi-Domain Datasets for Named Entity Recognition Evaluation. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 1889–1903, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
MELD: Melding Diverse Multilingual and Multi-Domain Datasets for Named Entity Recognition Evaluation (Glocker & Kuhlmann, LREC 2026)
Copy Citation: