Entity Image and Mixed-Modal Image Retrieval Datasets

Cristian-Ioan Blaga, Paul Suganthan G C, Sahil Dua, Krishna Srinivasan, Enrique Alfonseca, Peter Dornbach, Tom Duerig, Imed Zitouni, Zhe Dong


Abstract
Despite advances in multimodal learning, challenging benchmarks for mixed-modal image retrieval that combines visual and textual information are lacking. This paper introduces a novel benchmark to rigorously evaluate image retrieval that demands deep cross-modal contextual understanding. We present two new datasets: the Entity Image Dataset (EI), providing canonical images for Wikipedia entities, and the Mixed-Modal Image Retrieval Dataset (MMIR), derived from the WIT dataset. The MMIR benchmark features two challenging query types requiring models to ground textual descriptions in the context of provided visual entities: single entity-image queries (one entity image with descriptive text) and multi-entity-image queries (multiple entity images with relational text). We empirically validate the benchmark’s utility as both a training corpus and an evaluation set for mixed-modal retrieval. The quality of both datasets is further affirmed through crowd-sourced human annotations.
Anthology ID:
2026.lrec-1.734
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
9349–9357
Language:
External URL:
https://lrec.elra.info/lrec2026-main-734
DOI:
10.63317/2fnaa4f79qa5
Bibkey:
Cite (ACL):
Cristian-Ioan Blaga, Paul Suganthan G C, Sahil Dua, Krishna Srinivasan, Enrique Alfonseca, Peter Dornbach, Tom Duerig, Imed Zitouni, and Zhe Dong. 2026. Entity Image and Mixed-Modal Image Retrieval Datasets. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 9349–9357, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Entity Image and Mixed-Modal Image Retrieval Datasets (Blaga et al., LREC 2026)
Copy Citation: