JMTEB and JMTEB-lite: Japanese Massive Text Embedding Benchmark and Its Lightweight Version

Shengzhe Li, Masaya Ohagi, Ryokan Ri, Akihiko Fukuchi, Tomohide Shibata, Daisuke Kawahara


Abstract
We present JMTEB, a large-scale evaluation suite for Japanese text embedding models, designed to provide comprehensive coverage across multiple task types. The benchmark integrates 28 datasets across 5 tasks, enabling broad and challenging evaluation of model performance in diverse scenarios. While the full benchmark delivers thorough assessment, its scale poses practical challenges in terms of computation time and resource requirements. To address this, we construct JMTEB-lite, a lightweight version of JMTEB, by substantially reducing corpus size in retrieval-related tasks. JMTEB-lite significantly accelerates evaluation while maintaining high fidelity to the full benchmark. Together, JMTEB and JMTEB-lite form a flexible evaluation framework: the full version serves as a comprehensive standard for exhaustive benchmarking, while the lightweight version enables rapid iteration and efficient model selection. This dual approach facilitates both rigorous evaluation and practical development workflows, supporting the advancement of Japanese text embedding research.
Anthology ID:
2026.lrec-1.588
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
7423–7434
Language:
External URL:
https://lrec.elra.info/lrec2026-main-588
DOI:
10.63317/5ouzpv2f2f6k
Bibkey:
Cite (ACL):
Shengzhe Li, Masaya Ohagi, Ryokan Ri, Akihiko Fukuchi, Tomohide Shibata, and Daisuke Kawahara. 2026. JMTEB and JMTEB-lite: Japanese Massive Text Embedding Benchmark and Its Lightweight Version. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 7423–7434, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
JMTEB and JMTEB-lite: Japanese Massive Text Embedding Benchmark and Its Lightweight Version (Li et al., LREC 2026)
Copy Citation: