Lost in Translation: Repurposing semantic similarity benchmarks for evaluating lexical-semantic consistency in LLM-based machine translation

Quin Ye, Jelke Bloem


Abstract
We propose and demonstrate a repurposing of the lexical similarity benchmark Multi-SimLex and the SimLex-999 family of resources for assessing the cross-lingual lexical-semantic consistency of multilingual large language models. While originally gathered for evaluating word embedding models, the parallel nature of the word pairs enables their use in machine translation settings. Using a manually verified subset of 500 word pairs from the Multi-SimLex dataset, we evaluate models’ ability to assess semantic similarity and perform translation between English and Mandarin through zero-shot prompting. We compare BLOOMZ and GPT-4’s similarity ratings against human-annotated benchmarks and examine translation consistency using our and other metrics, with GPT-4 showing stronger human alignment. As SimLex-999 and Multi-SimLex together cover a range of at least 25 languages, this approach has the potential to be extended to many language pairs including ones that don’t involve English, though it requires some manual checks.
Anthology ID:
2026.resourceful-4.1
Volume:
Proceedings of the Fourth Workshop on the Role of Resources in the Age of Large Language Models (RESOURCEFUL 2026)
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Felix Morger, Nikolai Ilinykh, Barbara Scalvini, Simon Dobnik, Dana Dannélls
Venues:
RESOURCEFUL | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
1–12
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-resourceful-01
DOI:
10.63317/3n3847mvjzk2
Bibkey:
Cite (ACL):
Quin Ye and Jelke Bloem. 2026. Lost in Translation: Repurposing semantic similarity benchmarks for evaluating lexical-semantic consistency in LLM-based machine translation. In Proceedings of the Fourth Workshop on the Role of Resources in the Age of Large Language Models (RESOURCEFUL 2026), pages 1–12, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
Lost in Translation: Repurposing semantic similarity benchmarks for evaluating lexical-semantic consistency in LLM-based machine translation (Ye & Bloem, RESOURCEFUL 2026)
Copy Citation: