Prompting Instruction-tuned LLMs for Semantic Similarity Values

Xander Akiko Snelder, Yunchong Huang, Jelke Bloem


Abstract
The impressive few-shot performance of generative decoder transformer language models at novel tasks has raised interest in using them to estimate lexical-semantic properties of words, word pairs or multi-word expressions. We explore the task of eliciting semantic similarity scores between word pairs through prompting, comparing these scores to human benchmarks. We investigate different prompting approaches, different model architectures and different languages using the Dutch, English and Mandarin Chinese SimLex-999 benchmarks. The results show that prompting each word pair individually yields better correlations, and that models struggle with the distinction between similarity and relatedness, just as static and contextual word embedding models did. The new, open-weight gpt-oss-20b model yields the highest correlation with human ratings out of the models we evaluated.
Anthology ID:
2026.lrec-1.891
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
11390–11403
Language:
External URL:
https://lrec.elra.info/lrec2026-main-891
DOI:
10.63317/3kbjxx6989dg
Bibkey:
Cite (ACL):
Xander Akiko Snelder, Yunchong Huang, and Jelke Bloem. 2026. Prompting Instruction-tuned LLMs for Semantic Similarity Values. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 11390–11403, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Prompting Instruction-tuned LLMs for Semantic Similarity Values (Snelder et al., LREC 2026)
Copy Citation: