Leonidas Mylonadis


2026

Human judgements of word similarity have been a core benchmark for the intrinsic evaluation of word embedding models, and continue to be used for assessing the capabilities of large language models. While word similarity benchmarks have been collected for a range of languages, none existed for Greek. We develop a Modern Greek variant of the SimLex-999 word similarity dataset by gathering similarity judgements from 90 native speakers of Greek. We then use this as a benchmark for intrinsically evaluating several Greek language models.