Embedding Similarity Is Not Quality Estimation: Lessons from Replacing a Dedicated QE Model

Dimitrios Zaikis, Andrea Biondo, Matthew Dixon, Konstantinos Karageorgos, Aaron Schliem


Abstract
Machine translation quality estimation (QE) typically relies on dedicated neural models trained on human judgments. We evaluate whether cosine similarity over general-purpose embeddings can serve as a lightweight alternative, using Gemini embeddings as the scoring backbone. Through three experiments (rogue dimension analysis, score calibration, and a learned calibration head) and a root cause analysis, we find that cosine similarity between source and translation saturates in the 0.94–0.99 range because even poor translations preserve most of the source semantics, leaving an Area Under the ROC Curve (AUC) ceiling of approximately 0.63. However, a LightGBM classifier trained on normalized cosine and surface-level text features breaks through this ceiling (AUC 0.751), with the improvement driven primarily by features orthogonal to embedding similarity.
Anthology ID:
2026.eamt-2.26
Volume:
Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 2)
Month:
June
Year:
2026
Address:
Tilburg, The Netherlands
Editors:
Dimitar Shterionov, Eva Vanmassenhove, Mirella De Sisto, Fred Blain, Javad Pourmostafa Roshan Sharami, Lisa Lepp, Chiara Manna, Argentina Anna Rescigno, Alina Karakanta, Ayla Rigouts Terryn, Manuel Lardelli, Natalia Resende, Elena Murgolo, Janiça Hackenbuchner, Anna Zaretskaya, Miquel Esplà-Gomis, Thierry Etchegoyhen, Dagmar Gromann, Rachel Bawden, Barry Haddow, Sara Szoc, Mikel Forcada, Helena Moniz
Venue:
EAMT
SIG:
Publisher:
European Association for Machine Translation
Note:
Pages:
66–71
Language:
URL:
https://aclanthology.org/2026.eamt-2.26/
DOI:
Bibkey:
Cite (ACL):
Dimitrios Zaikis, Andrea Biondo, Matthew Dixon, Konstantinos Karageorgos, and Aaron Schliem. 2026. Embedding Similarity Is Not Quality Estimation: Lessons from Replacing a Dedicated QE Model. In Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 2), pages 66–71, Tilburg, The Netherlands. European Association for Machine Translation.
Cite (Informal):
Embedding Similarity Is Not Quality Estimation: Lessons from Replacing a Dedicated QE Model (Zaikis et al., EAMT 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.eamt-2.26.pdf