Andrea Biondo
Author directory2026
Embedding Similarity Is Not Quality Estimation: Lessons from Replacing a Dedicated QE Model
Dimitrios Zaikis | Andrea Biondo | Matthew Dixon | Konstantinos Karageorgos | Aaron Schliem
Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 2)
Dimitrios Zaikis | Andrea Biondo | Matthew Dixon | Konstantinos Karageorgos | Aaron Schliem
Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 2)
Machine translation quality estimation (QE) typically relies on dedicated neural models trained on human judgments. We evaluate whether cosine similarity over general-purpose embeddings can serve as a lightweight alternative, using Gemini embeddings as the scoring backbone. Through three experiments (rogue dimension analysis, score calibration, and a learned calibration head) and a root cause analysis, we find that cosine similarity between source and translation saturates in the 0.94–0.99 range because even poor translations preserve most of the source semantics, leaving an Area Under the ROC Curve (AUC) ceiling of approximately 0.63. However, a LightGBM classifier trained on normalized cosine and surface-level text features breaks through this ceiling (AUC 0.751), with the improvement driven primarily by features orthogonal to embedding similarity.