Semantic Fidelity Versus Literary Quality: A Construct Validity Study of Neural Machine Translation Metrics

Dmytro Chaplynskyi; Ivan Kulynych; Maria Shvedova; Lesia Ivashkevych

Semantic Fidelity Versus Literary Quality: A Construct Validity Study of Neural Machine Translation Metrics

Dmytro Chaplynskyi, Ivan Kulynych, Maria Shvedova, Lesia Ivashkevych

Abstract

Automatic machine translation metrics are the de facto standard for evaluating translation quality. Yet, it remains unclear what they actually measure. We investigate this question using a unique multilingual corpus: seven human Ukrainian translations of George Orwell’s Animal Farm, alongside three architecturally distinct AI systems (GPT-5.2, DeepL, and Lapa, a Ukrainian-tuned LLM). Across seven neural metrics, four reference-free and three reference-based, all three AI translations rank at the top. However, stylometric analysis exposes that these same AI translations are not as lexically rich as human ones ($-$18% MTLD), underuse Ukrainian particles (up to 2x fewer) and diminutive morphology (2.6x fewer), and converge on near-identical outputs (LaBSE pairwise similarity 0.941 vs. 0.711 for human pairs). A controlled LLM-as-a-judge experiment demonstrates a clear preference reversal: when the English source is visible, AI ranks first; when it is hidden and the judge evaluates literary quality alone, humans rise to the top and AI falls to the lower ranks. Human evaluation (1,034 pairwise judgments) is balanced across both patterns. We argue that current MT metrics reward semantic fidelity and surface fluency — properties optimized by AI systems — while failing to capture the lexical richness, cultural adaptation, and stylistic voice that characterize skilled literary translation.

Anthology ID:: 2026.unlp-1.8
Volume:: Proceedings of the Fifth Ukrainian Natural Language Processing Conference (UNLP 2026)
Month:: May
Year:: 2026
Address:: Lviv, Ukraine
Editor:: Mariana Romanyshyn
Venue:: UNLP
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 67–79
Language:
URL:: https://aclanthology.org/2026.unlp-1.8/
DOI:
Bibkey:
Cite (ACL):: Dmytro Chaplynskyi, Ivan Kulynych, Maria Shvedova, and Lesia Ivashkevych. 2026. Semantic Fidelity Versus Literary Quality: A Construct Validity Study of Neural Machine Translation Metrics. In Proceedings of the Fifth Ukrainian Natural Language Processing Conference (UNLP 2026), pages 67–79, Lviv, Ukraine. Association for Computational Linguistics.
Cite (Informal):: Semantic Fidelity Versus Literary Quality: A Construct Validity Study of Neural Machine Translation Metrics (Chaplynskyi et al., UNLP 2026)
Copy Citation:
PDF:: https://aclanthology.org/2026.unlp-1.8.pdf

PDF Cite Search Fix data