Is a Picture Worth a Thousand Words? Exploration and Implementation Considerations for Visual Context in Translation Workflows

Vera Senderowicz Guerra, Olesia Khrapunova


Abstract
Vision-language models (VLMs) have the potential to enhance machine translation (MT) by leveraging visual context alongside text, yet their real utility for production workflows remains unclear. We conduct a unified, multi-condition evaluation of six leading VLMs—both open and proprietary—on two challenging benchmarks (CoMMuTE and CaMMT), targeting lexical and cultural disambiguation respectively, with a domain-style case study simulating technical documentation localization. Results show that model performance varies widely, and the benefit of relevant images does not necessarily transfer across use cases. Proprietary models are notably sensitive to irrelevant images while open-source models are generally more stable; incorrect or contradicting visuals, by contrast, degrade translation across all models. Taken together, these findings make rigorous evaluation a necessary precondition for production deployment: metric gains can mask real accuracy losses in technical domains, model sensitivity to irrelevant images should inform model selection, and reliable image–text matching is a hard requirement for any pipeline.
Anthology ID:
2026.eamt-2.29
Volume:
Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 2)
Month:
June
Year:
2026
Address:
Tilburg, The Netherlands
Editors:
Dimitar Shterionov, Eva Vanmassenhove, Mirella De Sisto, Fred Blain, Javad Pourmostafa Roshan Sharami, Lisa Lepp, Chiara Manna, Argentina Anna Rescigno, Alina Karakanta, Ayla Rigouts Terryn, Manuel Lardelli, Natalia Resende, Elena Murgolo, Janiça Hackenbuchner, Anna Zaretskaya, Miquel Esplà-Gomis, Thierry Etchegoyhen, Dagmar Gromann, Rachel Bawden, Barry Haddow, Sara Szoc, Mikel Forcada, Helena Moniz
Venue:
EAMT
SIG:
Publisher:
European Association for Machine Translation
Note:
Pages:
91–103
Language:
URL:
https://aclanthology.org/2026.eamt-2.29/
DOI:
Bibkey:
Cite (ACL):
Vera Senderowicz Guerra and Olesia Khrapunova. 2026. Is a Picture Worth a Thousand Words? Exploration and Implementation Considerations for Visual Context in Translation Workflows. In Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 2), pages 91–103, Tilburg, The Netherlands. European Association for Machine Translation.
Cite (Informal):
Is a Picture Worth a Thousand Words? Exploration and Implementation Considerations for Visual Context in Translation Workflows (Guerra & Khrapunova, EAMT 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.eamt-2.29.pdf