Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 2)

Dimitar Shterionov, Eva Vanmassenhove, Mirella De Sisto, Fred Blain, Javad Pourmostafa Roshan Sharami, Lisa Lepp, Chiara Manna, Argentina Anna Rescigno, Alina Karakanta, Ayla Rigouts Terryn, Manuel Lardelli, Natalia Resende, Elena Murgolo, Janiça Hackenbuchner, Anna Zaretskaya, Miquel Esplà-Gomis, Thierry Etchegoyhen, Dagmar Gromann, Rachel Bawden, Barry Haddow, Sara Szoc, Mikel Forcada, Helena Moniz (Editors)


Anthology ID:
2026.eamt-2
Month:
June
Year:
2026
Address:
Tilburg, The Netherlands
Venue:
EAMT
Event:
Conference of the European Association for Machine Translation (2026)
SIG:
Publisher:
European Association for Machine Translation
URL:
https://aclanthology.org/2026.eamt-2/
DOI:
10.26116/9789403901404
ISBN:
9789403901404
Bib Export formats:
BibTeX MODS XML EndNote
PDF:
https://aclanthology.org/2026.eamt-2.pdf

he Erasmus+-funded international research consortium LT-LiDER develops a range of digital training resources which are grounded in the overarching frameworks of digital and AI literacy and oriented towards practical application contexts in the language and translation industry. These resources can be implemented on a component basis or as a complete curriculum in higher-education language and translation classrooms.
This paper reports on the final stages of the MaTIAS project. A functional prototype of the multilingual notification tool was deployed across seven Belgian reception centres, accompanied by training and technical support. Feedback was gathered through interviews and surveys. Two rounds of machine translation evaluation revealed considerable differences in quality across languages. The translation quality of Tigrinya in particular was deemed too low to be usable.
The Anglocentric nature of scholarly communication has many implications, such as limiting publication, discoverability and access from other language communities (even for major languages); putting minoritized languages at risk in the academic domain; and excluding many from peer review. The OSCAIL project addresses these challenges by exploring how machine translation (MT) enhanced by large language model (LLM)–based technologies can support access to scientific knowledge. Outputs will include evaluation datasets, protocols and best practices for MT in scholarly communication, and a prototype integration of MT tools into Open Journal Systems, the world’s most widely used open-source scholarly publishing platform.
We present AIDA Agents, a multi-agent translation platform that orchestrates LLM-based agents—for translation, rating, post-editing, and re-rating—delivering context-aware translations without model fine-tuning. Optional retrieval-augmented generation (RAG) injects translation memories, terminology, and style guidelines at every pipeline stage. On WMT24++ (Deutsch et al., 2025) (11 languages), AIDA Agents outperforms all systems on 10 of 11 pairs. On an industrial benchmark, 70–98% of segments are publication-ready without human post-editing. The platform is deployed with native XLIFF integration.
We present VERA, an easy-to-use platform for machine translation (MT) evaluation that integrates automatic and human evaluation within a single web environment. VERA supports standard reference based metrics and multiuser annotation following the Multidimensional Quality Metrics (MQM) Core framework. The platform enables the export of annotated corpora and the generation of final PDF reports summarizing both automatic and human evaluation results, including correlations between them.
This document presents an initial overview of the CLingS project, currently in its early development stage. It outlines a collaborative effort to build a cross-lingual information retrieval platform for scientific literature in underrepresented languages. The project CLingS aims to develop datasets, tools, and methods to improve multilingual access to scientific knowledge.
The ARTICULATE project is an ambitious and interdisciplinary initiative funded by the CHIST-ERA call 2025. Its vision is to revolutionize science education and democratize scientific knowledge beyond academia and English-speaking audiences through the integration of AI with self-regulated learning. The aim is to translate science not just across language but across language style, to create engaging spoken digital experiences. We present an introduction to this project, an overview of the consortium and research approach, and a number of expected impacts.
Human evaluation remains essential for reliable machine translation (MT) assessment, yet practical evaluation workflows are often difficult to reproduce and scale. Here we introduce HERMeS, a lightweight human evaluation platform designed to streamline systematic human evaluation and comparison of multiple MT systems across large translation sets. Unlike existing evaluation tools, HERMeS focuses specifically on scalable comparison of many anonymized systems through a hybrid ranking and direct assessment workflow, using a novel approach that reduces evaluator cognitive load while maintaining data quality, security, and integrity.
Domain adaptation remains a major challenge for machine translation, par-ticularly in institutional communica-tion. This paper presents the MULTI-TRAD project , which develops English–Spanish parallel corpora for the Third Social Sector communication. The project integrates three comple-mentary objectives: (i) the compilation of a domain-specific parallel corpus, (ii) the analysis of linguistic variation across human translation (HT), ma-chine translation (MT), and post-edited (PE) texts using Multidimensional Analysis (Biber, 1988), and (iii) the development of a domain-adapted neu-ral machine translation system. In par-ticular, the project investigates how dif-ferent translation processes give rise to distinct functional profiles, related to phenomena such as translationese and post-editese. This paper presents the project design and initial progress.
This paper presents PCDT, a web-based platform for collecting sentence-aligned parallel corpora through a community-driven approach to support machine translation for under-resourced languages. The tool decentralizes the translation task to the target community and subsequently reviewed by language experts.
This article describes the research project aimed at developing a Trilingual Machine translation System for English, Nepali, and Tamang language pairs. This project is expected to address knowledge and communication gaps caused by language barriers and mitigate disparities in the availability of information and knowledge sources in Tamang and Nepali.
The TELÓ project provides an open-source framework for automated subtitling in the performing arts. Integrating state-of-the-art ASR and NMT, the system enables bidirectional translation between Catalan, Spanish, English and French. Designed for live performances, it provides synchronized captions for multiple devices, enhancing cultural internationalization and accessibility.
DA + Criteria is a translation quality assessment method proposed based on a comprehensive systematic literature review on the concepts of quality in machine translation and translation studies. In the presented project the method was tested alongside MQM on the German translations by humans, DeepL and ChatGPT of English non-fiction texts, using the results of the study as well as the participants’ answers to further refine the method.
We present ACATMT, a compact bilingual encoder-decoder NMT system for English and Swedish, designed for professional computer-assisted translation (CAT) tools. It runs on-device in ONNX format, under 1 GB of RAM with no GPU needed, and features real-time post-edit based terminology adaptation. It also supports translation memory conditioning via decoder prefilling. Evaluation on 5,021 technical segments unseen during training shows significant improvements in COMET and BLEU when using glossaries.
Video-based sign language dictionary search – in which a user records a sign to retrieve its translation – has been increasingly studied, yet never deployed in a large-vocabulary setting. We present the first such deployment: a fully scalable video-based search system integrated into the Flemish Sign Language (VGT) Dictionary, comprising over 11,000 signs. The system, released on November 28th, 2025, requires no retraining as new signs are added, and was validated on data collected in the wild. It was developed through an equal partnership between the deaf-led Flemish Sign Language Centre (VGTC) and AI researchers from Ghent University, and shows that closing the gap between sign language research and community impact is both achievable and essential.
This paper presents the TaMTAS project (Terminology-Aware Machine Translation for Accessible Science), a research project coordinated by the Universitat Oberta de Catalunya (UOC) to develop an open- source translation ecosystem for the Life Sciences. While we provide a general overview of the project’s organization into seven Work Packages (WPs) and its col- laborative consortium, this article focuses specifically on the work of WP2. Led by the UOC, this package is responsible for the parallel corpus compilation for five lan- guages (English, Spanish, Catalan, Esto- nian, and Irish), the enhancement of TBX- Tools for terminology extraction, and the development of synthetic data augmenta- tion strategies. These linguistic assets are essential to power the downstream Large Reasoning Models (LRMs) and Automatic Post-Editing (APE) modules, ensuring ter- minological consistency in highly special- ized scientific domains.
Recently, generative artificial intelligence (GenAI) has been perceived as a "silver bullet" for achieving faster, cheaper, and better translation production. However, in professional localisation, AI capabilities alone are not enough, as the still time-consuming post-editing (PE) of machine translation (MT) and GenAI output proves. The features and processes presented in this work aim to reduce these ef-forts by enhancing terminological control and translation consistency within the CAT environment STAR Transit.
Translation 2.0 addresses a critical gap in accessible, up-to-date educational resources on recent developments in Machine Translation and Large Language Models for students of linguistics and translation. It develops an online module with open-access learning materials, including knowledge clips, a workbook with incremental exercises to consolidate conceptual understanding, practical coding guides, and industry professional videos. The module aims to build both subject knowledge and computational literacy, freeing up contact hours for deeper engagement and critical discussions on practical, professional and ethical aspects. Translation 2.0 is funded through an Educational Innovation grant by the Faculty of Humanities at Leiden University and ECOLe (Expert Centre for Education and Learning) and runs from February to December 2026.
Access to the primary labor market for people with cognitive impairments is hampered by barriers, notably the lack of workplace information in Easy Language (EL). Producing such texts is time- and cost-intensive and requires specialized translators. The project STARK-LS (Strengthening participation in the primary labor market through AI-generated Easy Language) addresses this gap by implementing an AI-translation tool to translate workplace materials into EL and embedding the approach in internships for people with cognitive impairments. An interdisciplinary project team conducts mixed-methods evaluations by testing the EL translations for applicability, comprehensibility, and acceptance using lab-based eye-tracking and questionnaire studies, qualitative interviews with interns with cognitive impairments and experts for EL, and a quantitative online survey with company representatives. The findings will lead to process models and best-practice recommendations for companies and rehabilitation agencies. The project advances scientific understanding of the perceived usefulness and potential barriers of EL in organizational contexts, while critically evaluating AI’s influence on the diffusion of high-quality EL texts in companies. Funded by the German Federal Ministry of Labour and Social Affairs, the project aims to scale high-quality accessible communication and promote sustainable inclusion in the German labor market.
Prompsit Language Engineering is launching an updated API and CLI for its open-source, planet-friendly machine translation services. Operating on a freemium model, the tools offer free limited access alongside tiered pricing for advanced features like MT evaluation, quality estimation, corpus scoring, and multilingual dataset annotation.
This paper presents the Multilingual, Multicultural, and Multimodal Medical Language Processing (4MLP) Project, funded through a competitive call of the University of Naples ”L’Orientale” (Italy). 4MLP aims at developing multilingual,multicultural, and multimodal language technologies for healthcare to bridge complex medical knowledge and patients’ needs, supporting inclusive and effective healthcare communication while advancing explainable and culturally aware Artificial Intelligence.
This paper introduces Ouvia, a research project to assess user-perceived usability and reliability of modern speech translation tools in EnPt scenarios. The project centers on a user study in which we simulate real-life daily interactions by recruiting crowdworkers online from different sociodemographic groups. We collect their spoken requests and self-assessments about quality, satisfaction, and reliability. Here, we describe the project’s motivation and objectives, the study design, and the expected outcomes we will provide to speech translation practitioners.
The CRITICS project addresses science accessibility and literacy through the convergence of advanced Machine Translation (MT) based on Large Language Models (LLMs) and educational technology. By leveraging MT systems specifically optimized for scientific content, educational institutions can provide accurate, culturally relevant translations of scientific materials in higher-education students’ native languages, ensuring that complex scientific concepts are comprehensible while maintaining technical accuracy. Novel research on MT specifically tailored for scientific documents aims to break down language barriers in accessing cutting-edge research and educational materials currently only available in high-resourced languages such as English, thereby facilitating the democratization of scientific knowledge across linguistic boundaries.
This study evaluates an AI post-editing (AIPE) system in a professional translation setting, covering translation from English into ten target languages across five domains. We evaluate the system using automatic metrics on 71,262 production segments and human evaluation on a stratified sample of 6,618 segments (approximately 600 segments per target language) assessed by 60 professional translators. AIPE refines machine translation output using a secure publicly available LLM, retrieving language-specific style guides and high-quality bilingual examples to guide edits. We compare it with direct LLM translation (LLMT), Google Translate, and DeepL. The two AIPE configurations evaluated consistently outperform the generic translation baselines in terms of quality. LLMT does not match this quality, though it may suit less quality-sensitive domains. We observe how AIPE’s gains vary according to pre-translation type, with fuzzy translation memory matches over-represented among severe errors, and discuss deployment implications.
We investigate whether reasoning information can enhance machine translation when incorporated as supportive context during training and inference. Using Hindi-Bengali translation as a case study, we define five reasoning components: Key Terms, Syntactic, Semantic, Pragmatic, and Paraphrase. We conduct a complete ablation across all 31 possible combinations using Gemma-3-1B-Instruct and evaluate on multi-domain benchmark with BLEU, chrF, and TER. Evaluation results show that reasoning effectiveness depends on its type and composition rather than quantity. Combining multiple heterogeneous signals causes objective diffusion, degrading performance. The compact Semantic and Paraphrase combination proves optimal, and providing it during inference yields 23.86 BLEU compared to 22.12 from standard fine-tuning a +1.74 BLEU gain across eight domains. These findings demonstrate that targeted semantic guidance consistently and meaningfully improves the compact translation models.
Machine translation quality estimation (QE) typically relies on dedicated neural models trained on human judgments. We evaluate whether cosine similarity over general-purpose embeddings can serve as a lightweight alternative, using Gemini embeddings as the scoring backbone. Through three experiments (rogue dimension analysis, score calibration, and a learned calibration head) and a root cause analysis, we find that cosine similarity between source and translation saturates in the 0.94–0.99 range because even poor translations preserve most of the source semantics, leaving an Area Under the ROC Curve (AUC) ceiling of approximately 0.63. However, a LightGBM classifier trained on normalized cosine and surface-level text features breaks through this ceiling (AUC 0.751), with the improvement driven primarily by features orthogonal to embedding similarity.
Valencian, the Western Catalan variety used in the Valencian Community of Spain, lacks a dedicated language code in most multilingual machine translation (MT) systems, and is systematically rendered closer to the standard written Eastern Catalan used in Catalonia. We address this gap by adapting TranslateGemma-4B-IT, a 4-billion-parameter instruction-tuned (IT) large language model (LLM) specialized for translation, via three post-training strategies using only public corpora and Quantized Low-Rank Adaptation (QLoRA): (i) supervised fine-tuning (SFT); (ii) Group Relative Policy Optimization (GRPO), a reinforcement learning (RL) technique, with chrF plus a naturalness reward (GRPOV1); and (iii) GRPO with a composite automatic-metric reward (GRPOV2). Our results suggest that reward-function alignment with the target dialect is a key determinant of RL success in low-resource dialectal MT.
Movie subtitle translation is inherently multimodal, yet text-only systems often miss visual cues needed to convey emotion, action, and social nuance, especially for low-resource Indic languages (English to Hindi, Bengali, Telugu, Tamil and Kannada). We present a case study on five full-length films and compare two lightweight visual grounding strategies: structured attribute summaries from a 5-minute sliding window and free-text summaries of inter-subtitle visual gaps. Our analysis shows that temporal misalignment between subtitles and frames is a major obstacle in long-form video, often rendering indiscriminate visual grounding ineffective. However, oracle selective grounding, which replaces only the lowest-quality 20-30
Vision-language models (VLMs) have the potential to enhance machine translation (MT) by leveraging visual context alongside text, yet their real utility for production workflows remains unclear. We conduct a unified, multi-condition evaluation of six leading VLMs—both open and proprietary—on two challenging benchmarks (CoMMuTE and CaMMT), targeting lexical and cultural disambiguation respectively, with a domain-style case study simulating technical documentation localization. Results show that model performance varies widely, and the benefit of relevant images does not necessarily transfer across use cases. Proprietary models are notably sensitive to irrelevant images while open-source models are generally more stable; incorrect or contradicting visuals, by contrast, degrade translation across all models. Taken together, these findings make rigorous evaluation a necessary precondition for production deployment: metric gains can mask real accuracy losses in technical domains, model sensitivity to irrelevant images should inform model selection, and reliable image–text matching is a hard requirement for any pipeline.
English dominates scientific publishing, which disadvantages researchers who are not native English speakers, especially those in the earlier stages of their careers. Being able to write and engage with scientific content written in their own language would clearly facilitate scientific production. The MaTOS project (Machine Translation for Open Science) seeks to reduce these barriers by developing machine translation tools for scientific documents in English and French. This article presents the design of the MaTOS pipeline for the HAL platform to automatically translate article abstracts, with author validation, to increase the number of bilingual abstracts on the platform. We also report preliminary experiments comparing translation of sentence, three-sentence chunks, and whole abstracts, evaluated using quality estimation metrics.
Style guides are a centrepiece of professional translation workflows. Yet, their integration into automatic pipelines remains underexplored. This paper presents exploratory work on information extraction from client style guides and application to a templated style guide, developed to be a system prompt. This template is then applied during an LLM-based translation to automatically produce outputs that are compliant to client’s requirements. The study focused on seven language pairs~(LP), evaluating the automatic extraction, and translation quality and compliance with the style guide. The extraction demonstrated reliable performance across languages and file formats. Translation quality and adherence were evaluated using human preference annotation, comparing two Tower models (Tower Zen 9B and Tower+ 72B). The results indicate a modest advantage for Tower+, but with mutual acceptability in certain instances. These findings establish a viable semi-automatic framework for style guide integration in translation workflows, and motivate further investigation across broader domains, clients, and LPs.
Since 2023, translators for the Parliament of Canada have had the option to use neural machine translation (NMT) technology provided by the National Research Council of Canada (NRC) to support their work in translating parliamentary publications between French and English. We present our analysis of an anonymized dataset of translators’ interactions with our Hawkeye MT systems, collected since their introduction and covering a period of 2.5 years. This data provides a unique perspective on how translators interact with the systems, how their use evolved over time and how it impacts the nature of their translations.