Proceedings of the 17th Conference of the Association for Machine Translation in the Americas (Volume 1: Research Track)
Eleftheria Briakou, Jeremy Gwinnup, Shivali Goel (Editors)
- Anthology ID:
- 2026.amta-research
- Month:
- August
- Year:
- 2026
- Address:
- Québec City, Canada
- Venue:
- AMTA
- Event:
- Conference of the Association for Machine Translation in the Americas (2026)
- SIG:
- Publisher:
- Association for Machine Translation in the Americas
- URL:
- https://aclanthology.org/2026.amta-research/
- DOI:
- PDF:
- https://aclanthology.org/2026.amta-research.pdf
Proceedings of the 17th Conference of the Association for Machine Translation in the Americas (Volume 1: Research Track)
Eleftheria Briakou | Jeremy Gwinnup | Shivali Goel
Eleftheria Briakou | Jeremy Gwinnup | Shivali Goel
AMTA Best Thesis Award Abstract: Overcoming Vocabulary Challenges in Natural Language Processing
Elizabeth Salesky
Elizabeth Salesky
This thesis addresses the vocabulary bottleneck in machine translation and other natural language processing applications, exploring more robust and flexible representations of text and their impact on translation quality and cross-lingual generalization. This volume includes a short summary of the thesis; the full thesis is available separately.
Translation-CoT: A Human-Inspired Chain-of-Thought Framework for Multilingual LLM Translation
Tabia Tanzin Prama | Juniper L Lovato | Chris Danforth | Peter Dodds
Tabia Tanzin Prama | Juniper L Lovato | Chris Danforth | Peter Dodds
Large language models (LLMs) have transformed machine translation, yet mistranslations, hallucinations, and unnatural phrasing still limit their effectiveness, particularly for low-resource languages. We propose Translation-CoT, a chain-of-thought prompting strategy that breaks translation into structured stages (lexical retrieval, grammatical analysis, and topic identification), followed by a refinement step to improve fluency, tone, and idiomatic expression. We evaluate Translation-CoT across 14 languages from 14 language families and multiple LLMs (GPT-4o, GPT-4o-mini, LLaMA 3.1, and Gemma 2), with GPT-4o performing best overall, in both English ↔ non-English (X) translation settings. Compared with zero-shot prompting, in-context learning, and existing chain-of-thought prompting methods (Tree-of-Thought (ToT) and Learning-Oriented Prompting (LOT)), Translation-CoT outperforms these prompting strategies on multilingual machine translation across BLEU, ChrF, and METEOR, with especially strong gains in the more difficult English→non-English (X) setting and in low-resource languages. Human evaluation further shows higher preference scores and lower MQM penalty scores, indicating fewer mistranslations, omissions, awkward phrasing, and hallucinations with Translation-CoT. Overall, our results show that structured, task-aware prompting is an effective approach for improving multilingual translation quality and robustness in LLMs.
Toward Equitable Machine Translation for Atypical Speech: An LLM Post-Correction Approach
Grace Pasion | Ammon Shurtz | Steve Richardson
Grace Pasion | Ammon Shurtz | Steve Richardson
Individuals with speech disabilities rely on everyday technologies powered by Automatic Speech Recognition (ASR) systems, yet these systems consistently fail them–producing significantly higher error rates that undermine the usefulness of voice assistants, hands-free devices, and machine translation pipelines. We conduct a multi-stage evaluation of speech impairment effects in cascaded speech-to-text translation, examining two distinct conditions: real dysarthric speech and simulated rhotacism. For dysarthria, we quantify ASR error rates. For rhotacism, we quantify error rates from a minimal-pair text substitution. We then analyze how impairment-induced errors propagate through downstream machine translation across three language directions (English to Spanish, Ukrainian, Khmer), and propose a training-free LLM-based post-correction methodology as an accessible intervention. We find that larger LLMs (70B parameters or more) consistently improve downstream translation quality even when surface-level corrections made by those LLMs remain modest, while smaller models lack the capacity to do so reliably. These results reveal a promising but scale-dependent path toward more equitable speech technology for users with atypical speech.
Peer-review–based translator training promotes reflection and collaborative critique. This study examines whether GPT5, guided by MQM-like prompts, can function as a peer-review training partner rather than a grading tool. Using translated passages from a practice group, we compared GPT5’s feedback with human evaluations of the same segments, including both negative and positive judgments. GPT5 aligned with human evaluators on 77.8% of negative flags and 88.9% of positive flags, and achieved an F1 score of 0.875 for detailed rationales supporting the flags. The results suggest that GPT5 can provide useful analyses and alternative perspectives that support learner reflection, although its occasional poor judgments indicate that it should be used as a supplementary training partner rather than a standalone evaluator.
Predict and Fix: A Unified Model for Translation Quality Estimation and Post-Editing
Maciej Modrzejewski | Yash Bhaskar | Chinmay Pateria
Maciej Modrzejewski | Yash Bhaskar | Chinmay Pateria
We propose a unified architecture for jointly modeling Translation Quality Estimation (QE) and Automatic Post-Editing (APE) within a single lightweight language model. Our approach integrates quality prediction and correction generation in a single decoding process using a decoder-only Qwen2.5 model (0.5B parameters), augmented with a dedicated QE regression head operating on hidden states at a special token position. The model produces structured outputs that include a continuous quality score, an edit decision, and a corrected translation when necessary. We train on datasets of 100K, 1M, and 1.84M manually annotated samples across eight language pairs, enabling analysis of both data scale and distribution. Experimental results show that the proposed model achieves strong QE performance (r=0.907) and high post-editing decision accuracy (88.4%), while reducing over-editing compared to both autoregressive baselines and large commercial LLMs.
BRIDGE-MT: A Benchmark for Role Interactions and Dependencies in Machine Translation Gender Evaluation
Neha Gajakos | Christopher Staff | Brenda Murphy | John D. Kelleher | Rejwanul Haque
Neha Gajakos | Christopher Staff | Brenda Murphy | John D. Kelleher | Rejwanul Haque
This paper investigates gender behavior in Hindi–English machine translation (MT) within multi-entity settings, especially when two occupational roles appear within the same sentence. Existing benchmarks often focus on single-referenced entities, leaving cross-role dependencies largely unexplored. We define a taxonomy of thirteen role-gender configurations covering masculine (m), feminine (f), and neutral (n) assignments and introduce BRIDGE-MT, a manually created dataset of 351 Hindi–English sentence pairs (594 role-level instances) in order to evaluate dual-role interactions. We evaluated commercial MT systems and multilingual LLMs, and propose neutral-comparison asymmetry metrics and a conditional interaction metric to analyze cross-role dependencies. Our results show that explicitly gendered roles achieve higher F1 scores than neutral-labelled roles. We also observe a consistent position effect, where Role B (the second role) tends to have lower accuracy and greater gender asymmetry than Role A (the first role) across all evaluated systems. Conditional interaction analysis further indicates that the gender assigned to one role can influence the translation of the other. These findings highlight the importance of evaluating gender behavior in multi-entity settings to better understand interaction-driven asymmetries in MT.
Improving Term Evaluation in Machine Translation: Variation Matters
Nicolas Dahan | Ziqian Peng | François Yvon | Rachel Bawden
Nicolas Dahan | Ziqian Peng | François Yvon | Rachel Bawden
Terminology evaluation in machine translation (MT) usually assumes a single correct target form per source term, yet human translators routinely introduce variation that current metrics penalize as inconsistency. We examine how to account for this variation in document-level MT evaluation in English–French scientific translation, combining glossary-based accuracy, translation consistency, and a new cross-term variation (CTV) diagnostic measure that captures whether variation relationships are preserved across languages. On two parallel corpora translated by four MT systems, we find that (1) MT systems generate less target-side variation than human translators; (2) transfer patterns strongly depend on the variation type; (3) consistency rankings vary with the choice of metric; and (4) constraining MT with a glossary improves accuracy and consistency but degrades CTV by suppressing valid variation. We argue for variation-aware evaluation that conditions consistency penalties on whether target-side variation mirrors source-side variation.
Layout-Based Chunk Alignment: Utilizing Visual Information to Collect Parallel Texts From Image Documents
Masaki Kinouchi | Kayoko Nohara | Xinru Zhu | Yuma Miura
Masaki Kinouchi | Kayoko Nohara | Xinru Zhu | Yuma Miura
This study proposes a layout-based chunk alignment method (Layout-CA) as an intermediate step between document- and sentence-level parallel text alignment for bilingual document images. Visually rich printed materials, such as institutional reports and magazines, often contain high-quality translations and are valuable sources of parallel data, yet their layout cues are underutilized. Layout-CA aligns semantically coherent text chunks across document pairs by integrating multi-modal cues from textual content and layout, and sentence alignment is then performed within the aligned chunk pairs. Experiments on English UNESCO reports and their Japanese translations show that introducing chunk alignment improves downstream sentence alignment for both Bleualign and Vecalign. When document order is disrupted, Layout-CA preserves alignment coverage by restricting sentence matching to corresponding chunks, enabling robust alignment in multilingual image documents.
Crosslingual Disparities in LLM Performance: Challenges for MT as Mitigation
Rebecca Knowles | Cyril Goutte
Rebecca Knowles | Cyril Goutte
We examine crosslingual performance disparities in large language models (LLMs) in the context of safety- and regulation-related queries in Canada. We manually build a set of English and French query pairs with gold standard answers and collect LLM-generated answers, which are manually annotated for correctness. We find that LLMs are more likely to produce errors in their answers in French than in English. We investigate a machine translation pipeline, translating the French query, producing an English LLM response, and translating the response back to French. We find that, while it can mitigate some of these performance disparities, additional challenges such as the reliability and language of the cited sources or technical terms greatly impact that mitigation strategy.
Translators’ Perceptions and Edit Traces: Quality of MT as a Tool in Canadian Parliamentary Translation
Jeniffer Leal-Wyss | Gabriel Bernier-Colborne | Delaney Lothian | Michel Simard | Rebecca Knowles
Jeniffer Leal-Wyss | Gabriel Bernier-Colborne | Delaney Lothian | Michel Simard | Rebecca Knowles
The Parliament of Canada’s translation workflow includes access to a specialized neural machine translation (NMT) system. This study analyzes post-editing (PE) activity to identify the types of edits translators make when interacting with the NMT system, as well as the frequency, nature, and severity of errors encountered. We compare translations produced with and without the use of this NMT system to evaluate potential differences in edit patterns. To complement this analysis, we draw on insights from a user study. Our findings explore how translators’ perceptions align with observed PE patterns and how their feedback can inform strategies to better understand, and possibly mitigate, some of the errors observed.
Document Summarization for AI-based Post-Editing
Vera Senderowicz Guerra | Dimitrios Pavlou | Peter Bourgonje | Olesia Khrapunova | Konstantinos Karageorgos | Aaron Schliem
Vera Senderowicz Guerra | Dimitrios Pavlou | Peter Bourgonje | Olesia Khrapunova | Konstantinos Karageorgos | Aaron Schliem
Post-Editing (PE) is typically performed on isolated segments or small batches, without access to broader document context. In this paper, we investigate whether pre-generated, document-level summaries can improve PE quality. Using a purpose-built summarization prompt evaluated across nine LLMs from OpenAI and Google, we select two models with contrasting summary styles for downstream experiments on 448 documents covering 37 target locales and 13 content domains. Summaries generated by gemini-2.5-flash-lite, which are directive and domain-specific, yield gains in edit distance and modest gains in COMET, whereas those generated by GPT-4o, which tend to be more generic and descriptive, degrade performance across most metrics. The positive effect appears most pronounced in terminologically dense domains and lower-resource locales. A qualitative analysis shows that improvements arise when summaries provide specific, actionable guidance on terminology, domain conventions, and style, and that performance decreases when summaries are underspecified or conflicting. These findings suggest that summary specificity and actionability, rather than the mere addition of context, determine whether document-level information benefits post-editing.
Benchmarking Large Language Models for Game Localization Quality Assurance: A Cross-Model, Cross-Lingual Analysis
Mao Tian | Na Wu
Mao Tian | Na Wu
Localization quality assurance (LQA) is a critical component of game development, where manual review of large volumes of translated text is time-consuming and costly. Recent advances in large language models (LLMs) suggest strong potential for automated LQA, yet their effectiveness across different models, target languages, and game domains remains insufficiently understood. We present a comprehensive benchmark evaluating eight LLMs, including both closed-source and open-weight models, on English-to-six-language gaming LQA tasks across two game genres. Our dataset comprises 96 evaluation settings with a total of 48,000 translation samples. The results show that Claude Sonnet 4 achieves the best overall performance (F1 = 0.766), followed by Qwen-2.5-72B (F1 = 0.711) and Gemini 2.0 Flash (F1 = 0.691). We observe that (1) the target language does not significantly affect model performance (p = 0.285), (2) models achieve their most consistent performance on French, while Japanese is the most challenging target language, and (3) game genre (RPG vs. strategy) has minimal impact on accuracy. While closed-source models achieve the highest overall performance, open-weight alternatives such as Qwen-2.5-72B provide competitive quality at substantially lower cost. These findings provide practical guidance for deploying LLM-based LQA systems in production game localization workflows.