Joachim Minder
Author directory2026
On the Use of LLMs for Specialised Terminology: A Good Alternative to Corpora?
Joachim Minder | Guillaume Wisniewski | Natalie Kübler
Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 1)
Joachim Minder | Guillaume Wisniewski | Natalie Kübler
Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 1)
Specialised translation relies on the use of documentary and terminological resources, including corpora. These resources are particularly useful for terminology. However, their compilation and exploitation have several limitations: they require time, technical skills and access to data that can be difficult to collect. This study examines the extent to which LLMs can assist specialised translators in finding equivalents from English to French. We evaluate four proprietary models, GPT-4o, GPT-5.2, Claude Sonnet 4.5 and DeepSeek, in two specialised domains, Earth, Environmental and Planetary Sciences (EEPS) and Natural Language Processing (NLP). The experiment is based on 80 terms per domain and compares two prompting strategies: a terminology and a translation mode. The results highlight clear differences between models, prompting strategies and, to a lesser extent, domains. Claude Sonnet 4.5 achieves the best results in the most favourable configuration, while DeepSeek stands out for its greater stability. Analysis of confidence estimates also shows that they are only a partial indicator of terminological accuracy. Overall, the findings suggest that LLMs can be useful tools for specialised translators, but cannot, at this stage, replace specialised corpora. This research therefore paves the way for future work on the real practical usefulness of LLMs for specialised translators in work and educational contexts.
2025
MaTOS: Machine Translation for Open Science
Rachel Bawden | Maud Bénard | Éric de la Clergerie | José Cornejo Cárcamo | Nicolas Dahan | Manon Delorme | Mathilde Huguin | Natalie Kübler | Paul Lerner | Alexandra Mestivier | Joachim Minder | Jean-François Nominé | Ziqian Peng | Laurent Romary | Panagiotis Tsolakis | Lichao Zhu | François Yvon
Proceedings of Machine Translation Summit XX: Volume 2
Rachel Bawden | Maud Bénard | Éric de la Clergerie | José Cornejo Cárcamo | Nicolas Dahan | Manon Delorme | Mathilde Huguin | Natalie Kübler | Paul Lerner | Alexandra Mestivier | Joachim Minder | Jean-François Nominé | Ziqian Peng | Laurent Romary | Panagiotis Tsolakis | Lichao Zhu | François Yvon
Proceedings of Machine Translation Summit XX: Volume 2
This paper is a short presentation of MaTOS, a project focusing on the automatic translation of scholarly documents. Its main aims are threefold: (a) to develop resources (term lists and corpora) for high-quality machine translation; (b) to study methods for translating complete, structured documents in a cohesive and consistent manner; (c) to propose novel metrics to evaluate machine translation in technical domains. Publications and resources are available on the project web site: https://anr-matos.gihub.io.
Testing LLMs’ Capabilities in Annotating Translations Based on an Error Typology Designed for LSP Translation: First Experiments with ChatGPT
Joachim Minder | Guillaume Wisniewski | Natalie Kübler
Proceedings of Machine Translation Summit XX: Volume 1
Joachim Minder | Guillaume Wisniewski | Natalie Kübler
Proceedings of Machine Translation Summit XX: Volume 1
This study investigates the capabilities of large language models (LLMs), specifically ChatGPT, in annotating MT outputs based on an error typology. In contrast to previous work focusing mainly on general language, we explore ChatGPT’s ability to identify and categorise errors in specialised translations. By testing two different prompts and based on a customised error typology, we compare ChatGPT annotations with human expert evaluations of translations produced by DeepL and ChatGPT itself. The results show that, for translations generated by DeepL, recall and precision are quite high. However, the degree of accuracy in error categorisation depends on the prompt’s specific features and its level of detail, ChatGPT performing very well with a detailed prompt. When evaluating its own translations, ChatGPT achieves significantly poorer results, revealing limitations with self-assessment. These results highlight both the potential and the limitations of LLMs for translation evaluation, particularly in specialised domains. Our experiments pave the way for future research on open-source LLMs, which could produce annotations of comparable or even higher quality. In the future, we also aim to test the practical effectiveness of this automated evaluation in the context of translation training, particularly by optimising the process of human evaluation by teachers and by exploring the impact of annotations by LLMs on students’ post-editing and translation learning.