Human-in-the-loop Machine Translation with Large Language Model

Xinyi Yang; Runzhe Zhan; Derek F. Wong; Junchao Wu; Lidia S. Chao

Human-in-the-loop Machine Translation with Large Language Model

Xinyi Yang, Runzhe Zhan, Derek F. Wong, Junchao Wu, Lidia S. Chao

Abstract

The large language model (LLM) has garnered significant attention due to its in-context learning mechanisms and emergent capabilities. The research community has conducted several pilot studies to apply LLMs to machine translation tasks and evaluate their performance from diverse perspectives. However, previous research has primarily focused on the LLM itself and has not explored human intervention in the inference process of LLM. The characteristics of LLM, such as in-context learning and prompt engineering, closely mirror human cognitive abilities in language tasks, offering an intuitive solution for human-in-the-loop generation. In this study, we propose a human-in-the-loop pipeline that guides LLMs to produce customized outputs with revision instructions. The pipeline initiates by prompting the LLM to produce a draft translation, followed by the utilization of automatic retrieval or human feedback as supervision signals to enhance the LLM’s translation through in-context learning. The human-machine interactions generated in this pipeline are also stored in an external database to expand the in-context retrieval database, enabling us to leverage human supervision in an offline setting. We evaluate the proposed pipeline using the GPT-3.5-turbo API on five domain-specific benchmarks for German-English translation. The results demonstrate the effectiveness of the pipeline in tailoring in-domain translations and improving translation performance compared to direct translation instructions. Additionally, we discuss the experimental results from the following perspectives: 1) the effectiveness of different in-context retrieval methods; 2) the construction of a retrieval database under low-resource scenarios; 3) the observed differences across selected domains; 4) the quantitative analysis of sentence-level and word-level statistics; and 5) the qualitative analysis of representative translation cases.

Anthology ID:: 2023.mtsummit-users.8
Volume:: Proceedings of Machine Translation Summit XIX, Vol. 2: Users Track
Month:: September
Year:: 2023
Address:: Macau SAR, China
Editors:: Masaru Yamada, Felix do Carmo
Venue:: MTSummit
SIG:
Publisher:: Asia-Pacific Association for Machine Translation
Note:
Pages:: 88–98
Language:
URL:: https://aclanthology.org/2023.mtsummit-users.8
DOI:
Bibkey:
Cite (ACL):: Xinyi Yang, Runzhe Zhan, Derek F. Wong, Junchao Wu, and Lidia S. Chao. 2023. Human-in-the-loop Machine Translation with Large Language Model. In Proceedings of Machine Translation Summit XIX, Vol. 2: Users Track, pages 88–98, Macau SAR, China. Asia-Pacific Association for Machine Translation.
Cite (Informal):: Human-in-the-loop Machine Translation with Large Language Model (Yang et al., MTSummit 2023)
Copy Citation:
PDF:: https://aclanthology.org/2023.mtsummit-users.8.pdf

PDF Cite Search