DecoMT: Decomposed Prompting for Machine Translation Between Related Languages using Large Language Models

Ratish Puduppully, Anoop Kunchukuttan, Raj Dabre, Ai Ti Aw, Nancy Chen


Abstract
This study investigates machine translation between related languages i.e., languages within the same family that share linguistic characteristics such as word order and lexical similarity. Machine translation through few-shot prompting leverages a small set of translation pair examples to generate translations for test sentences. This procedure requires the model to learn how to generate translations while simultaneously ensuring that token ordering is maintained to produce a fluent and accurate translation. We propose that for related languages, the task of machine translation can be simplified by leveraging the monotonic alignment characteristic of such languages. We introduce DecoMT, a novel approach of few-shot prompting that decomposes the translation process into a sequence of word chunk translations. Through automatic and human evaluation conducted on multiple related language pairs across various language families, we demonstrate that our proposed approach of decomposed prompting surpasses multiple established few-shot baseline approaches. For example, DecoMT outperforms the strong few-shot prompting BLOOM model with an average improvement of 8 chrF++ scores across the examined languages.
Anthology ID:
2023.emnlp-main.279
Volume:
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing
Month:
December
Year:
2023
Address:
Singapore
Editors:
Houda Bouamor, Juan Pino, Kalika Bali
Venue:
EMNLP
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
4586–4602
Language:
URL:
https://aclanthology.org/2023.emnlp-main.279
DOI:
10.18653/v1/2023.emnlp-main.279
Bibkey:
Cite (ACL):
Ratish Puduppully, Anoop Kunchukuttan, Raj Dabre, Ai Ti Aw, and Nancy Chen. 2023. DecoMT: Decomposed Prompting for Machine Translation Between Related Languages using Large Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4586–4602, Singapore. Association for Computational Linguistics.
Cite (Informal):
DecoMT: Decomposed Prompting for Machine Translation Between Related Languages using Large Language Models (Puduppully et al., EMNLP 2023)
Copy Citation:
PDF:
https://aclanthology.org/2023.emnlp-main.279.pdf
Video:
 https://aclanthology.org/2023.emnlp-main.279.mp4