Genre Transfer in NMT:Creating Synthetic Spoken Parallel Sentences using Written Parallel Data

Nalin Kumar, Ondrej Bojar


Abstract
Text style transfer (TST) aims to control attributes in a given text without changing the content. The matter gets complicated when the boundary separating two styles gets blurred. We can notice similar difficulties in the case of parallel datasets in spoken and written genres. Genuine spoken features like filler words and repetitions in the existing spoken genre parallel datasets are often cleaned during transcription and translation, making the texts closer to written datasets. This poses several problems for spoken genre-specific tasks like simultaneous speech translation. This paper seeks to address the challenge of improving spoken language translations. We start by creating a genre classifier for individual sentences and then try two approaches for data augmentation using written examples:(1) a novel method that involves assembling and disassembling spoken and written neural machine translation (NMT) models, and (2) a rule-based method to inject spoken features. Though the observed results for (1) are not promising, we get some interesting insights into the solution. The model proposed in (1) fine-tuned on the synthesized data from (2) produces naturally looking spoken translations for written-to-spoken genre transfer in En-Hi translation systems. We use this system to produce a second-stage En-Hi synthetic corpus, which however lacks appropriate alignments of explicit spoken features across the languages. For the final evaluation, we fine-tune Hi-En spoken translation systems on the synthesized parallel corpora. We observe that the parallel corpus synthesized using our rule-based method produces the best results.
Anthology ID:
2022.icon-main.28
Volume:
Proceedings of the 19th International Conference on Natural Language Processing (ICON)
Month:
December
Year:
2022
Address:
New Delhi, India
Editors:
Md. Shad Akhtar, Tanmoy Chakraborty
Venue:
ICON
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
224–233
Language:
URL:
https://aclanthology.org/2022.icon-main.28
DOI:
Bibkey:
Cite (ACL):
Nalin Kumar and Ondrej Bojar. 2022. Genre Transfer in NMT:Creating Synthetic Spoken Parallel Sentences using Written Parallel Data. In Proceedings of the 19th International Conference on Natural Language Processing (ICON), pages 224–233, New Delhi, India. Association for Computational Linguistics.
Cite (Informal):
Genre Transfer in NMT:Creating Synthetic Spoken Parallel Sentences using Written Parallel Data (Kumar & Bojar, ICON 2022)
Copy Citation:
PDF:
https://aclanthology.org/2022.icon-main.28.pdf