Scaling Back-Translation with Domain Text Generation for Sign Language Gloss Translation

Jinhui Ye; Wenxiang Jiao; Xing Wang; Zhaopeng Tu

doi:10.18653/v1/2023.eacl-main.34

Scaling Back-Translation with Domain Text Generation for Sign Language Gloss Translation

Jinhui Ye, Wenxiang Jiao, Xing Wang, Zhaopeng Tu

Abstract

Sign language gloss translation aims to translate the sign glosses into spoken language texts, which is challenging due to the scarcity of labeled gloss-text parallel data. Back translation (BT), which generates pseudo-parallel data by translating in-domain spoken language texts into sign glosses, has been applied to alleviate the data scarcity problem. However, the lack of large-scale high-quality in-domain spoken language text data limits the effect of BT. In this paper, to overcome the limitation, we propose a Prompt based domain text Generation (PGen) approach to produce the large-scale in-domain spoken language text data. Specifically, PGen randomly concatenates sentences from the original in-domain spoken language text data as prompts to induce a pre-trained language model (i.e., GPT-2) to generate spoken language texts in a similar style. Experimental results on three benchmarks of sign language gloss translation in varied languages demonstrate that BT with spoken language texts generated by PGen significantly outperforms the compared methods. In addition, as the scale of spoken language texts generated by PGen increases, the BT technique can achieve further improvements, demonstrating the effectiveness of our approach. We release the code and data for facilitating future research in this field.

Anthology ID:: 2023.eacl-main.34
Volume:: Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics
Month:: May
Year:: 2023
Address:: Dubrovnik, Croatia
Editors:: Andreas Vlachos, Isabelle Augenstein
Venue:: EACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 463–476
Language:
URL:: https://aclanthology.org/2023.eacl-main.34
DOI:: 10.18653/v1/2023.eacl-main.34
Bibkey:
Cite (ACL):: Jinhui Ye, Wenxiang Jiao, Xing Wang, and Zhaopeng Tu. 2023. Scaling Back-Translation with Domain Text Generation for Sign Language Gloss Translation. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 463–476, Dubrovnik, Croatia. Association for Computational Linguistics.
Cite (Informal):: Scaling Back-Translation with Domain Text Generation for Sign Language Gloss Translation (Ye et al., EACL 2023)
Copy Citation:
PDF:: https://aclanthology.org/2023.eacl-main.34.pdf
Video:: https://aclanthology.org/2023.eacl-main.34.mp4

PDF Cite Search Video