Word Alignment by Fine-tuning Embeddings on Parallel Corpora

Zi-Yi Dou; Graham Neubig

doi:10.18653/v1/2021.eacl-main.181

Word Alignment by Fine-tuning Embeddings on Parallel Corpora

Abstract

Word alignment over parallel corpora has a wide variety of applications, including learning translation lexicons, cross-lingual transfer of language processing tools, and automatic evaluation or analysis of translation outputs. The great majority of past work on word alignment has worked by performing unsupervised learning on parallel text. Recently, however, other work has demonstrated that pre-trained contextualized word embeddings derived from multilingually trained language models (LMs) prove an attractive alternative, achieving competitive results on the word alignment task even in the absence of explicit training on parallel data. In this paper, we examine methods to marry the two approaches: leveraging pre-trained LMs but fine-tuning them on parallel text with objectives designed to improve alignment quality, and proposing methods to effectively extract alignments from these fine-tuned models. We perform experiments on five language pairs and demonstrate that our model can consistently outperform previous state-of-the-art models of all varieties. In addition, we demonstrate that we are able to train multilingual word aligners that can obtain robust performance on different language pairs.

Anthology ID:: 2021.eacl-main.181
Volume:: Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume
Month:: April
Year:: 2021
Address:: Online
Editors:: Paola Merlo, Jorg Tiedemann, Reut Tsarfaty
Venue:: EACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 2112–2128
Language:
URL:: https://aclanthology.org/2021.eacl-main.181/
DOI:: 10.18653/v1/2021.eacl-main.181
Bibkey:
Cite (ACL):: Zi-Yi Dou and Graham Neubig. 2021. Word Alignment by Fine-tuning Embeddings on Parallel Corpora. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2112–2128, Online. Association for Computational Linguistics.
Cite (Informal):: Word Alignment by Fine-tuning Embeddings on Parallel Corpora (Dou & Neubig, EACL 2021)
Copy Citation:
PDF:: https://aclanthology.org/2021.eacl-main.181.pdf
Code: neulab/awesome-align + additional community code

PDF Cite Search Code Fix data