Alex Root


pdf bib
Williams College’s Submission for the Coco4MT 2023 Shared Task
Alex Root | Mark Hopkins
Proceedings of the Second Workshop on Corpus Generation and Corpus Augmentation for Machine Translation

Professional translation is expensive. As a consequence, when developing a translation system in the absence of a pre-existing parallel corpus, it is important to strategically choose sentences to have professionally translated for the training corpus. In our contribution to the Coco4MT 2023 Shared Task, we explore how sentence embeddings can be leveraged to choose an impactful set of sentences to translate. Based on six language pairs of the JHU Bible corpus, we demonstrate that a technique based on SimCSE embeddings outperforms a competitive suite of baselines.