Missed opportunities in translation memory matching

Friedel Wolff; Laurette Pretorius; Paul Buitelaar

Missed opportunities in translation memory matching

Friedel Wolff, Laurette Pretorius, Paul Buitelaar

Abstract

A translation memory system stores a data set of source-target pairs of translations. It attempts to respond to a query in the source language with a useful target text from the data set to assist a human translator. Such systems estimate the usefulness of a target text suggestion according to the similarity of its associated source text to the source text query. This study analyses two data sets in two language pairs each to find highly similar target texts, which would be useful mutual suggestions. We further investigate which of these useful suggestions can not be selected through source text similarity, and we do a thorough analysis of these cases to categorise and quantify them. This analysis provides insight into areas where the recall of translation memory systems can be improved. Specifically, source texts with an omission, and semantically very similar source texts are some of the more frequent cases with useful target text suggestions that are not selected with the baseline approach of simple edit distance between the source texts.

Anthology ID:: L14-1044
Volume:: Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14)
Month:: May
Year:: 2014
Address:: Reykjavik, Iceland
Editors:: Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Hrafn Loftsson, Bente Maegaard, Joseph Mariani, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:: LREC
SIG:
Publisher:: European Language Resources Association (ELRA)
Note:
Pages:: 4401–4406
Language:
URL:: http://www.lrec-conf.org/proceedings/lrec2014/pdf/1061_Paper.pdf
DOI:
Bibkey:
Cite (ACL):: Friedel Wolff, Laurette Pretorius, and Paul Buitelaar. 2014. Missed opportunities in translation memory matching. In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14), pages 4401–4406, Reykjavik, Iceland. European Language Resources Association (ELRA).
Cite (Informal):: Missed opportunities in translation memory matching (Wolff et al., LREC 2014)
Copy Citation:
PDF:: http://www.lrec-conf.org/proceedings/lrec2014/pdf/1061_Paper.pdf

PDF Cite Search Fix data