WAGS: A Beautiful English-Italian Benchmark Supporting Word Alignment Evaluation on Rare Words

Luisa Bentivogli, Mauro Cettolo, M. Amin Farajian, Marcello Federico


Abstract
This paper presents WAGS (Word Alignment Gold Standard), a novel benchmark which allows extensive evaluation of WA tools on out-of-vocabulary (OOV) and rare words. WAGS is a subset of the Common Test section of the Europarl English-Italian parallel corpus, and is specifically tailored to OOV and rare words. WAGS is composed of 6,715 sentence pairs containing 11,958 occurrences of OOV and rare words up to frequency 15 in the Europarl Training set (5,080 English words and 6,878 Italian words), representing almost 3% of the whole text. Since WAGS is focused on OOV/rare words, manual alignments are provided for these words only, and not for the whole sentences. Two off-the-shelf word aligners have been evaluated on WAGS, and results have been compared to those obtained on an existing benchmark tailored to full text alignment. The results obtained confirm that WAGS is a valuable resource, which allows a statistically sound evaluation of WA systems’ performance on OOV and rare words, as well as extensive data analyses. WAGS is publicly released under a Creative Commons Attribution license.
Anthology ID:
L16-1562
Volume:
Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16)
Month:
May
Year:
2016
Address:
Portorož, Slovenia
Editors:
Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Sara Goggi, Marko Grobelnik, Bente Maegaard, Joseph Mariani, Helene Mazo, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association (ELRA)
Note:
Pages:
3535–3542
Language:
URL:
https://aclanthology.org/L16-1562
DOI:
Bibkey:
Cite (ACL):
Luisa Bentivogli, Mauro Cettolo, M. Amin Farajian, and Marcello Federico. 2016. WAGS: A Beautiful English-Italian Benchmark Supporting Word Alignment Evaluation on Rare Words. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16), pages 3535–3542, Portorož, Slovenia. European Language Resources Association (ELRA).
Cite (Informal):
WAGS: A Beautiful English-Italian Benchmark Supporting Word Alignment Evaluation on Rare Words (Bentivogli et al., LREC 2016)
Copy Citation:
PDF:
https://aclanthology.org/L16-1562.pdf
Data
Europarl