Lexical Substitution Dataset for German

Kostadin Cholakov, Chris Biemann, Judith Eckle-Kohler, Iryna Gurevych


Abstract
This article describes a lexical substitution dataset for German. The whole dataset contains 2,040 sentences from the German Wikipedia, with one target word in each sentence. There are 51 target nouns, 51 adjectives, and 51 verbs randomly selected from 3 frequency groups based on the lemma frequency list of the German WaCKy corpus. 200 sentences have been annotated by 4 professional annotators and the remaining sentences by 1 professional annotator and 5 additional annotators who have been recruited via crowdsourcing. The resulting dataset can be used to evaluate not only lexical substitution systems, but also different sense inventories and word sense disambiguation systems.
Anthology ID:
L14-1450
Volume:
Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14)
Month:
May
Year:
2014
Address:
Reykjavik, Iceland
Editors:
Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Hrafn Loftsson, Bente Maegaard, Joseph Mariani, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association (ELRA)
Note:
Pages:
1406–1411
Language:
URL:
http://www.lrec-conf.org/proceedings/lrec2014/pdf/545_Paper.pdf
DOI:
Bibkey:
Cite (ACL):
Kostadin Cholakov, Chris Biemann, Judith Eckle-Kohler, and Iryna Gurevych. 2014. Lexical Substitution Dataset for German. In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14), pages 1406–1411, Reykjavik, Iceland. European Language Resources Association (ELRA).
Cite (Informal):
Lexical Substitution Dataset for German (Cholakov et al., LREC 2014)
Copy Citation:
PDF:
http://www.lrec-conf.org/proceedings/lrec2014/pdf/545_Paper.pdf