GerDraCor-Coref: A Coreference Corpus for Dramatic Texts in German

Janis Pagel, Nils Reiter


Abstract
Dramatic texts are a highly structured literary text type. Their quantitative analysis so far has relied on analysing structural properties (e.g., in the form of networks). Resolving coreferences is crucial for an analysis of the content of the character speech, but developing automatic coreference resolution (CR) systems depends on the existence of annotated corpora. In this paper, we present an annotated corpus of German dramatic texts, a preliminary analysis of the corpus as well as some baseline experiments on automatic CR. The analysis shows that with respect to the reference structure, dramatic texts are very different from news texts, but more similar to other dialogical text types such as interviews. Baseline experiments show a performance of 28.8 CoNLL score achieved by the rule-based CR system CorZu. In the future, we plan to integrate the (partial) information given in the dramatis personae into the CR model.
Anthology ID:
2020.lrec-1.7
Volume:
Proceedings of the Twelfth Language Resources and Evaluation Conference
Month:
May
Year:
2020
Address:
Marseille, France
Editors:
Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association
Note:
Pages:
55–64
Language:
English
URL:
https://aclanthology.org/2020.lrec-1.7
DOI:
Bibkey:
Cite (ACL):
Janis Pagel and Nils Reiter. 2020. GerDraCor-Coref: A Coreference Corpus for Dramatic Texts in German. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 55–64, Marseille, France. European Language Resources Association.
Cite (Informal):
GerDraCor-Coref: A Coreference Corpus for Dramatic Texts in German (Pagel & Reiter, LREC 2020)
Copy Citation:
PDF:
https://aclanthology.org/2020.lrec-1.7.pdf