Cora Haiber
2026
Domain-Specific Considerations in the Preparation of Specialized Corpora: A Case Study on a Corpus of German Sermons
Cora Haiber | Adam Roussel | Stefanie Dipper
Proceedings of the Workshop on Structured Linguistic Data and Evaluation (SLiDE)
Cora Haiber | Adam Roussel | Stefanie Dipper
Proceedings of the Workshop on Structured Linguistic Data and Evaluation (SLiDE)
We present a new corpus of contemporary German sermons and describe the steps taken in its preparation. We apply a semi-automatic approach to sentence segmentation, tokenization, and lemmatization, utilizing annotation guidelines that are specialized to this domain. In the process of preparing these data, we find that state-of-the-art tools for these tasks still make problematic errors, especially with non-standard data, despite apparently very high performance on common benchmarks. We obtain test scores of F1 = 96.69 % for sentence segmentation, F1 = 99.99 % for tokenization, and acc = 64.00 % for lemmatization with our domain-adapted models and show that domain-adaptation improves performance over state-of-the-art models for the token and sentence segmentation tasks.
2024
Universal Dependencies: Extensions for Modern and Historical German
Stefanie Dipper | Cora Haiber | Anna Maria Schröter | Alexandra Wiemann | Maike Brinkschulte
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
Stefanie Dipper | Cora Haiber | Anna Maria Schröter | Alexandra Wiemann | Maike Brinkschulte
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
In this paper we present extensions of the UD scheme for modern and historical German. The extensions relate in part to fundamental differences such as those between different kinds of arguments and modifiers. We illustrate the extensions with examples from the MHG data and discuss a number of MHG-specific constructions. At the current time, we have annotated a corpus of Middle High German with almost 29K tokens using this scheme, which to our knowledge is the first UD treebank for Middle High German. Inter-annotator agreement is very high: the annotators achieve a score of α = 0.85. A statistical analysis of the annotations shows some interesting differences in the distribution of labels between modern and historical German.