Harald Lüngen
Also published as: Harald Lungen
2026
EuReCo, KorAP and DeReKo: Updates on Ingestion and Annotation Pipelines, Backend, Interfaces, Operation, and Corpora
Marc Kupietz | Nils Diewald | Harald Lüngen | Eliza Margaretha Illig | Helge Stallkamp | Uyen-Nhu Tran | Rameela Yaddehige
Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora
Marc Kupietz | Nils Diewald | Harald Lüngen | Eliza Margaretha Illig | Helge Stallkamp | Uyen-Nhu Tran | Rameela Yaddehige
Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora
This paper reports on recent technical developments in the European Reference Corpus EuReCo and its current technical implementation based on the corpus search and analysis platform KorAP. We describe updates to the ingestion pipeline, including extensions to the TEI-to-KorAP-XML converter tei2korapxml and the KorAP tokenizer, as well as the newly introduced korapxmltool for annotation and index conversion. We further present Koral-Mapper, a service that enables cross-schema comparability of annotations and metadata at query time, and report on developments in the backend access control system Kustvakt, the web user interface Kalamar, API client libraries for R and Python that promote reproducibility and methodologically sound AI-assisted analysis, and containerized deployment. The corpora and languages currently represented in EuReCo are outlined, and the role of the German Reference Corpus DeReKo, including its metadata-driven virtual corpus design, predefined useful subcorpora, and TEI encoding, is discussed in detail. We further present the National Libraries as Corpus approach and DeLiKo-2025@DNB as its first full-scale proof of concept, and discuss the potential of this approach for extending EuReCo with comparable contemporary fiction corpora across European countries.
2022
Proceedings of the Workshop on Challenges in the Management of Large Corpora (CMLC-10)
Piotr Banski | Adrien Barbaresi | Simon Clematide | Marc Kupietz | Harald Lüngen
Proceedings of the Workshop on Challenges in the Management of Large Corpora (CMLC-10)
Piotr Banski | Adrien Barbaresi | Simon Clematide | Marc Kupietz | Harald Lüngen
Proceedings of the Workshop on Challenges in the Management of Large Corpora (CMLC-10)
2020
Proceedings of the 8th Workshop on Challenges in the Management of Large Corpora
Piotr Bański | Adrien Barbaresi | Simon Clematide | Marc Kupietz | Harald Lüngen | Ines Pisetta
Proceedings of the 8th Workshop on Challenges in the Management of Large Corpora
Piotr Bański | Adrien Barbaresi | Simon Clematide | Marc Kupietz | Harald Lüngen | Ines Pisetta
Proceedings of the 8th Workshop on Challenges in the Management of Large Corpora
2018
The German Reference Corpus DeReKo: New Developments – New Opportunities
Marc Kupietz | Harald Lüngen | Paweł Kamocki | Andreas Witt
Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018)
Marc Kupietz | Harald Lüngen | Paweł Kamocki | Andreas Witt
Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018)
2014
Recent Developments in DeReKo
Marc Kupietz | Harald Lüngen
Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14)
Marc Kupietz | Harald Lüngen
Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14)
This paper gives an overview of recent developments in the German Reference Corpus DeReKo in terms of growth, maximising relevant corpus strata, metadata, legal issues, and its current and future research interface. Due to the recent acquisition of new licenses, DeReKo has grown by a factor of four in the first half of 2014, mostly in the area of newspaper text, and presently contains over 24 billion word tokens. Other strata, like fictional texts, web corpora, in particular CMC texts, and spoken but conceptually written texts have also increased significantly. We report on the newly acquired corpora that led to the major increase, on the principles and strategies behind our corpus acquisition activities, and on our solutions for the emerging legal, organisational, and technical challenges.
2004
Text Type Structure and Logical Document Structure
Hagen Langer | Harald Lungen | Petra Saskia Bayerl
Proceedings of the Workshop on Discourse Annotation
Hagen Langer | Harald Lungen | Petra Saskia Bayerl
Proceedings of the Workshop on Discourse Annotation