Arden Zimmermann
2026
Text+: A National Hub Including Legacy Language Data
Florian Barth | Christoph Draxler | Jennifer Ecker | Stefan Fischer | Philippe Genêt | Alina Hemmer | Timm Lehmberg | Thorsten Trippel | Andreas Witt | Arden Zimmermann | Claus Zinn
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Florian Barth | Christoph Draxler | Jennifer Ecker | Stefan Fischer | Philippe Genêt | Alina Hemmer | Timm Lehmberg | Thorsten Trippel | Andreas Witt | Arden Zimmermann | Claus Zinn
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Text+ is the German distributed research data infrastructure for literary studies, linguistics, and spoken and written language. Its resources consist of contemporary and historical literary and media texts, deeply annotated material, transcripts of spoken and sign language, and original recordings. Text+ provides access to its resources according to the FAIR guidelines: Findable due to standard-conformant metadata, Accessible with single sign-on authentication, Interoperable via open data formats, and Reproducible through web services and extensive documentation. The 30+ partners of Text+ are archives, libraries, universities, and other research institutions. The partners are autonomous, and they differ in the amount of data and processing capabilities they provide. In this paper, we describe the hub architecture of Text+, which gives users a central and FAIR point of access to research data that continues to be distributed across the Text+ partner institutions. The architecture serves as a blueprint to evolving research infrastructures that aim at maintaining (and empowering) their research data contributors.
DIN 19461: A National Standard for Derived Text Formats
Thorsten Trippel | Florian Barth | Jose Calvo Tello | Keli Du | Philippe Genêt | Daniel Kurzawe | Peter Leinen | Piroska Lendvai | Christof Schöch | Andreas Witt | Arden Zimmermann
Proceedings of Leveraging Derived Text Formats to Unlock Copyrighted Collections for Open Science (DTF) @ LREC 2026
Thorsten Trippel | Florian Barth | Jose Calvo Tello | Keli Du | Philippe Genêt | Daniel Kurzawe | Peter Leinen | Piroska Lendvai | Christof Schöch | Andreas Witt | Arden Zimmermann
Proceedings of Leveraging Derived Text Formats to Unlock Copyrighted Collections for Open Science (DTF) @ LREC 2026
We present DIN 19461:2026-06 (E), a German draft national standard that defines categories, terminology, and process requirements for Derived Text Formats (DTFs) created from text documents in natural language. The standard specifies enrichment and information reduction operations, requirements for combining multiple DTFs, and documentation obligations for publication, archiving, and reuse. Its aim is to enable legally compliant sharing and analysis of texts–especially where copyright or data protection prevents distributing originals–while maintaining scientific utility and reproducibility through explicit process and parameter recording. We outline the scope, the key concepts, the four core reduction operations (retain, delete, replace, randomise), together with examples across token-, structure-, and vector-based DTFs, and implications for infrastructures (e.g., ISO 24622-based metadata). Finally, we discuss limitations, open questions (e.g., reconstruction risks with modern ML models), and next steps for adoption and maintenance.