Addressing Cha(lle)nges in Long-Term Archiving of Large Corpora

Denis Arnold, Bernhard Fisseni, Pawel Kamocki, Oliver Schonefeld, Marc Kupietz, Thomas Schmidt


Abstract
This paper addresses long-term archival for large corpora. Three aspects specific to language resources are focused, namely (1) the removal of resources for legal reasons, (2) versioning of (unchanged) objects in constantly growing resources, especially where objects can be part of multiple releases but also part of different collections, and (3) the conversion of data to new formats for digital preservation. It is motivated why language resources may have to be changed, and why formats may need to be converted. As a solution, the use of an intermediate proxy object called a signpost is suggested. The approach will be exemplified with respect to the corpora of the Leibniz Institute for the German Language in Mannheim, namely the German Reference Corpus (DeReKo) and the Archive for Spoken German (AGD).
Anthology ID:
2020.cmlc-1.1
Volume:
Proceedings of the 8th Workshop on Challenges in the Management of Large Corpora
Month:
May
Year:
2020
Address:
Marseille, France
Editors:
Piotr Bański, Adrien Barbaresi, Simon Clematide, Marc Kupietz, Harald Lüngen, Ines Pisetta
Venue:
CMLC
SIG:
Publisher:
European Language Ressources Association
Note:
Pages:
1–9
Language:
English
URL:
https://aclanthology.org/2020.cmlc-1.1
DOI:
Bibkey:
Cite (ACL):
Denis Arnold, Bernhard Fisseni, Pawel Kamocki, Oliver Schonefeld, Marc Kupietz, and Thomas Schmidt. 2020. Addressing Cha(lle)nges in Long-Term Archiving of Large Corpora. In Proceedings of the 8th Workshop on Challenges in the Management of Large Corpora, pages 1–9, Marseille, France. European Language Ressources Association.
Cite (Informal):
Addressing Cha(lle)nges in Long-Term Archiving of Large Corpora (Arnold et al., CMLC 2020)
Copy Citation:
PDF:
https://aclanthology.org/2020.cmlc-1.1.pdf