Towards an Interoperable Corpus of Austrian Historical Newspapers: The case of PressMint-AT

Tanja Wissik, Jona Hassenbach, Hannes Pirker, Claudia Resch, Stefan Resch


Abstract
In this paper the PressMint-AT project is presented, which aims to create a historical newspaper corpus based on the Wiener Abendpost. The quality of automatic text recognition (ATR) is a key factor in creating historical newspaper corpora. Therefore, the performance of established ATR tools, multimodal large language models (LLMS), and existing full-text transcriptions provided by the Austrian National Library via ANNO is evaluated in order identify the most suitable approach for the PressMint-AT project. Even though recent research has demonstrated promising results for OCR tasks using multimodal LLMs, the experiments presented in this paper show, that PERO OCR achieves the best performance for the PressMint-AT dataset.
Anthology ID:
2026.pressmint-1.7
Volume:
Proceedings of the First Workshop on Creating Interoperable Corpora of Historical Newspapers
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Maciej Ogrodniczuk, Petya Osenova, Tanja Wissik
Venues:
PressMint | WS
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
34–39
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-pressmint-07
DOI:
10.63317/4ionpc6w2rmt
Bibkey:
Cite (ACL):
Tanja Wissik, Jona Hassenbach, Hannes Pirker, Claudia Resch, and Stefan Resch. 2026. Towards an Interoperable Corpus of Austrian Historical Newspapers: The case of PressMint-AT. In Proceedings of the First Workshop on Creating Interoperable Corpora of Historical Newspapers, pages 34–39, Palma de Mallorca, Spain. Association for Computational Linguistics.
Cite (Informal):
Towards an Interoperable Corpus of Austrian Historical Newspapers: The case of PressMint-AT (Wissik et al., PressMint 2026)
Copy Citation: