Towards a Bulgarian Historical Newspaper Corpus – Construction of Reading Order over the Text in Searchable PDFs

Nikolay Paev, Stefan Marinov, Ivan Kratchanov, Petya Osenova, Kiril Simov


Abstract
The determine the reading order of the text extracted from a searchable PDF produced by an OCR software from an old newspaper is the first task in the process of preparation of corpora of old newspapers. In the paper we present an algorithm for generation of reading order of black selected from the corresponding PDF. Also we performed a tuning of the parameters of the algorithm. The optimization provides 10 % improvement.
Anthology ID:
2026.pressmint-1.10
Volume:
Proceedings of the First Workshop on Creating Interoperable Corpora of Historical Newspapers
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Maciej Ogrodniczuk, Petya Osenova, Tanja Wissik
Venues:
PressMint | WS
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
56–64
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-pressmint-10
DOI:
10.63317/336sd7ixv3fw
Bibkey:
Cite (ACL):
Nikolay Paev, Stefan Marinov, Ivan Kratchanov, Petya Osenova, and Kiril Simov. 2026. Towards a Bulgarian Historical Newspaper Corpus – Construction of Reading Order over the Text in Searchable PDFs. In Proceedings of the First Workshop on Creating Interoperable Corpora of Historical Newspapers, pages 56–64, Palma de Mallorca, Spain. Association for Computational Linguistics.
Cite (Informal):
Towards a Bulgarian Historical Newspaper Corpus – Construction of Reading Order over the Text in Searchable PDFs (Paev et al., PressMint 2026)
Copy Citation: