PressMint-PT - Compiling a Portuguese Historical Newspaper Corpus

Jose Aires, Amália Mendes


Abstract
We present a new European Portuguese corpus of newspapers from the 19th and early 20th centuries, integrated in the recent PressMint project, whose goal is to provide a set of comparable newspaper corpora for European languages in that time frame. We discuss the raw data that was previously available, as well as new data specifically compiled for the project, and the challenges involving OCR, text recognition, and different orthographical norms. We describe the pipeline setup for XML encoding and annotation, partially based on work developed for the ParlaMint corpora. The corpus is currently under development and will be made freely available at the end of the project, as part of the PressMint corpora.
Anthology ID:
2026.pressmint-1.4
Volume:
Proceedings of the First Workshop on Creating Interoperable Corpora of Historical Newspapers
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Maciej Ogrodniczuk, Petya Osenova, Tanja Wissik
Venues:
PressMint | WS
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
16–20
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-pressmint-04
DOI:
10.63317/2os9smose6kj
Bibkey:
Cite (ACL):
Jose Aires and Amália Mendes. 2026. PressMint-PT - Compiling a Portuguese Historical Newspaper Corpus. In Proceedings of the First Workshop on Creating Interoperable Corpora of Historical Newspapers, pages 16–20, Palma de Mallorca, Spain. Association for Computational Linguistics.
Cite (Informal):
PressMint-PT - Compiling a Portuguese Historical Newspaper Corpus (Aires & Mendes, PressMint 2026)
Copy Citation: