Data Matters: Looking for High-Quality Corpora to Build Robust and Reliable Models for Humanists

Jaione Macicior-Mitxelena, Ana García-Serrano


Abstract
The digitization of Spanish historical newspapers poses significant challenges due to low scan quality, typographical diversity, complex layouts and linguistic variation from contemporary Spanish. While advances in Optical Character Recognition (OCR) and layout-aware models offer promising results, their effectiveness strongly depends on the quality and consistency of the underlying training corpora. This work focuses on corpus construction and evaluation for historical document processing. Two experiments were conducted. In the first corpus los101 was used, a manually curated and structurally annotated subcorpus derived from historical Spanish newspapers, designed to ensure coherent ground truth under heterogeneous real-world conditions. This corpus enables systematic experimentation across OCR and document layout analysis tasks. In a second experimental phase, we apply an additional layout-focused corpus characterized by structural regularity and consistent page organization, allowing us to isolate the impact of layout homogeneity on segmentation performance. State-of-the-art OCR models and a layout detection model are evaluated as validation instruments to assess corpus adequacy rather than as primary contributions. Quantitative and qualitative analyses based on (1) relationship between annotation quality, (2) structural variability, and (3) model behavior, show that heterogeneous corpora challenge both transcription and segmentation stability, while layout-consistent data significantly improves structural detection reliability.
Anthology ID:
2026.pressmint-1.13
Volume:
Proceedings of the First Workshop on Creating Interoperable Corpora of Historical Newspapers
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Maciej Ogrodniczuk, Petya Osenova, Tanja Wissik
Venues:
PressMint | WS
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
82–91
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-pressmint-13
DOI:
10.63317/28atsf4aiory
Bibkey:
Cite (ACL):
Jaione Macicior-Mitxelena and Ana García-Serrano. 2026. Data Matters: Looking for High-Quality Corpora to Build Robust and Reliable Models for Humanists. In Proceedings of the First Workshop on Creating Interoperable Corpora of Historical Newspapers, pages 82–91, Palma de Mallorca, Spain. Association for Computational Linguistics.
Cite (Informal):
Data Matters: Looking for High-Quality Corpora to Build Robust and Reliable Models for Humanists (Macicior-Mitxelena & García-Serrano, PressMint 2026)
Copy Citation: