Building the v4 of the Croatian National Corpus

Marko Tadić, Vanja Štefanec, Daša Farkaš


Abstract
It has been thirteen years since the release of the current version (v3) of the Croatian National Corpus (HNK). In terms of synchronicity in corpus linguistics, that many years may be considered quite some time. The preparatory phase for the composition of the new version of HNK (v4) has been going already for several years and in this paper we touch on several issues of concern. Apart of regular corpus parameters, e.g. text sources, text genres, coverage of language varieties, time span, we also discuss about metadata and linguistic annotation schemata. One of important technical prerequisites was the development of CorpRepo, a custom corpus data management system and file system, which enable us to do sustainable long-term maintenance of the data, and to produce newer versions of corpus more easily and more often. The selection of IPR-cleared data entails some restrictions and we give several examples of that kind of textual sources, but also discuss possible weaknesses of such approach to data selection. Regarding the linguistic annotation, the important shift is the decision to abandon the MulText East morphosyntacting descriptions and use solutions recommended by UD-initiative.
Anthology ID:
2026.cmlc-1.13
Volume:
Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Piotr Bański, Dawn Knight, Marc Kupietz, Andreas Witt, Alina Wróblewska
Venues:
CMLC | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
80–83
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-cmlc-13
DOI:
10.63317/48si6yrozisf
Bibkey:
Cite (ACL):
Marko Tadić, Vanja Štefanec, and Daša Farkaš. 2026. Building the v4 of the Croatian National Corpus. In Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora, pages 80–83, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
Building the v4 of the Croatian National Corpus (Tadić et al., CMLC 2026)
Copy Citation: