Recent Developments of the Bulgarian National Corpus

Svetla Peneva Koeva, Ivelina Stoyanova


Abstract
We present recent developments in the Bulgarian National Corpus, including data collection from various sources, cleaning of diverse datasets, enrichment with multimodal data, and extensive metadata, which resulted in the development of IfGPT, a large BulNC-based dataset. Typical methods for distributing the BulNC-based dataset are briefly described, with emphasis on effective searching within the metadata stored in a graph database.
Anthology ID:
2026.cmlc-1.10
Volume:
Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Piotr Bański, Dawn Knight, Marc Kupietz, Andreas Witt, Alina Wróblewska
Venues:
CMLC | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
71–75
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-cmlc-10
DOI:
10.63317/3m95ohtw7mjs
Bibkey:
Cite (ACL):
Svetla Peneva Koeva and Ivelina Stoyanova. 2026. Recent Developments of the Bulgarian National Corpus. In Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora, pages 71–75, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
Recent Developments of the Bulgarian National Corpus (Koeva & Stoyanova, CMLC 2026)
Copy Citation: