A Large Dataset Representing Bulgarian, with the Bulgarian National Corpus as Its Core

Svetla Peneva Koeva, Ivelina Stoyanova


Abstract
The paper introduces the IfGPT dataset, which integrates several Bulgarian text collections, including the Bulgarian National Corpus, and applies cleaning, deduplication, and LLM-oriented metadata such as personally identifiable information and bias scores. The composition of the IfGPT dataset is presented, along with the unified metadata schema and metadata management in a graph database, enabling efficient querying and document selection for specific tasks. The main contributions are the integration of multiple Bulgarian text collections into a unified dataset, the development of a standardised metadata schema with graph-based organisation, and the provision of efficient metadata querying mechanisms to support LLM development.
Anthology ID:
2026.cmlc-1.2
Volume:
Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Piotr Bański, Dawn Knight, Marc Kupietz, Andreas Witt, Alina Wróblewska
Venues:
CMLC | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
12–24
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-cmlc-02
DOI:
10.63317/4gxqpsnk5jz5
Bibkey:
Cite (ACL):
Svetla Peneva Koeva and Ivelina Stoyanova. 2026. A Large Dataset Representing Bulgarian, with the Bulgarian National Corpus as Its Core. In Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora, pages 12–24, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
A Large Dataset Representing Bulgarian, with the Bulgarian National Corpus as Its Core (Koeva & Stoyanova, CMLC 2026)
Copy Citation: