Svetla Peneva Koeva
2026
Bulgarian Massive Multitask Language Understanding Benchmark
Svetla Peneva Koeva | Ivelina Stoyanova | Dimiter Georgiev | Svetlozara Leseva | Valentina Stefanova | Maria Todorova | Tsvetana Ivanova Dimitrova | Hristina Kukova | Mihaela Moskova | Tinko Tinchev
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Svetla Peneva Koeva | Ivelina Stoyanova | Dimiter Georgiev | Svetlozara Leseva | Valentina Stefanova | Maria Todorova | Tsvetana Ivanova Dimitrova | Hristina Kukova | Mihaela Moskova | Tinko Tinchev
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Assessing the broad general knowledge of Large Language Models (LLMs) across multiple domains in Bulgarian remains challenging due to the limited availability of Bulgarian evaluation benchmarks. To address this gap, we introduce the Bulgarian Massive Multitask Language Understanding benchmark (MMLU-BG), designed to evaluate whether LLMs possess generalised knowledge capabilities beyond simple text prediction in Bulgarian. This paper presents the structure, the development protocol, and the size of the MMLU-BG benchmark. It is tested in comparison with the original MMLU for English across seven LLMs selected according to specific criteria. The experiments demonstrate that the MMLU-BG benchmark assesses multi-domain versatility and highlights the models’ strengths and weaknesses across different subject areas.
A Large Dataset Representing Bulgarian, with the Bulgarian National Corpus as Its Core
Svetla Peneva Koeva | Ivelina Stoyanova
Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora
Svetla Peneva Koeva | Ivelina Stoyanova
Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora
The paper introduces the IfGPT dataset, which integrates several Bulgarian text collections, including the Bulgarian National Corpus, and applies cleaning, deduplication, and LLM-oriented metadata such as personally identifiable information and bias scores. The composition of the IfGPT dataset is presented, along with the unified metadata schema and metadata management in a graph database, enabling efficient querying and document selection for specific tasks. The main contributions are the integration of multiple Bulgarian text collections into a unified dataset, the development of a standardised metadata schema with graph-based organisation, and the provision of efficient metadata querying mechanisms to support LLM development.
Recent Developments of the Bulgarian National Corpus
Svetla Peneva Koeva | Ivelina Stoyanova
Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora
Svetla Peneva Koeva | Ivelina Stoyanova
Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora
We present recent developments in the Bulgarian National Corpus, including data collection from various sources, cleaning of diverse datasets, enrichment with multimodal data, and extensive metadata, which resulted in the development of IfGPT, a large BulNC-based dataset. Typical methods for distributing the BulNC-based dataset are briefly described, with emphasis on effective searching within the metadata stored in a graph database.
2025
IfGPT: A Dataset in Bulgarian for Large Language Models
Svetla Peneva Koeva | Ivelina Stoyanova | Jordan Konstantinov Kralev
Proceedings of the First Workshop on Advancing NLP for Low-Resource Languages
Svetla Peneva Koeva | Ivelina Stoyanova | Jordan Konstantinov Kralev
Proceedings of the First Workshop on Advancing NLP for Low-Resource Languages
The paper presents the large dataset IfGPT, which contains available corpora and datasets for Bulgarian, and describes methods to continuously expand it with unduplicated and unbiased Bulgarian data. The samples in the dataset are annotated with metadata that enable effective extraction of domain- and application-oriented datasets for fine-tuning or Retrieval Augmented Generation (RAG) of large language models (LLMs). The paper focuses on the description of the extended metadata of the IfGPT dataset and its management in a graph database.
2023
XL-WA: a Gold Evaluation Benchmark for Word Alignment in 14 Language Pairs
Federico Martelli | Andrei Stefan Bejgu | Cesare Campagnano | Jaka Čibej | Rute Costa | Apolonija Gantar | Jelena Kallas | Svetla Peneva Koeva | Kristina Koppel | Simon Krek | Margit Langemets | Veronika Lipp | Sanni Nimb | Sussi Olsen | Bolette Sanford Pedersen | Valeria Quochi | Ana Salgado | László Simon | Carole Tiberius | Rafael-J Ureña-Ruiz | Roberto Navigli
Proceedings of the Ninth Italian Conference on Computational Linguistics (CLiC-it 2023)
Federico Martelli | Andrei Stefan Bejgu | Cesare Campagnano | Jaka Čibej | Rute Costa | Apolonija Gantar | Jelena Kallas | Svetla Peneva Koeva | Kristina Koppel | Simon Krek | Margit Langemets | Veronika Lipp | Sanni Nimb | Sussi Olsen | Bolette Sanford Pedersen | Valeria Quochi | Ana Salgado | László Simon | Carole Tiberius | Rafael-J Ureña-Ruiz | Roberto Navigli
Proceedings of the Ninth Italian Conference on Computational Linguistics (CLiC-it 2023)
Search
Fix author
Co-authors
- Ivelina Stoyanova 4
- Andrei Stefan Bejgu 1
- Cesare Campagnano 1
- Rute Costa 1
- Tsvetana Ivanova Dimitrova 1
- Apolonija Gantar 1
- Dimiter Georgiev 1
- Jelena Kallas 1
- Kristina Koppel 1
- Jordan Konstantinov Kralev 1
- Simon Krek 1
- Hristina Kukova 1
- Margit Langemets 1
- Svetlozara Leseva 1
- Veronika Lipp 1
- Federico Martelli 1
- Mihaela Moskova 1
- Roberto Navigli 1
- Sanni Nimb 1
- Sussi Olsen 1
- Valeria Quochi 1
- Ana Salgado 1
- Bolette Sanford Pedersen 1
- László Simon 1
- Valentina Stefanova 1
- Carole Tiberius 1
- Tinko Tinchev 1
- Maria Todorova 1
- Rafael-J. Ureña-Ruiz 1
- Jaka Čibej 1