Śmigiel Dataset: Laying Foundations for Investigating Machine-Generated Text Detection in Polish

Jakub Strebeyko, Alina Wróblewska, Piotr Przybyła


Abstract
We present Śmigiel, the first open dataset for training and evaluating machine-generated text (MGT) in Polish. The dataset includes a collection of human-written text fragments from six domains, which are used to prompt text generation by eight language models capable of producing credible Polish text. In addition to the raw corpus of over 462K generated texts, we also release a cleaned source- and domain-balanced dataset suitable for training and evaluating MGT detectors. Finally, we conduct preliminary experiments with text classifiers, showing that task difficulty depends on the text domain, the generating language model, and the availability of similar data in training. The results indicate that MGT detection in Polish can be approached with general-purpose classifiers that generalize well to new LLMs, but struggle to adapt to genres not represented in the training data.
Anthology ID:
2026.lrec-1.828
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
10556–10568
Language:
External URL:
https://lrec.elra.info/lrec2026-main-828
DOI:
10.63317/3p7ghe9pfm8v
Bibkey:
Cite (ACL):
Jakub Strebeyko, Alina Wróblewska, and Piotr Przybyła. 2026. Śmigiel Dataset: Laying Foundations for Investigating Machine-Generated Text Detection in Polish. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 10556–10568, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Śmigiel Dataset: Laying Foundations for Investigating Machine-Generated Text Detection in Polish (Strebeyko et al., LREC 2026)
Copy Citation: