A Dialectal Corpus for Ukrainian: Collection, Classification, and Standardization

Yuliia Frund, Sina Ahmadi


Abstract
Ukrainian dialects remain largely excluded from the digital linguistic landscape despite their active everyday use. We present a regional dialect corpus covering 18 administrative regions of Ukraine, compiled from digitized fieldwork collections and an online dialect atlas. The corpus comprises over 284,000 tokens of dialect text, annotated by region and partially accompanied by manually standardized translations. Using these resources, we investigate language identification and dialect-to-standard standardization. Baseline language identification yields an F-score of 0.75, rising to 0.99 with dialect-inclusive training. Dialect classification reaches 0.58, with confusion patterns reflecting known regional boundaries. For standardization, the best-performing LLM achieves a COMET score of 0.80, though BLEU scores remain low (0.21–0.23) across all models. We release the corpus, labelled datasets, model outputs, and reference translations to support future work on inclusive language technologies for non-standard varieties.
Anthology ID:
2026.dialres-1.14
Volume:
Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective
Month:
May
Year:
2026
Address:
Palma de Mallorca
Editors:
Antonis Anastasopoulos, Stella Markantonatou, Angela Ralli, Marcos Zampieri, Stavros Bompolas, Vivian Stamou
Venues:
DialRes | WS
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
135–143
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-dialres-14
DOI:
10.63317/4mkaru7y2op5
Bibkey:
Cite (ACL):
Yuliia Frund and Sina Ahmadi. 2026. A Dialectal Corpus for Ukrainian: Collection, Classification, and Standardization. In Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective, pages 135–143, Palma de Mallorca. Association for Computational Linguistics.
Cite (Informal):
A Dialectal Corpus for Ukrainian: Collection, Classification, and Standardization (Frund & Ahmadi, DialRes 2026)
Copy Citation: