Southern Kurdish Speech Recognition Resources and Benchmarking

Mohammad Mohammadamini, Marie Tahon


Abstract
This article introduces a dedicated speech recognition dataset for Southern Kurdish, which is a threatened variant of Kurdish macrolanguage. We present 30 hours of validated read speech for training and an evaluation benchmark for Southern Kurdish Automatic Speech Recognition (ASR). Both the training data and evaluation benchmark are read speech recorded by crowdsourcing campaigns. Besides a detailed description of the provided resources, we provide the ASR baselines using Whisper-turbo and wav2vec-bert CTC architectures. We achieved a 4.09 CER and 24.26 WER on our benchmark using wav2vec-bert model. We also provide a categorization of errors to support further improvements in future studies.The resources and trained models are released under the CC BY-NC-ND 4.0 license and are publicly available at https://huggingface.co/datasets/aranemini/southern-kurdish-asr
Anthology ID:
2026.lrec-1.432
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
5538–5544
Language:
External URL:
https://lrec.elra.info/lrec2026-main-432
DOI:
10.63317/2rkqhw7hmo2d
Bibkey:
Cite (ACL):
Mohammad Mohammadamini and Marie Tahon. 2026. Southern Kurdish Speech Recognition Resources and Benchmarking. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 5538–5544, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Southern Kurdish Speech Recognition Resources and Benchmarking (Mohammadamini & Tahon, LREC 2026)
Copy Citation: