Using Songs to Improve Kazakh Automatic Speech Recognition

Rustem Yeshpanov


Abstract
Developing automatic speech recognition (ASR) systems for low-resource languages is hindered by the scarcity of transcribed corpora. This proof-of-concept study explores songs as an unconventional yet promising data source for Kazakh ASR. We curate a dataset of 3,013 audio-text pairs (about 4.5 hours) from 195 songs by 36 artists, segmented at the lyric-line level. Using Whisper as the base recogniser, we fine-tune models under seven training scenarios involving Songs, Common Voice Corpus (CVC), and FLEURS, and evaluate them on three benchmarks: CVC, FLEURS, and Kazakh Speech Corpus 2 (KSC2). Results show that song-based fine-tuning improves performance over zero-shot baselines. For instance, Whisper Large-V3 Turbo trained on a mixture of Songs, CVC, and FLEURS achieves 27.6% normalised WER on CVC and 11.8% on FLEURS, while halving the error on KSC2 (39.3% vs. 81.2%) relative to the zero-shot model. Although these gains remain below those of models trained on the 1,100-hour KSC2 corpus, they demonstrate that even modest song-speech mixtures can yield meaningful adaptation improvements in low-resource ASR. The dataset is released on Hugging Face for research purposes under a gated, non-commercial licence.
Anthology ID:
2026.lrec-1.431
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
5528–5537
Language:
External URL:
https://lrec.elra.info/lrec2026-main-431
DOI:
10.63317/5hqonmz5roum
Bibkey:
Cite (ACL):
Rustem Yeshpanov. 2026. Using Songs to Improve Kazakh Automatic Speech Recognition. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 5528–5537, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Using Songs to Improve Kazakh Automatic Speech Recognition (Yeshpanov, LREC 2026)
Copy Citation: