Creating Task-Specific Speech Recognition Datasets from Scratch for Low-Resource Languages: Assessing the Impact of Token Sequence Overlap

Adwoa Asantewaa Bremang, Dennis Asamoah Owusu, Victor Kow Quagraine, Leanne M.M. Annor-Adjaye


Abstract
Creating a task-specific speech recognition dataset is essential for developing speech recognition applications in low-resource languages. Such applications have uses in agriculture, finance, healthcare, and others, and benefit individuals with low literacy. However, a significant challenge is the high cost of data creation. While there is some work around cost-effective dataset selection, there is little to no work on building a cost-effective dataset for a task from scratch. Our work contributes to the latter. We created a speech recognition dataset from scratch and conducted two major sets of experiments. The first aimed to observe the effect of different datasets of the same size on model performance. Our results confirmed that the same amount spent collecting data can have vastly different results. The second experiment analyzed the effect of token sequence overlap between target and training data since a natural and intuitive approach to building a dataset from scratch for task would be having the task tokens occur in the training data. Our experiments showed that token sequence overlap was not the primary factor influencing model performance. Our work provides a counter-intuitive insight into building speech recognition datasets from scratch in low-resource settings and shows the need for further investigation.
Anthology ID:
2026.lrec-1.240
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
3075–3082
Language:
External URL:
https://lrec.elra.info/lrec2026-main-240
DOI:
10.63317/3myb33sgskfb
Bibkey:
Cite (ACL):
Adwoa Asantewaa Bremang, Dennis Asamoah Owusu, Victor Kow Quagraine, and Leanne M.M. Annor-Adjaye. 2026. Creating Task-Specific Speech Recognition Datasets from Scratch for Low-Resource Languages: Assessing the Impact of Token Sequence Overlap. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 3075–3082, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Creating Task-Specific Speech Recognition Datasets from Scratch for Low-Resource Languages: Assessing the Impact of Token Sequence Overlap (Bremang et al., LREC 2026)
Copy Citation: