Investigating the Role of Synthetic Data Augmentation and Training Strategies on Improving Low-Resource Language ASR

Yun Hao, Reihaneh Amooie, Wietse de Vries, Rik van Noord, Martijn Wieling


Abstract
Low-resource automatic speech recognition (ASR) is challenging due to a scarcity of annotated data. While synthetic data from text-to-speech (TTS) systems can augment ASR training, its efficacy for low-resource languages remains unclear. In this study, we investigate under which conditions TTS-based data augmentation is most effective for low-resource languages. Experiments on six low-resource languages in Common Voice show that synthetic data is most beneficial under extremely low-resource ASR conditions (i.e., less than one hour of available real speech data), or for languages with larger amounts of TTS data (i.e., more than 10 hours). Additionally, increasing the amount and diversity of synthetic data while keeping an appropriate ratio of synthetic-to-real data can further improve ASR performance.
Anthology ID:
2026.lrec-1.442
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
5633–5639
Language:
External URL:
https://lrec.elra.info/lrec2026-main-442
DOI:
10.63317/4c4ad7t5i967
Bibkey:
Cite (ACL):
Yun Hao, Reihaneh Amooie, Wietse de Vries, Rik van Noord, and Martijn Wieling. 2026. Investigating the Role of Synthetic Data Augmentation and Training Strategies on Improving Low-Resource Language ASR. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 5633–5639, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Investigating the Role of Synthetic Data Augmentation and Training Strategies on Improving Low-Resource Language ASR (Hao et al., LREC 2026)
Copy Citation: