Reihaneh Amooie
2026
Improving Low-resource ASR Using Bilingual Fine-tuning with Language Identification: A Cross-linguistic Evaluation
Reihaneh Amooie | Yun Hao | Wietse de Vries | Jelske Dijkstra | Matt Coler | Martijn Wieling
Proceedings of Speech Language Models in Low-Resource Settings: Performance, Evaluation, and Bias Analysis (SPEAKABLE) @ LREC 2026
Reihaneh Amooie | Yun Hao | Wietse de Vries | Jelske Dijkstra | Matt Coler | Martijn Wieling
Proceedings of Speech Language Models in Low-Resource Settings: Performance, Evaluation, and Bias Analysis (SPEAKABLE) @ LREC 2026
This study explores how bilingual fine-tuning affects automatic speech recognition (ASR) in low-resource languages. We evaluate this method across nine linguistically and geographically diverse language pairs, covering a range of language families and writing systems. To distinguish the two languages, during training, we pre-pend each input text with a language identification token. At inference, the model jointly predicts both the language and transcription from the speech input alone. As texts for which the language is incorrectly determined show low ASR performance, we also conduct a follow-up experiment in which the language identification token is provided both during training and inference. Our results show that bilingual fine-tuning can be beneficial when language identification accuracy is high, and that in cases where language identification performance is low, including the language identification token at inference helps to improve ASR performance.
Investigating the Role of Synthetic Data Augmentation and Training Strategies on Improving Low-Resource Language ASR
Yun Hao | Reihaneh Amooie | Wietse de Vries | Rik van Noord | Martijn Wieling
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Yun Hao | Reihaneh Amooie | Wietse de Vries | Rik van Noord | Martijn Wieling
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Low-resource automatic speech recognition (ASR) is challenging due to a scarcity of annotated data. While synthetic data from text-to-speech (TTS) systems can augment ASR training, its efficacy for low-resource languages remains unclear. In this study, we investigate under which conditions TTS-based data augmentation is most effective for low-resource languages. Experiments on six low-resource languages in Common Voice show that synthetic data is most beneficial under extremely low-resource ASR conditions (i.e., less than one hour of available real speech data), or for languages with larger amounts of TTS data (i.e., more than 10 hours). Additionally, increasing the amount and diversity of synthetic data while keeping an appropriate ratio of synthetic-to-real data can further improve ASR performance.