SpeechLM for Automatic Speech Recognition in Low-resource Languages

Md Abdur Razzaq Riyadh, Eneko Agirre, Eva Navas, Claudia Borg


Abstract
Multi-modal Speech Language Models (SpeechLMs) are a recent advancement in natural language processing. These SpeechLMs are instruction-tuned and optimized for general tasks. Their usefulness for Automatic Speech Recognition (ASR), particularly in relatively low-resource scenarios, remains largely understudied. This work developed SpeechLM for ASR in Basque and Maltese and studied the impact of language-adapted Large Language Model (LLM) and speech encoder within the SpeechLM for ASR. Using supervised learning, we fine-tuned LLaMA-Omni, a SpeechLM, for ASR. We have conducted comprehensive hyperparameter tuning and experimented with language-adapted SpeechLM components to improve performance and evaluated our best models on in-distribution datasets for both languages and an out-of-distribution dataset for Basque. LLaMA-Omni achieved 8.09% WER in Basque and 25.65% WER for Maltese on average across multiple test splits. The in-distribution results show that SpeechLM outperforms a fine-tuned ASR system under specific constraints, whereas it underperforms the baseline model on out-of-distribution Basque, indicating weaker overall robustness. We also find that a language-adapted LLM within SpeechLM improves in out-of-distribution settings when compared to the off-the-shelf LLM within SpeechLM.
Anthology ID:
2026.speakable-1.14
Volume:
Proceedings of Speech Language Models in Low-Resource Settings: Performance, Evaluation, and Bias Analysis (SPEAKABLE) @ LREC 2026
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Nina Hosseini-Kivanani, Alessio Brutti, Marco Matassoni, Sandipana Dowerah, Davide Liga, Christoph Schommer
Venues:
SPEAKABLE | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
125–131
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-speakable-14
DOI:
10.63317/2634e2pv97js
Bibkey:
Cite (ACL):
Md Abdur Razzaq Riyadh, Eneko Agirre, Eva Navas, and Claudia Borg. 2026. SpeechLM for Automatic Speech Recognition in Low-resource Languages. In Proceedings of Speech Language Models in Low-Resource Settings: Performance, Evaluation, and Bias Analysis (SPEAKABLE) @ LREC 2026, pages 125–131, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
SpeechLM for Automatic Speech Recognition in Low-resource Languages (Riyadh et al., SPEAKABLE 2026)
Copy Citation: