HARNESS: Lightweight Distilled Arabic Speech Foundation Models

Vrunda Nileshkumar Sukhadia, Shammur Absar Chowdhury


Abstract
Large self-supervised speech (SSL) models achieve strong downstream performance, but their size limits deployment in resource-constrained settings. We present HArnESS, an Arabic-centric self-supervised speech model family trained from scratch with iterative self-distillation, together with lightweight student variants that offer strong accuracy-efficiency trade-offs on Automatic Speech Recognition (ASR), Dialect Identification (DID), and Speech Emotion Recognition (SER). Our approach begins with a large bilingual Arabic-English teacher and progressively distills its knowledge into compressed student models while preserving Arabic-relevant acoustic and paralinguistic representations. We further study PCA-based compression of the teacher supervision signal to better match the capacity of shallow and thin students. Compared with HuBERT and XLS-R, HArnESS consistently improves performance on Arabic downstream tasks, while the compressed models remain competitive under substantial structural reduction. These results position HArnESS as a practical and accessible Arabic-centric SSL foundation for real-world speech applications.
Anthology ID:
2026.speakable-1.12
Volume:
Proceedings of Speech Language Models in Low-Resource Settings: Performance, Evaluation, and Bias Analysis (SPEAKABLE) @ LREC 2026
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Nina Hosseini-Kivanani, Alessio Brutti, Marco Matassoni, Sandipana Dowerah, Davide Liga, Christoph Schommer
Venues:
SPEAKABLE | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
109–117
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-speakable-12
DOI:
10.63317/2x7bxm8g49ju
Bibkey:
Cite (ACL):
Vrunda Nileshkumar Sukhadia and Shammur Absar Chowdhury. 2026. HARNESS: Lightweight Distilled Arabic Speech Foundation Models. In Proceedings of Speech Language Models in Low-Resource Settings: Performance, Evaluation, and Bias Analysis (SPEAKABLE) @ LREC 2026, pages 109–117, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
HARNESS: Lightweight Distilled Arabic Speech Foundation Models (Sukhadia & Chowdhury, SPEAKABLE 2026)
Copy Citation: