Towards Fair Speech Recognition: Mitigating Demographic Bias in End-to-End ASR Systems

Maliha Jahan, Thomas Thebaud, Zsuzsanna Fagyal, Jesus Villalba, Mark Hasegawa-Johnson, Laureano Moro Velazquez, Najim Dehak


Abstract
Demographic bias in the performance of speech and language technology has been an active area of recent research. A lot of studies have shown the existence of demographic biases in Automatic Speech Recognition (ASR) systems. In this work, we propose a novel model-agnostic and demographic label-agnostic approach, called DARe, to mitigate any existing bias in an ASR system towards certain speaker groups. We built a debiasing module that goes between the feature extractor of an ASR and the rest of that ASR. The module includes content-group disentanglers to separate content and group, a demographic classifier, and adversarial reweighting. To eliminate the need for demographic labels, we generated pseudo-group labels by extracting speaker embeddings and clustering them. We worked with three ASR systems–Wav2Vec2 base, SEW tiny, and Whisper small. We used the FAI dataset, which contains naturalistic conversations with speakers who self-identify their demographic attributes. We used Word Error Rate (WER) as a metric of ASR performance and a Poisson regression-based approach to evaluate the racial fairness of the models. We compared the racial bias of the models before and after applying our proposed approach and observed a significant improvement in fairness.
Anthology ID:
2026.lrec-1.325
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
4115–4125
Language:
External URL:
https://lrec.elra.info/lrec2026-main-325
DOI:
10.63317/5e5f93jkhzt5
Bibkey:
Cite (ACL):
Maliha Jahan, Thomas Thebaud, Zsuzsanna Fagyal, Jesus Villalba, Mark Hasegawa-Johnson, Laureano Moro Velazquez, and Najim Dehak. 2026. Towards Fair Speech Recognition: Mitigating Demographic Bias in End-to-End ASR Systems. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 4115–4125, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Towards Fair Speech Recognition: Mitigating Demographic Bias in End-to-End ASR Systems (Jahan et al., LREC 2026)
Copy Citation: