Common Voice for Pakistan: Developing an Open Speech Corpus for Low-Resource Pakistani Languages

Meesum Alam, Francis Tyers


Abstract
Pakistan is home to more than 70 languages out of which 30 languages are endangered. Most of Pakistani languages remain absent from modern speech and text technologies, with resources focused on Urdu and a few major tongues. Through Mozilla’s Open Multilingual Speech Fund, this paper documents one year project for the development of an open, community driven speech corpus for 39 indigenous languages of Pakistan. The dataset includes locally authored texts, daily life sentences, poetry, and folk songs to make a culturally balanced. The project not only supports Automatic Speech Recognition but also promote linguistic preservation and digital inclusion.
Anthology ID:
2026.lrec-1.265
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
3355–3359
Language:
External URL:
https://lrec.elra.info/lrec2026-main-265
DOI:
10.63317/4r3mie85u8cq
Bibkey:
Cite (ACL):
Meesum Alam and Francis Tyers. 2026. Common Voice for Pakistan: Developing an Open Speech Corpus for Low-Resource Pakistani Languages. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 3355–3359, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Common Voice for Pakistan: Developing an Open Speech Corpus for Low-Resource Pakistani Languages (Alam & Tyers, LREC 2026)
Copy Citation: