Sociolinguistic aspects of crowdsourcing for a vocal corpus of Alsatian

Pascale Erhart, Lucile Hamm, Sam Bigeard, Carole Werner, Malek Yaich, Slim Ouni


Abstract
Alsatian is a regional low-resource language spoken in a majority-language context. In order to create a voice dataset suited for training automatic speech recognition and speech-to-text models, we launched a crowdsourcing campaign on the platform Mozilla Common Voice. We describe sociolinguistic issues we ran into, such as participants’ perception of their own language and its role in the AI landscape, which are vital to address to raise the participation in the crowdsourcing effort. We found that the participants are often confused about NLP and AI tools, and have a strong interested in preserving their language.
Anthology ID:
2026.dialres-1.25
Volume:
Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective
Month:
May
Year:
2026
Address:
Palma de Mallorca
Editors:
Antonis Anastasopoulos, Stella Markantonatou, Angela Ralli, Marcos Zampieri, Stavros Bompolas, Vivian Stamou
Venues:
DialRes | WS
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
256–264
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-dialres-25
DOI:
10.63317/5ch7cwah438g
Bibkey:
Cite (ACL):
Pascale Erhart, Lucile Hamm, Sam Bigeard, Carole Werner, Malek Yaich, and Slim Ouni. 2026. Sociolinguistic aspects of crowdsourcing for a vocal corpus of Alsatian. In Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective, pages 256–264, Palma de Mallorca. Association for Computational Linguistics.
Cite (Informal):
Sociolinguistic aspects of crowdsourcing for a vocal corpus of Alsatian (Erhart et al., DialRes 2026)
Copy Citation: