South Tyrolean Dialect-to-Standard Speech Translation: A Resource

Greta H. Franzini, Luca Ducceschi


Abstract
This paper presents a developing oral resource for South Tyrolean, a German dialect spoken in Northern Italy. The dialect is ubiquitous in spoken communication but lacks a standardised orthography. In this context, strict transcription into dialect is of limited to no utility to the local community. Instead, there is a distinct and strong demand for technology capable of directly translating spoken dialect into Standard German. To address this specific need, we introduce a dynamic, incrementally growing dataset designed to fine-tune ASR models for this translation task. Our corpus aggregates diverse sources, including media and research interviews, totalling over 13 hours of aligned audio. We describe a collaborative workflow where community partners contribute audio archives in exchange for automated transcriptions, creating a virtuous cycle of data improvement. Additionally, we detail our iterative model fine-tuning strategy, data collection challenges and the resulting improvements in model performance.
Anthology ID:
2026.dialres-1.19
Volume:
Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective
Month:
May
Year:
2026
Address:
Palma de Mallorca
Editors:
Antonis Anastasopoulos, Stella Markantonatou, Angela Ralli, Marcos Zampieri, Stavros Bompolas, Vivian Stamou
Venues:
DialRes | WS
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
188–194
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-dialres-19
DOI:
10.63317/3visgk9f8s7z
Bibkey:
Cite (ACL):
Greta H. Franzini and Luca Ducceschi. 2026. South Tyrolean Dialect-to-Standard Speech Translation: A Resource. In Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective, pages 188–194, Palma de Mallorca. Association for Computational Linguistics.
Cite (Informal):
South Tyrolean Dialect-to-Standard Speech Translation: A Resource (Franzini & Ducceschi, DialRes 2026)
Copy Citation: