The Denglisch Corpus of German-English Code-Switching

Doreen Osmelak, Shuly Wintner


Abstract
When multilingual speakers involve in a conversation they inevitably introduce code-switching (CS), i.e., mixing of more than one language between and within utterances. CS is still an understudied phenomenon, especially in the written medium, and relatively few computational resources for studying it are available. We describe a corpus of German-English code-switching in social media interactions. We focus on some challenges in annotating CS, especially due to words whose language ID cannot be easily determined. We introduce a novel schema for such word-level annotation, with which we manually annotated a subset of the corpus. We then trained classifiers to predict and identify switches, and applied them to the remainder of the corpus. Thereby, we created a large scale corpus of German-English mixed utterances with precise indications of CS points.
Anthology ID:
2023.sigtyp-1.5
Volume:
Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual NLP
Month:
May
Year:
2023
Address:
Dubrovnik, Croatia
Editors:
Lisa Beinborn, Koustava Goswami, Saliha Muradoğlu, Alexey Sorokin, Ritesh Kumar, Andreas Shcherbakov, Edoardo M. Ponti, Ryan Cotterell, Ekaterina Vylomova
Venue:
SIGTYP
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
42–51
Language:
URL:
https://aclanthology.org/2023.sigtyp-1.5
DOI:
10.18653/v1/2023.sigtyp-1.5
Bibkey:
Cite (ACL):
Doreen Osmelak and Shuly Wintner. 2023. The Denglisch Corpus of German-English Code-Switching. In Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual NLP, pages 42–51, Dubrovnik, Croatia. Association for Computational Linguistics.
Cite (Informal):
The Denglisch Corpus of German-English Code-Switching (Osmelak & Wintner, SIGTYP 2023)
Copy Citation:
PDF:
https://aclanthology.org/2023.sigtyp-1.5.pdf
Video:
 https://aclanthology.org/2023.sigtyp-1.5.mp4