Dhiraj Kumar Sah

2026

Improving Language Identification for Code-Switched Speech: The Pivotal Role of Accented English
Adyasha Patra | Dhiraj Kumar Sah | Preethi Jyothi
Findings of the Association for Computational Linguistics: EACL 2026

Code-switching, where speakers alternate between languages within a single utterance, poses unique challenges for language identification (LID). Existing LID models often fail to reliably identify English spoken with the accent of the matrix (dominant) language. We show that finetuning LID models with small amounts of such accented English significantly improves code-switched LID, without degrading performance on standard monolingual speech—a limitation observed with direct finetuning on code-switched utterances. This is achieved via low-rank adaptation (LoRA) on limited accented data, which allows models to adapt efficiently. To better evaluate performance, we introduce LangRank, a metric that captures the relative ranking of identified languages often overlooked by traditional metrics. Our method generalizes across multiple language pairs, including Hindi-English, Bengali-English, Mandarin-English, and Arabic-English, providing robust LID in code-switched multilingual contexts.

Co-authors

Preethi Jyothi 1
Adyasha Patra 1

Venues

Findings1

Fix author