NE-LID: A Fast and Accurate Language Identification System for Northeast Indian Languages

Badal Nyalang


Abstract
Language identification (LID) is crucial for natural language processing systems, yet Northeast Indian languages remain severely underserved by existing multilingual LID models. We present NE-LID, a fast and accurate language identification system specifically designed for eleven languages of Northeast India. Built using character n-gram features with fastText, NE-LID achieves 99.09% accuracy on a balanced test set, significantly outperforming existing multilingual systems including GlotLID (73.12%), OpenLID (42.03%), IndicLID (39.30%), and LangDetect (24.33%). Our model processes predictions in 0.084 milliseconds on average, enabling real-time applications. We demonstrate that character-level modeling outperforms transformer-based approaches for script-diverse, low-resource languages
Anthology ID:
2026.wildre-1.14
Volume:
Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation
Month:
May
Year:
2026
Address:
Palma, Mallorca, Spain
Editors:
Girish Nath Jha, Kalika Bali, Sobha L, Devendr Kumar
Venues:
WILDRE | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
104–108
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-wildre-14
DOI:
10.63317/2q69fiyg73p6
Bibkey:
Cite (ACL):
Badal Nyalang. 2026. NE-LID: A Fast and Accurate Language Identification System for Northeast Indian Languages. In Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation, pages 104–108, Palma, Mallorca, Spain. ELRA Language Resources Association (ELRA).
Cite (Informal):
NE-LID: A Fast and Accurate Language Identification System for Northeast Indian Languages (Nyalang, WILDRE 2026)
Copy Citation: