Steven Limcorn
2025
NusaBERT: Teaching IndoBERT to be Multilingual and Multicultural
Wilson Wongso
|
David Samuel Setiawan
|
Steven Limcorn
|
Ananto Joyoadikusumo
Proceedings of the Second Workshop in South East Asian Language Processing
We present NusaBERT, a multilingual model built on IndoBERT and tailored for Indonesia’s diverse languages. By expanding vocabulary and pre-training on a regional corpus, NusaBERT achieves state-of-the-art performance on Indonesian NLU benchmarks, enhancing IndoBERT’s multilingual capability. This study also addresses NusaBERT’s limitations and encourages further research on Indonesia’s underrepresented languages.