Efficient methods for building LLMs for low-resourced languages

Nalin Kumar


Abstract
Large Language Models (LLMs) excel in many NLP tasks but remain biased toward high-resource languages. This position paper discusses the author’s current findings on efficient strategies for low-resource settings: (i) modular training, where only non-embedding parameters are tuned after learning language-specific tokenizers and embeddings, and (ii) artificial language initialization, which leverages structurally biased synthetic languages for faster, parameter-efficient pretraining. The paper also shares plans for future research and topics that the author would like to discuss during the round-table.
Anthology ID:
2025.ynlg-main.5
Volume:
Proceedings of the 1st Workshop for Young Researchers in Natural Language Generation
Month:
October
Year:
2025
Address:
Hanoi, Vietnam
Editors:
Alyssa Allen, Nils Feldhus, Rudali Huidrom, Michela Lorandi, Adarsa Sivaprasad, Patrícia Schmidtová
Venue:
YNLG
SIG:
SIGGEN
Publisher:
Association for Computational Linguistics
Note:
Pages:
21–23
Language:
URL:
https://aclanthology.org/2025.ynlg-main.5/
DOI:
Bibkey:
Cite (ACL):
Nalin Kumar. 2025. Efficient methods for building LLMs for low-resourced languages. In Proceedings of the 1st Workshop for Young Researchers in Natural Language Generation, pages 21–23, Hanoi, Vietnam. Association for Computational Linguistics.
Cite (Informal):
Efficient methods for building LLMs for low-resourced languages (Kumar, YNLG 2025)
Copy Citation:
PDF:
https://aclanthology.org/2025.ynlg-main.5.pdf