How Well Do Large Language Models Reason in Under-Resourced Languages? Evidence from Vietnamese

Tuan Anh Do, Jelke Bloem


Abstract
Despite advancements in Large Language Models, reasoning benchmarks remain centered on high-resource languages, leaving languages like Vietnamese under-evaluated. In this study, we aim to address this gap by evaluating four models: PhoGPT (native), Vistral and VBD-Llama (adapted), and Llama-2 (English-centric), on commonsense reasoning and arithmetic reasoning. As Vietnamese benchmarks for these tasks are lacking, we adapt two analogy datasets from English to Vietnamese and construct two sequence datasets, ensuring a range of structural complexity and difficulty levels. We evaluate diverse prompting strategies, including Chain-of-Thought, role-playing guidance, cross-lingual prompting, and few-shot learning. Our results reveal a baseline proficiency in analogical and arithmetic reasoning among the models, with Vistral and Llama-2 outperforming other models in multiple tasks. The effects of Chain-of-Thought and contextual guidance are limited in Vietnamese, while cross-lingual prompting and few-shot learning show promising performance improvements. The findings underscore the feasibility of adapting benchmarks to less-resourced languages and provide insights into strengths and weaknesses in the performance of Vietnamese LLMs, suggesting directions for model improvements.
Anthology ID:
2026.sigul-1.1
Volume:
Proceedings of the SIGUL 2026 Joint Workshop with ELE, EURALI, and DCLRL: Towards Inclusivity and Equality: Language Resources and Technologies for Under-Resourced and Endangered Languages
Month:
May
Year:
2026
Address:
Palma, Mallorca, Spain
Editors:
Atul Kr. Ojha, Sakriani Sakti, Claudia Soria, Maite Melero, John P. McCrae, Constantine Lignos, Chao-Hong Liu, German Rigau Claramunt, Georg Rehm
Venues:
SIGUL | EURALI | DCLRL | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
1–18
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-sigul-01
DOI:
10.63317/2a43bkurpywk
Bibkey:
Cite (ACL):
Tuan Anh Do and Jelke Bloem. 2026. How Well Do Large Language Models Reason in Under-Resourced Languages? Evidence from Vietnamese. In Proceedings of the SIGUL 2026 Joint Workshop with ELE, EURALI, and DCLRL: Towards Inclusivity and Equality: Language Resources and Technologies for Under-Resourced and Endangered Languages, pages 1–18, Palma, Mallorca, Spain. ELRA Language Resources Association (ELRA).
Cite (Informal):
How Well Do Large Language Models Reason in Under-Resourced Languages? Evidence from Vietnamese (Do & Bloem, SIGUL-EURALI-DCLRL 2026)
Copy Citation: