BURMESE-SAN: Burmese NLP Benchmark for Evaluating Large Language Models

Thura Aung, Jann Railey Montalan, Jian Gang Ngui, Peerat Limkonchotiwat


Abstract
We introduce BURMESE-SAN, the first holistic benchmark that systematically evaluates large language models (LLMs) for Burmese across three core NLP competencies: understanding (NLU), reasoning (NLR), and generation (NLG). BURMESE-SAN consolidates seven subtasks spanning these competencies, including Question Answering, Sentiment Analysis, Toxicity Detection, Causal Reasoning, Natural Language Inference, Abstractive Summarization, and Machine Translation, several of which were previously unavailable for Burmese. The benchmark is constructed through a rigorous native-speaker-driven process to ensure linguistic naturalness, fluency, and cultural authenticity while minimizing translation-induced artifacts. We conduct a large-scale evaluation of both open-weight and commercial LLMs to examine challenges in Burmese modeling arising from limited pretraining coverage, rich morphology, and syntactic variation. Our results show that Burmese performance depends more on architectural design, language representation, and instruction tuning than on model scale alone. In particular, Southeast Asia regional fine-tuning and newer model generations yield substantial gains. Finally, we release BURMESE-SAN as a public leaderboard to support systematic evaluation and sustained progress in Burmese and other low-resource languages. https://leaderboard.sea-lion.ai/detailed/MY
Anthology ID:
2026.lrec-1.16
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
224–245
Language:
External URL:
https://lrec.elra.info/lrec2026-main-016
DOI:
10.63317/54dxgzy8h77c
Bibkey:
Cite (ACL):
Thura Aung, Jann Railey Montalan, Jian Gang Ngui, and Peerat Limkonchotiwat. 2026. BURMESE-SAN: Burmese NLP Benchmark for Evaluating Large Language Models. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 224–245, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
BURMESE-SAN: Burmese NLP Benchmark for Evaluating Large Language Models (Aung et al., LREC 2026)
Copy Citation: