LLM as a Morphological Disambiguator for Belarusian: A Preliminary Study

Vladislav Poritski, Oksana Volchek, Ilia Afanasev


Abstract
We explore the use of large language models (LLMs) for morphological disambiguation in Belarusian, a low-resource language. The pipeline has two stages: a rule-based analyzer generates candidate lemmas and grammatical tags, which an LLM then disambiguates in context. Initial evaluation of ChatGPT, Claude, and Gemini on a gold-standard sample shows high accuracy. We scale this approach to a 375K-word corpus using Gemini and compare the results against a neural baseline (Stanza). Manual review of discrepancies suggests that the LLM-based approach outperforms the baseline, offering a solution for corpus annotation in Belarusian.
Anthology ID:
2026.sigul-1.4
Volume:
Proceedings of the SIGUL 2026 Joint Workshop with ELE, EURALI, and DCLRL: Towards Inclusivity and Equality: Language Resources and Technologies for Under-Resourced and Endangered Languages
Month:
May
Year:
2026
Address:
Palma, Mallorca, Spain
Editors:
Atul Kr. Ojha, Sakriani Sakti, Claudia Soria, Maite Melero, John P. McCrae, Constantine Lignos, Chao-Hong Liu, German Rigau Claramunt, Georg Rehm
Venues:
SIGUL | EURALI | DCLRL | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
42–48
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-sigul-04
DOI:
10.63317/3skazxbd27m8
Bibkey:
Cite (ACL):
Vladislav Poritski, Oksana Volchek, and Ilia Afanasev. 2026. LLM as a Morphological Disambiguator for Belarusian: A Preliminary Study. In Proceedings of the SIGUL 2026 Joint Workshop with ELE, EURALI, and DCLRL: Towards Inclusivity and Equality: Language Resources and Technologies for Under-Resourced and Endangered Languages, pages 42–48, Palma, Mallorca, Spain. ELRA Language Resources Association (ELRA).
Cite (Informal):
LLM as a Morphological Disambiguator for Belarusian: A Preliminary Study (Poritski et al., SIGUL-EURALI-DCLRL 2026)
Copy Citation: