Yet Another Format of Universal Dependencies for Korean

Yige Chen, Eunkyul Leah Jo, Yundong Yao, KyungTae Lim, Miikka Silfverberg, Francis M. Tyers, Jungyeul Park


Abstract
In this study, we propose a morpheme-based scheme for Korean dependency parsing and adopt the proposed scheme to Universal Dependencies. We present the linguistic rationale that illustrates the motivation and the necessity of adopting the morpheme-based format, and develop scripts that convert between the original format used by Universal Dependencies and the proposed morpheme-based format automatically. The effectiveness of the proposed format for Korean dependency parsing is then testified by both statistical and neural models, including UDPipe and Stanza, with our carefully constructed morpheme-based word embedding for Korean. morphUD outperforms parsing results for all Korean UD treebanks, and we also present detailed error analysis.
Anthology ID:
2022.coling-1.482
Volume:
Proceedings of the 29th International Conference on Computational Linguistics
Month:
October
Year:
2022
Address:
Gyeongju, Republic of Korea
Venue:
COLING
SIG:
Publisher:
International Committee on Computational Linguistics
Note:
Pages:
5432–5437
Language:
URL:
https://aclanthology.org/2022.coling-1.482
DOI:
Bibkey:
Cite (ACL):
Yige Chen, Eunkyul Leah Jo, Yundong Yao, KyungTae Lim, Miikka Silfverberg, Francis M. Tyers, and Jungyeul Park. 2022. Yet Another Format of Universal Dependencies for Korean. In Proceedings of the 29th International Conference on Computational Linguistics, pages 5432–5437, Gyeongju, Republic of Korea. International Committee on Computational Linguistics.
Cite (Informal):
Yet Another Format of Universal Dependencies for Korean (Chen et al., COLING 2022)
Copy Citation:
PDF:
https://aclanthology.org/2022.coling-1.482.pdf
Code
 jungyeul/morphud-korean
Data
Universal Dependencies