ICON: Building a Large-Scale Benchmark Constituency Treebank for the Indonesian Language

Ee Suan Lim; Wei Qi Leong; Ngan Thanh Nguyen; Dea Adhista; Wei Ming Kng; William Chandra Tjh; Ayu Purwarianti

ICON: Building a Large-Scale Benchmark Constituency Treebank for the Indonesian Language

Ee Suan Lim, Wei Qi Leong, Ngan Thanh Nguyen, Dea Adhista, Wei Ming Kng, William Chandra Tjh, Ayu Purwarianti

Abstract

Constituency parsing is an important task of informing how words are combined to form sentences. While constituency parsing in English has seen significant progress in the last few years, tools for constituency parsing in Indonesian remain few and far between. In this work, we publish ICON (Indonesian CONstituency treebank), the hitherto largest publicly-available manually-annotated benchmark constituency treebank for the Indonesian language with a size of 10,000 sentences and approximately 124,000 constituents and 182,000 tokens, which can support the training of state-of-the-art transformer-based models. We establish strong baselines on the ICON dataset using the Berkeley Neural Parser with transformer-based pre-trained embeddings, with the best performance of 88.85% F1 score coming from our own version of SpanBERT (IndoSpanBERT). We further analyze the predictions made by our best-performing model to reveal certain idiosyncrasies in the Indonesian language that pose challenges for constituency parsing.

Anthology ID:: 2023.tlt-1.5
Volume:: Proceedings of the 21st International Workshop on Treebanks and Linguistic Theories (TLT, GURT/SyntaxFest 2023)
Month:: March
Year:: 2023
Address:: Washington, D.C.
Editors:: Daniel Dakota, Kilian Evang, Sandra Kübler, Lori Levin
Venues:: TLT | SyntaxFest
SIG:: SIGPARSE
Publisher:: Association for Computational Linguistics
Note:
Pages:: 37–53
Language:
URL:: https://aclanthology.org/2023.tlt-1.5/
DOI:
Bibkey:
Cite (ACL):: Ee Suan Lim, Wei Qi Leong, Ngan Thanh Nguyen, Dea Adhista, Wei Ming Kng, William Chandra Tjh, and Ayu Purwarianti. 2023. ICON: Building a Large-Scale Benchmark Constituency Treebank for the Indonesian Language. In Proceedings of the 21st International Workshop on Treebanks and Linguistic Theories (TLT, GURT/SyntaxFest 2023), pages 37–53, Washington, D.C.. Association for Computational Linguistics.
Cite (Informal):: ICON: Building a Large-Scale Benchmark Constituency Treebank for the Indonesian Language (Suan Lim et al., TLT-SyntaxFest 2023)
Copy Citation:
PDF:: https://aclanthology.org/2023.tlt-1.5.pdf
Video:: https://aclanthology.org/2023.tlt-1.5.mp4

PDF Cite Search Video Fix data