A Coarse-to-Fine Labeling Framework for Joint Word Segmentation, POS Tagging, and Constituent Parsing

Yang Hou; Houquan Zhou (周厚全); Zhenghua Li (李正华); Yu Zhang; Min Zhang; Zhefeng Wang; Baoxing Huai; Nicholas Jing Yuan

doi:10.18653/v1/2021.conll-1.23

A Coarse-to-Fine Labeling Framework for Joint Word Segmentation, POS Tagging, and Constituent Parsing

Yang Hou, Houquan Zhou, Zhenghua Li, Yu Zhang, Min Zhang, Zhefeng Wang, Baoxing Huai, Nicholas Jing Yuan

Abstract

The most straightforward approach to joint word segmentation (WS), part-of-speech (POS) tagging, and constituent parsing is converting a word-level tree into a char-level tree, which, however, leads to two severe challenges. First, a larger label set (e.g., ≥ 600) and longer inputs both increase computational costs. Second, it is difficult to rule out illegal trees containing conflicting production rules, which is important for reliable model evaluation. If a POS tag (like VV) is above a phrase tag (like VP) in the output tree, it becomes quite complex to decide word boundaries. To deal with both challenges, this work proposes a two-stage coarse-to-fine labeling framework for joint WS-POS-PAR. In the coarse labeling stage, the joint model outputs a bracketed tree, in which each node corresponds to one of four labels (i.e., phrase, subphrase, word, subword). The tree is guaranteed to be legal via constrained CKY decoding. In the fine labeling stage, the model expands each coarse label into a final label (such as VP, VP*, VV, VV*). Experiments on Chinese Penn Treebank 5.1 and 7.0 show that our joint model consistently outperforms the pipeline approach on both settings of w/o and w/ BERT, and achieves new state-of-the-art performance.

Anthology ID:: 2021.conll-1.23
Volume:: Proceedings of the 25th Conference on Computational Natural Language Learning
Month:: November
Year:: 2021
Address:: Online
Editors:: Arianna Bisazza, Omri Abend
Venue:: CoNLL
SIG:: SIGNLL
Publisher:: Association for Computational Linguistics
Note:
Pages:: 290–299
Language:
URL:: https://aclanthology.org/2021.conll-1.23/
DOI:: 10.18653/v1/2021.conll-1.23
Bibkey:
Cite (ACL):: Yang Hou, Houquan Zhou, Zhenghua Li, Yu Zhang, Min Zhang, Zhefeng Wang, Baoxing Huai, and Nicholas Jing Yuan. 2021. A Coarse-to-Fine Labeling Framework for Joint Word Segmentation, POS Tagging, and Constituent Parsing. In Proceedings of the 25th Conference on Computational Natural Language Learning, pages 290–299, Online. Association for Computational Linguistics.
Cite (Informal):: A Coarse-to-Fine Labeling Framework for Joint Word Segmentation, POS Tagging, and Constituent Parsing (Hou et al., CoNLL 2021)
Copy Citation:
PDF:: https://aclanthology.org/2021.conll-1.23.pdf
Video:: https://aclanthology.org/2021.conll-1.23.mp4

PDF Cite Search Video Fix data