Unsupervised Morphological Paradigm Completion

Huiming Jin, Liwei Cai, Yihui Peng, Chen Xia, Arya McCarthy, Katharina Kann


Abstract
We propose the task of unsupervised morphological paradigm completion. Given only raw text and a lemma list, the task consists of generating the morphological paradigms, i.e., all inflected forms, of the lemmas. From a natural language processing (NLP) perspective, this is a challenging unsupervised task, and high-performing systems have the potential to improve tools for low-resource languages or to assist linguistic annotators. From a cognitive science perspective, this can shed light on how children acquire morphological knowledge. We further introduce a system for the task, which generates morphological paradigms via the following steps: (i) EDIT TREE retrieval, (ii) additional lemma retrieval, (iii) paradigm size discovery, and (iv) inflection generation. We perform an evaluation on 14 typologically diverse languages. Our system outperforms trivial baselines with ease and, for some languages, even obtains a higher accuracy than minimally supervised systems.
Anthology ID:
2020.acl-main.598
Volume:
Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics
Month:
July
Year:
2020
Address:
Online
Editors:
Dan Jurafsky, Joyce Chai, Natalie Schluter, Joel Tetreault
Venue:
ACL
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
6696–6707
Language:
URL:
https://aclanthology.org/2020.acl-main.598
DOI:
10.18653/v1/2020.acl-main.598
Bibkey:
Cite (ACL):
Huiming Jin, Liwei Cai, Yihui Peng, Chen Xia, Arya McCarthy, and Katharina Kann. 2020. Unsupervised Morphological Paradigm Completion. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6696–6707, Online. Association for Computational Linguistics.
Cite (Informal):
Unsupervised Morphological Paradigm Completion (Jin et al., ACL 2020)
Copy Citation:
PDF:
https://aclanthology.org/2020.acl-main.598.pdf
Video:
 http://slideslive.com/38929159
Code
 cai-lw/morpho-baseline