Cross-lingual Parsing with Polyglot Training and Multi-treebank Learning: A Faroese Case Study

James Barry, Joachim Wagner, Jennifer Foster


Abstract
Cross-lingual dependency parsing involves transferring syntactic knowledge from one language to another. It is a crucial component for inducing dependency parsers in low-resource scenarios where no training data for a language exists. Using Faroese as the target language, we compare two approaches using annotation projection: first, projecting from multiple monolingual source models; second, projecting from a single polyglot model which is trained on the combination of all source languages. Furthermore, we reproduce multi-source projection (Tyers et al., 2018), in which dependency trees of multiple sources are combined. Finally, we apply multi-treebank modelling to the projected treebanks, in addition to or alternatively to polyglot modelling on the source side. We find that polyglot training on the source languages produces an overall trend of better results on the target language but the single best result for the target language is obtained by projecting from monolingual source parsing models and then training multi-treebank POS tagging and parsing models on the target side.
Anthology ID:
D19-6118
Volume:
Proceedings of the 2nd Workshop on Deep Learning Approaches for Low-Resource NLP (DeepLo 2019)
Month:
November
Year:
2019
Address:
Hong Kong, China
Editors:
Colin Cherry, Greg Durrett, George Foster, Reza Haffari, Shahram Khadivi, Nanyun Peng, Xiang Ren, Swabha Swayamdipta
Venue:
WS
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
163–174
Language:
URL:
https://aclanthology.org/D19-6118
DOI:
10.18653/v1/D19-6118
Bibkey:
Cite (ACL):
James Barry, Joachim Wagner, and Jennifer Foster. 2019. Cross-lingual Parsing with Polyglot Training and Multi-treebank Learning: A Faroese Case Study. In Proceedings of the 2nd Workshop on Deep Learning Approaches for Low-Resource NLP (DeepLo 2019), pages 163–174, Hong Kong, China. Association for Computational Linguistics.
Cite (Informal):
Cross-lingual Parsing with Polyglot Training and Multi-treebank Learning: A Faroese Case Study (Barry et al., 2019)
Copy Citation:
PDF:
https://aclanthology.org/D19-6118.pdf
Code
 Jbar-ry/multilingual-parsing