Ensemble Transfer Learning for Multilingual Coreference Resolution

Tuan Lai, Heng Ji


Abstract
Entity coreference resolution is an important research problem with many applications, including information extraction and question answering. Coreference resolution for English has been studied extensively. However, there is relatively little work for other languages. A problem that frequently occurs when working with a non-English language is the scarcity of annotated training data. To overcome this challenge, we design a simple but effective ensemble-based framework that combines various transfer learning (TL) techniques. We first train several models using different TL methods. Then, during inference, we compute the unweighted average scores of the models’ predictions to extract the final set of predicted clusters. Furthermore, we also propose a low-cost TL method that bootstraps coreference resolution models by utilizing Wikipedia anchor texts. Leveraging the idea that the coreferential links naturally exist between anchor texts pointing to the same article, our method builds a sizeable distantly-supervised dataset for the target language that consists of tens of thousands of documents. We can pre-train a model on the pseudo-labeled dataset before finetuning it on the final target dataset. Experimental results on two benchmark datasets, OntoNotes and SemEval, confirm the effectiveness of our methods. Our best ensembles consistently outperform the baseline approach of simple training by up to 7.68% in the F1 score. These ensembles also achieve new state-of-the-art results for three languages: Arabic, Dutch, and Spanish.
Anthology ID:
2023.codi-1.3
Volume:
Proceedings of the 4th Workshop on Computational Approaches to Discourse (CODI 2023)
Month:
July
Year:
2023
Address:
Toronto, Canada
Editors:
Michael Strube, Chloe Braud, Christian Hardmeier, Junyi Jessy Li, Sharid Loaiciga, Amir Zeldes
Venue:
CODI
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
24–36
Language:
URL:
https://aclanthology.org/2023.codi-1.3
DOI:
10.18653/v1/2023.codi-1.3
Bibkey:
Cite (ACL):
Tuan Lai and Heng Ji. 2023. Ensemble Transfer Learning for Multilingual Coreference Resolution. In Proceedings of the 4th Workshop on Computational Approaches to Discourse (CODI 2023), pages 24–36, Toronto, Canada. Association for Computational Linguistics.
Cite (Informal):
Ensemble Transfer Learning for Multilingual Coreference Resolution (Lai & Ji, CODI 2023)
Copy Citation:
PDF:
https://aclanthology.org/2023.codi-1.3.pdf