Analyzing BERT Cross-lingual Transfer Capabilities in Continual Sequence Labeling

Juan Manuel Coria, Mathilde Veron, Sahar Ghannay, Guillaume Bernard, Hervé Bredin, Olivier Galibert, Sophie Rosset


Abstract
Knowledge transfer between neural language models is a widely used technique that has proven to improve performance in a multitude of natural language tasks, in particular with the recent rise of large pre-trained language models like BERT. Similarly, high cross-lingual transfer has been shown to occur in multilingual language models. Hence, it is of great importance to better understand this phenomenon as well as its limits. While most studies about cross-lingual transfer focus on training on independent and identically distributed (i.e. i.i.d.) samples, in this paper we study cross-lingual transfer in a continual learning setting on two sequence labeling tasks: slot-filling and named entity recognition. We investigate this by training multilingual BERT on sequences of 9 languages, one language at a time, on the MultiATIS++ and MultiCoNER corpora. Our first findings are that forward transfer between languages is retained although forgetting is present. Additional experiments show that lost performance can be recovered with as little as a single training epoch even if forgetting was high, which can be explained by a progressive shift of model parameters towards a better multilingual initialization. We also find that commonly used metrics might be insufficient to assess continual learning performance.
Anthology ID:
2022.mmmpie-1.3
Volume:
Proceedings of the First Workshop on Performance and Interpretability Evaluations of Multimodal, Multipurpose, Massive-Scale Models
Month:
October
Year:
2022
Address:
Virtual
Venue:
MMMPIE
SIG:
Publisher:
International Conference on Computational Linguistics
Note:
Pages:
15–25
Language:
URL:
https://aclanthology.org/2022.mmmpie-1.3
DOI:
Bibkey:
Cite (ACL):
Juan Manuel Coria, Mathilde Veron, Sahar Ghannay, Guillaume Bernard, Hervé Bredin, Olivier Galibert, and Sophie Rosset. 2022. Analyzing BERT Cross-lingual Transfer Capabilities in Continual Sequence Labeling. In Proceedings of the First Workshop on Performance and Interpretability Evaluations of Multimodal, Multipurpose, Massive-Scale Models, pages 15–25, Virtual. International Conference on Computational Linguistics.
Cite (Informal):
Analyzing BERT Cross-lingual Transfer Capabilities in Continual Sequence Labeling (Coria et al., MMMPIE 2022)
Copy Citation:
PDF:
https://aclanthology.org/2022.mmmpie-1.3.pdf
Code
 juanmc2005/continualnlu
Data
ATISMultiCoNER