Investigating Language Relationships in Multilingual Sentence Encoders Through the Lens of Linguistic Typology

Rochelle Choenni; Ekaterina Shutova

doi:10.1162/coli_a_00444

Investigating Language Relationships in Multilingual Sentence Encoders Through the Lens of Linguistic Typology

Abstract

Multilingual sentence encoders have seen much success in cross-lingual model transfer for downstream NLP tasks. The success of this transfer is, however, dependent on the model’s ability to encode the patterns of cross-lingual similarity and variation. Yet, we know relatively little about the properties of individual languages or the general patterns of linguistic variation that the models encode. In this article, we investigate these questions by leveraging knowledge from the field of linguistic typology, which studies and documents structural and semantic variation across languages. We propose methods for separating language-specific subspaces within state-of-the-art multilingual sentence encoders (LASER, M-BERT, XLM, and XLM-R) with respect to a range of typological properties pertaining to lexical, morphological, and syntactic structure. Moreover, we investigate how typological information about languages is distributed across all layers of the models. Our results show interesting differences in encoding linguistic variation associated with different pretraining strategies. In addition, we propose a simple method to study how shared typological properties of languages are encoded in two state-of-the-art multilingual models—M-BERT and XLM-R. The results provide insight into their information-sharing mechanisms and suggest that these linguistic properties are encoded jointly across typologically similar languages in these models.

Anthology ID:: 2022.cl-3.5
Volume:: Computational Linguistics, Volume 48, Issue 3 - September 2022
Month:: September
Year:: 2022
Address:: Cambridge, MA
Venue:: CL
SIG:
Publisher:: MIT Press
Note:
Pages:: 635–672
Language:
URL:: https://aclanthology.org/2022.cl-3.5
DOI:: 10.1162/coli_a_00444
Bibkey:
Cite (ACL):: Rochelle Choenni and Ekaterina Shutova. 2022. Investigating Language Relationships in Multilingual Sentence Encoders Through the Lens of Linguistic Typology. Computational Linguistics, 48(3):635–672.
Cite (Informal):: Investigating Language Relationships in Multilingual Sentence Encoders Through the Lens of Linguistic Typology (Choenni & Shutova, CL 2022)
Copy Citation:
PDF:: https://aclanthology.org/2022.cl-3.5.pdf
Video:: https://aclanthology.org/2022.cl-3.5.mp4
Data: Tatoeba, XNLI

PDF Cite Search Video