mPLM-Sim: Better Cross-Lingual Similarity and Transfer in Multilingual Pretrained Language Models

Peiqin Lin, Chengzhi Hu, Zheyu Zhang, Andre Martins, Hinrich Schuetze


Abstract
Recent multilingual pretrained language models (mPLMs) have been shown to encode strong language-specific signals, which are not explicitly provided during pretraining. It remains an open question whether it is feasible to employ mPLMs to measure language similarity, and subsequently use the similarity results to select source languages for boosting cross-lingual transfer. To investigate this, we propose mPLM-Sim, a language similarity measure that induces the similarities across languages from mPLMs using multi-parallel corpora. Our study shows that mPLM-Sim exhibits moderately high correlations with linguistic similarity measures, such as lexicostatistics, genealogical language family, and geographical sprachbund. We also conduct a case study on languages with low correlation and observe that mPLM-Sim yields more accurate similarity results. Additionally, we find that similarity results vary across different mPLMs and different layers within an mPLM. We further investigate whether mPLM-Sim is effective for zero-shot cross-lingual transfer by conducting experiments on both low-level syntactic tasks and high-level semantic tasks. The experimental results demonstrate that mPLM-Sim is capable of selecting better source languages than linguistic measures, resulting in a 1%-2% improvement in zero-shot cross-lingual transfer performance.
Anthology ID:
2024.findings-eacl.20
Volume:
Findings of the Association for Computational Linguistics: EACL 2024
Month:
March
Year:
2024
Address:
St. Julian’s, Malta
Editors:
Yvette Graham, Matthew Purver
Venue:
Findings
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
276–310
Language:
URL:
https://aclanthology.org/2024.findings-eacl.20
DOI:
Bibkey:
Cite (ACL):
Peiqin Lin, Chengzhi Hu, Zheyu Zhang, Andre Martins, and Hinrich Schuetze. 2024. mPLM-Sim: Better Cross-Lingual Similarity and Transfer in Multilingual Pretrained Language Models. In Findings of the Association for Computational Linguistics: EACL 2024, pages 276–310, St. Julian’s, Malta. Association for Computational Linguistics.
Cite (Informal):
mPLM-Sim: Better Cross-Lingual Similarity and Transfer in Multilingual Pretrained Language Models (Lin et al., Findings 2024)
Copy Citation:
PDF:
https://aclanthology.org/2024.findings-eacl.20.pdf