CCMatrix: Mining Billions of High-Quality Parallel Sentences on the Web

Holger Schwenk; Guillaume Wenzek; Sergey Edunov; Édouard Grave; Armand Joulin; Angela Fan

doi:10.18653/v1/2021.acl-long.507

CCMatrix: Mining Billions of High-Quality Parallel Sentences on the Web

Holger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave, Armand Joulin, Angela Fan

Abstract

We show that margin-based bitext mining in a multilingual sentence space can be successfully scaled to operate on monolingual corpora of billions of sentences. We use 32 snapshots of a curated common crawl corpus (Wenzel et al, 2019) totaling 71 billion unique sentences. Using one unified approach for 90 languages, we were able to mine 10.8 billion parallel sentences, out of which only 2.9 billions are aligned with English. We illustrate the capability of our scalable mining system to create high quality training sets from one language to any other by training hundreds of different machine translation models and evaluating them on the many-to-many TED benchmark. Further, we evaluate on competitive translation benchmarks such as WMT and WAT. Using only mined bitext, we set a new state of the art for a single system on the WMT’19 test set for English-German/Russian/Chinese. In particular, our English/German and English/Russian systems outperform the best single ones by over 4 BLEU points and are on par with best WMT’19 systems, which train on the WMT training data and augment it with backtranslation. We also achieve excellent results for distant languages pairs like Russian/Japanese, outperforming the best submission at the 2020 WAT workshop. All of the mined bitext will be freely available.

Anthology ID:: 2021.acl-long.507
Volume:: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)
Month:: August
Year:: 2021
Address:: Online
Editors:: Chengqing Zong, Fei Xia, Wenjie Li, Roberto Navigli
Venues:: ACL | IJCNLP
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 6490–6500
Language:
URL:: https://aclanthology.org/2021.acl-long.507/
DOI:: 10.18653/v1/2021.acl-long.507
Bibkey:
Cite (ACL):: Holger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave, Armand Joulin, and Angela Fan. 2021. CCMatrix: Mining Billions of High-Quality Parallel Sentences on the Web. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6490–6500, Online. Association for Computational Linguistics.
Cite (Informal):: CCMatrix: Mining Billions of High-Quality Parallel Sentences on the Web (Schwenk et al., ACL-IJCNLP 2021)
Copy Citation:
PDF:: https://aclanthology.org/2021.acl-long.507.pdf
Optionalsupplementarymaterial:: 2021.acl-long.507.OptionalSupplementaryMaterial.zip
Video:: https://aclanthology.org/2021.acl-long.507.mp4

PDF Cite Search Optionalsupplementarymaterial Video Fix data