Yao Lu
Author directoryOther people with similar names: Yao Lu, Yao Lu
Unverified author pages with similar names: Yao Lu
2026
The Role of Mixed-Language Documents for Multilingual Large Language Model Pretraining
Jiandong Shao | Raphael Tang | Crystina Zhang | Karin Sevegnani | Pontus Stenetorp | Jianfei Yang | Yao Lu
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Jiandong Shao | Raphael Tang | Crystina Zhang | Karin Sevegnani | Pontus Stenetorp | Jianfei Yang | Yao Lu
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Multilingual large language models achieve impressive cross-lingual performance despite largely monolingual pretraining. While bilingual data in pretraining corpora is widely believed to enable these abilities, details of its contributions remain unclear. We investigate this question by pretraining models from scratch under controlled conditions, comparing the standard web corpus with a monolingual-only version that removes all multilingual documents. Despite constituting only 2% of the corpus, removing bilingual data causes translation performance to drop 56% in BLEU, while behaviour on cross-lingual QA and general reasoning tasks remains stable, with training curves largely overlapping the baseline. To understand this asymmetry, we categorize bilingual data into parallel (14%), code-switching (72%), and miscellaneous documents (14%) based on the semantic relevance of content in different languages. We then conduct granular ablations by reintroducing parallel or code-switching data into the monolingual-only corpus. Our experiments reveal that parallel data almost fully restores translation performance (91% of the unfiltered baseline), whereas code-switching contributes minimally. Other cross-lingual tasks remain largely unaffected by either type. These findings reveal that translation critically depends on systematic token-level alignments from parallel data, whereas cross-lingual understanding and reasoning appear to be achievable even without bilingual data.
2025
Multilingual Language Model Pretraining using Machine-translated Data
Jiayi Wang | Yao Lu | Maurice Weber | Max Ryabinin | David Ifeoluwa Adelani | Yihong Chen | Raphael Tang | Pontus Stenetorp
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Jiayi Wang | Yao Lu | Maurice Weber | Max Ryabinin | David Ifeoluwa Adelani | Yihong Chen | Raphael Tang | Pontus Stenetorp
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
English, as a very high-resource language, enables the pretraining of high-quality large language models (LLMs). However, the same can not be said for most other languages, likely due to a gap in the quality and diversity of available multilingual pretraining corpora. In this work, we find that documents machine-translated from a high-quality English corpus, can contribute significantly to the pretraining quality of multilingual LLMs. Concretely, we translate FineWeb-Edu, a high-quality English web corpus, into nine languages. resulting in a 1.7-trillion-token corpus, which we call TransWebEdu and pretrain a 1.3B-parameter model, TransWebLLM, from scratch on this corpus. Across Non-English understanding and reasoning tasks, we show that TransWebLLM matches or even outperforms multilingual LLMs of similar size, including Llama3.2, Qwen2.5, and Gemma3, despite being trained on an order of magnitude less data. Moreover, we show that adding fewer than 5% of TransWebLLM’s training tokens as domain-specific data for continued pretraining yields state-of-the-art results in Arabic, Indonesian, Swahili, and Welsh for understanding and commonsense reasoning tasks. To promote reproducibility, we release our corpus and models under Open Source Initiative-approved licenses.