Nexus: Adaptive Upcycling to Efficiently Pretrain Mixture of Experts

Nikolas Gritsch; Qizhen Zhang; Acyr Locatelli; Sara Hooker; Ahmet Üstün

doi:10.18653/v1/2025.findings-emnlp.1323

Nexus: Adaptive Upcycling to Efficiently Pretrain Mixture of Experts

Nikolas Gritsch, Qizhen Zhang, Acyr Locatelli, Sara Hooker, Ahmet Üstün

Abstract

Frontier language models are increasingly based on the Mixture of Experts (MoE) architecture, boosting the efficiency of training and inference by sparsely activating parameters. Nevertheless, training from scratch on trillions of tokens remains so expensive that most users can only finetune these models. In this work, we combine parameter reuse of dense models for the MoE layers ("*upcycling*”) with a novel, *adaptive* Nexus router that can integrate new experts into an existing trained model without hurting the performance on previous domains. Our router leverages the knowledge of each expert’s training data distribution via domain embeddings to initialize the router, improving specialization and allowing it to adapt faster to new domains than a standard MoE router. Nexus overturns the strict sequential separation between training and finetuning in classical approaches, allowing more powerful improvements to existing models at a later stage through long token-horizon trainings on new pretraining data. Our experiments show that Nexus achieves a relative gain of up to 2.1% over the baseline for initial upcycling, and an 18.8% relative gain for extending the MoE to a new domain with a new expert by using limited finetuning data. This flexibility of Nexus can power an open-source ecosystem where every user continuously assembles their own MoE-mix from a multitude of dense models.

Anthology ID:: 2025.findings-emnlp.1323
Volume:: Findings of the Association for Computational Linguistics: EMNLP 2025
Month:: November
Year:: 2025
Address:: Suzhou, China
Editors:: Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 24364–24381
Language:
URL:: https://aclanthology.org/2025.findings-emnlp.1323/
DOI:: 10.18653/v1/2025.findings-emnlp.1323
Bibkey:
Cite (ACL):: Nikolas Gritsch, Qizhen Zhang, Acyr Locatelli, Sara Hooker, and Ahmet Üstün. 2025. Nexus: Adaptive Upcycling to Efficiently Pretrain Mixture of Experts. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 24364–24381, Suzhou, China. Association for Computational Linguistics.
Cite (Informal):: Nexus: Adaptive Upcycling to Efficiently Pretrain Mixture of Experts (Gritsch et al., Findings 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.findings-emnlp.1323.pdf
Checklist:: 2025.findings-emnlp.1323.checklist.pdf

PDF Cite Search Checklist Fix data