Multi-Scale Model Compression via Nested Matrix Learning

Xiangjue Dong, Aditya Anantharaman, Hemant Pugaliya, Kai Zhong


Abstract
Large language models (LLMs) have been widely deployed and have achieved remarkable success in downstream tasks. However, their high latency continues to pose challenges for real-time applications that require fast inference, and the need to train and deploy distinct models for different hardware constraints increases both financial and computational costs. To address this, we propose Nested Matrix Learning (NML), a method that trains a single, flexible model capable of generating multiple high-performing student models of varying sizes. This is achieved by simultaneously optimizing a pre-trained teacher model and its nested sub-models in a single training process, without sacrificing the teacher’s performance. NML provides a flexible and scalable solution, allowing models to adapt to different computational budgets. Our extensive experiments show that student models produced by NML, which can be up to 10x smaller than the full-size model, can be directly deployed for efficient inference or serve as superior initialization points for further fine-tuning in downstream tasks. By preserving the performance of the teacher model while delivering compact and efficient student models of various sizes, NML enhances the usability and adaptability of LLMs in real-world scenarios.
Anthology ID:
2026.lrec-1.196
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
2501–2511
Language:
External URL:
https://lrec.elra.info/lrec2026-main-196
DOI:
10.63317/5o97c4anqod5
Bibkey:
Cite (ACL):
Xiangjue Dong, Aditya Anantharaman, Hemant Pugaliya, and Kai Zhong. 2026. Multi-Scale Model Compression via Nested Matrix Learning. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 2501–2511, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Multi-Scale Model Compression via Nested Matrix Learning (Dong et al., LREC 2026)
Copy Citation: