Low-Rank Compression of Language Models via Differentiable Rank Selection

Sidhant Sundrani, Francesco Tudisco, Pasquale Minervini


Abstract
Approaches for compressing large-language models using low-rank decomposition have made strides, particularly with the introduction of activation and loss-aware SVD, which improves the trade-off between decomposition rank and downstream task performance. Despite these advancements, a persistent challenge remains–selecting the optimal ranks for each layer to jointly optimise compression rate and downstream task accuracy. Current methods either rely on heuristics that can yield sub-optimal results due to their limited discrete search space or are gradient-based but are not as performant as heuristic approaches without post-compression fine-tuning. To address these issues, we propose Learning to Low-Rank Compress (LLRC), a gradient-based approach that directly learns the weights of masks that select singular values in a fine-tuning-free setting. Using a calibration dataset, we train only the mask weights to select fewer and fewer singular values while minimising the divergence of intermediate activations from the original model. Our approach outperforms competing methods that similarly require no post-compression fine-tuning across various compression rates on common-sense reasoning and open-domain question-answering tasks. For instance, with a compression rate of 20% on Llama-2-13B, LLRC outperforms the competitive Sensitivity-based Truncation Rank Searching (STRS) on MMLU, BoolQ, and OpenbookQA by 12%, 3.5%, and 4.4%, respectively. Compared to other compression techniques, our approach consistently outperforms fine-tuning-free variants of SVD-LLM and LLM-Pruner across datasets and compression rates. Our approach also performs competitively with LLM-Pruner after fine-tuning on Llama-2-7B and Llama-2-13B.
Anthology ID:
2026.lrec-1.787
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
10031–10045
Language:
External URL:
https://lrec.elra.info/lrec2026-main-787
DOI:
10.63317/2xbs948bhby9
Bibkey:
Cite (ACL):
Sidhant Sundrani, Francesco Tudisco, and Pasquale Minervini. 2026. Low-Rank Compression of Language Models via Differentiable Rank Selection. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 10031–10045, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Low-Rank Compression of Language Models via Differentiable Rank Selection (Sundrani et al., LREC 2026)
Copy Citation: