LangCompress: Language-Aware Compression of Large Language Models

Dieu-Hien Nguyen; Nguyen-Khang Le; Truong Dinh Do; Minh Le Nguyen

LangCompress: Language-Aware Compression of Large Language Models

Dieu-Hien Nguyen, Nguyen-Khang Le, Truong Dinh Do, Le-Minh Nguyen

Abstract

Large Language Models (LLMs) demonstrate strong multilingual capabilities but are costly to deploy due to their size and computational demands. To mitigate this, compression techniques such as pruning and quantization are widely used. However, these methods face two key limitations: (1) they assume access to high-quality instruction or calibration data, which is often unavailable for low-resource languages; and (2) they aim to preserve multilingual generality, making them inefficient for language-specific applications. We introduce LangCompress, a language-aware compression framework that enhances existing compression methods for targeted deployment. LangCompress is method-agnostic and improves state-of-the-art pruning and quantization approaches. It features two core components: an iterative self-supervised pipeline for generating instruction data in the target language, and a vocabulary simplification strategy that reduces the LM head to focus on key tokens. Experiments on perplexity, translation, and summarization tasks show that LangCompress improves performance in the target language. The code and data are publicly available.

Anthology ID:: 2025.ijcnlp-long.112
Volume:: Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics
Month:: December
Year:: 2025
Address:: Mumbai, India
Editors:: Kentaro Inui, Sakriani Sakti, Haofen Wang, Derek F. Wong, Pushpak Bhattacharyya, Biplab Banerjee, Asif Ekbal, Tanmoy Chakraborty, Dhirendra Pratap Singh
Venues:: IJCNLP | AACL
SIG:
Publisher:: The Asian Federation of Natural Language Processing and The Association for Computational Linguistics
Note:
Pages:: 2068–2077
Language:
URL:: https://aclanthology.org/2025.ijcnlp-long.112/
DOI:
Bibkey:
Cite (ACL):: Dieu-Hien Nguyen, Nguyen-Khang Le, Truong Dinh Do, and Le-Minh Nguyen. 2025. LangCompress: Language-Aware Compression of Large Language Models. In Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics, pages 2068–2077, Mumbai, India. The Asian Federation of Natural Language Processing and The Association for Computational Linguistics.
Cite (Informal):: LangCompress: Language-Aware Compression of Large Language Models (Nguyen et al., IJCNLP-AACL 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.ijcnlp-long.112.pdf

PDF Cite Search Fix data