LLM Compression: How Far Can We Go in Balancing Size and Performance?

Sahil Sk; Debashish Dhal; Sonal Khosla; Akash Dhaka; Shantipriya Parida; Sk Shahid; Sambit Shekhar; Dilip Prasad; Ondřej Bojar

LLM Compression: How Far Can We Go in Balancing Size and Performance?

Sahil Sk, Debashish Dhal, Sonal Khosla, Akash Dhaka, Shantipriya Parida, Sk Shahid, Sambit Shekhar, Dilip Prasad, Ondrej Bojar

Abstract

Quantization is an essential and popular technique for improving the accessibility of large language models (LLMs) by reducing memory usage and computational costs while maintaining performance. In this study, we apply 4-bit Group Scaling Quantization (GSQ) and Generative Pretrained Transformer Quantization (GPTQ) to LLaMA 1B, Qwen 0.5B, and PHI 1.5B, evaluating their impact across multiple NLP tasks. We benchmark these models on MS MARCO (Information Retrieval), BoolQ (Boolean Question Answering), and GSM8K (Mathematical Reasoning) datasets, assessing both accuracy and efficiency accross various tasks. The study measures the trade-offs between model compression and task performance, analyzing key evaluation metrics namely: accuracy, inference latency, and throughput, providing insights into the suitability of low-bit quantization for real-world deployment and highlight the tradeoffs between memory, computing and latency in such settings, helping a user make suitable decisions

Anthology ID:: 2025.ranlp-1.136
Volume:: Proceedings of the 15th International Conference on Recent Advances in Natural Language Processing - Natural Language Processing in the Generative AI Era
Month:: September
Year:: 2025
Address:: Varna, Bulgaria
Editors:: Galia Angelova, Maria Kunilovskaya, Marie Escribe, Ruslan Mitkov
Venue:: RANLP
SIG:
Publisher:: INCOMA Ltd., Shoumen, Bulgaria
Note:
Pages:: 1183–1187
Language:
URL:: https://aclanthology.org/2025.ranlp-1.136/
DOI:
Bibkey:
Cite (ACL):: Sahil Sk, Debashish Dhal, Sonal Khosla, Akash Dhaka, Shantipriya Parida, Sk Shahid, Sambit Shekhar, Dilip Prasad, and Ondrej Bojar. 2025. LLM Compression: How Far Can We Go in Balancing Size and Performance?. In Proceedings of the 15th International Conference on Recent Advances in Natural Language Processing - Natural Language Processing in the Generative AI Era, pages 1183–1187, Varna, Bulgaria. INCOMA Ltd., Shoumen, Bulgaria.
Cite (Informal):: LLM Compression: How Far Can We Go in Balancing Size and Performance? (Sk et al., RANLP 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.ranlp-1.136.pdf

PDF Cite Search Fix data