AttnComp: Attention-Guided Adaptive Context Compression for Retrieval-Augmented Generation

Lvzhou Luo; Yixuan Cao; Ping Luo

doi:10.18653/v1/2025.findings-emnlp.449

AttnComp: Attention-Guided Adaptive Context Compression for Retrieval-Augmented Generation

Abstract

Retrieval-augmented generation improves the factual accuracy of Large Language Models (LLMs) by incorporating external context, but often suffers from irrelevant retrieved content that hinders effectiveness. Context compression addresses this issue by filtering out irrelevant information from context before LLM generation. However, existing methods struggle to adaptively adjust compression rates for different context, maintain low latency and integrate information across multiple documents. To overcome these limitations, We introduce AttnComp, an adaptive, efficient and context-aware compression framework. By leveraging the attention mechanism of LLMs to identify relevant information, AttnComp employs a Top-P compression algorithm to retain the minimal set of documents whose cumulative attention weights exceeds a predefined threshold. In addition to compression, AttnComp estimates response confidence by assessing the overall relevance of the retrieved content, enabling users to gauge response reliability. Experiments demonstrate that AttnComp outperforms existing compression methods and uncompressed baselines, achieving higher accuracy with substantial compression rates and lower latency.

Anthology ID:: 2025.findings-emnlp.449
Volume:: Findings of the Association for Computational Linguistics: EMNLP 2025
Month:: November
Year:: 2025
Address:: Suzhou, China
Editors:: Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 8456–8472
Language:
URL:: https://aclanthology.org/2025.findings-emnlp.449/
DOI:: 10.18653/v1/2025.findings-emnlp.449
Bibkey:
Cite (ACL):: Lvzhou Luo, Yixuan Cao, and Ping Luo. 2025. AttnComp: Attention-Guided Adaptive Context Compression for Retrieval-Augmented Generation. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 8456–8472, Suzhou, China. Association for Computational Linguistics.
Cite (Informal):: AttnComp: Attention-Guided Adaptive Context Compression for Retrieval-Augmented Generation (Luo et al., Findings 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.findings-emnlp.449.pdf
Checklist:: 2025.findings-emnlp.449.checklist.pdf

PDF Cite Search Checklist Fix data