MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents

Haoran Tan; Zeyu Zhang; Chen Ma; Xu Chen; Quanyu Dai; Zhenhua Dong

doi:10.18653/v1/2025.findings-acl.989

MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents

Haoran Tan, Zeyu Zhang, Chen Ma, Xu Chen, Quanyu Dai, Zhenhua Dong

Abstract

Recent works have highlighted the significance of memory mechanisms in LLM-based agents, which enable them to store observed information and adapt to dynamic environments. However, evaluating their memory capabilities still remains challenges. Previous evaluations are commonly limited by the diversity of memory levels and interactive scenarios. They also lack comprehensive metrics to reflect the memory capabilities from multiple aspects. To address these problems, in this paper, we construct a more comprehensive dataset and benchmark to evaluate the memory capability of LLM-based agents. Our dataset incorporates factual memory and reflective memory as different levels, and proposes participation and observation as various interactive scenarios. Based on our dataset, we present a benchmark, named MemBench, to evaluate the memory capability of LLM-based agents from multiple aspects, including their effectiveness, efficiency, and capacity. To benefit the research community, we release our dataset and project at https://github.com/import-myself/Membench.

Anthology ID:: 2025.findings-acl.989
Volume:: Findings of the Association for Computational Linguistics: ACL 2025
Month:: July
Year:: 2025
Address:: Vienna, Austria
Editors:: Wanxiang Che, Joyce Nabende, Ekaterina Shutova, Mohammad Taher Pilehvar
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 19336–19352
Language:
URL:: https://aclanthology.org/2025.findings-acl.989/
DOI:: 10.18653/v1/2025.findings-acl.989
Bibkey:
Cite (ACL):: Haoran Tan, Zeyu Zhang, Chen Ma, Xu Chen, Quanyu Dai, and Zhenhua Dong. 2025. MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents. In Findings of the Association for Computational Linguistics: ACL 2025, pages 19336–19352, Vienna, Austria. Association for Computational Linguistics.
Cite (Informal):: MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents (Tan et al., Findings 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.findings-acl.989.pdf

PDF Cite Search Fix data