MIRAGE: Exploring How Large Language Models Perform in Complex Social Interactive Environments

Yin Cai; Zhouhong Gu; Zhaohan Du; Zheyu Ye; Shaosheng Cao; Yiqian Xu; Hongwei Feng; Ping Chen

doi:10.18653/v1/2025.acl-short.2

MIRAGE: Exploring How Large Language Models Perform in Complex Social Interactive Environments

Yin Cai, Zhouhong Gu, Zhaohan Du, Zheyu Ye, Shaosheng Cao, Yiqian Xu, Hongwei Feng, Ping Chen

Abstract

Large Language Models (LLMs) have shown remarkable capabilities in environmental perception, reasoning-based decision-making, and simulating complex human behaviors, particularly in interactive role-playing contexts. This paper introduces the Multiverse Interactive Role-play Ability General Evaluation (MIRAGE), a comprehensive framework designed to assess LLMs’ proficiency in portraying advanced human behaviors through murder mystery games. MIRAGE features eight intricately crafted scripts encompassing diverse themes and styles, providing a rich simulation. To evaluate LLMs’ performance, MIRAGE employs four distinct methods: the Trust Inclination Index (TII) to measure dynamics of trust and suspicion, the Clue Investigation Capability (CIC) to measure LLMs’ capability of conducting information, the Interactivity Capability Index (ICI) to assess role-playing capabilities and the Script Compliance Index (SCI) to assess LLMs’ capability of understanding and following instructions. Our experiments indicate that even popular models like GPT-4 face significant challenges in navigating the complexities presented by the MIRAGE. The datasets and simulation codes are available in https://github.com/lime728/MIRAGE.

Anthology ID:: 2025.acl-short.2
Volume:: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)
Month:: July
Year:: 2025
Address:: Vienna, Austria
Editors:: Wanxiang Che, Joyce Nabende, Ekaterina Shutova, Mohammad Taher Pilehvar
Venue:: ACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 14–40
Language:
URL:: https://aclanthology.org/2025.acl-short.2/
DOI:: 10.18653/v1/2025.acl-short.2
Bibkey:
Cite (ACL):: Yin Cai, Zhouhong Gu, Zhaohan Du, Zheyu Ye, Shaosheng Cao, Yiqian Xu, Hongwei Feng, and Ping Chen. 2025. MIRAGE: Exploring How Large Language Models Perform in Complex Social Interactive Environments. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 14–40, Vienna, Austria. Association for Computational Linguistics.
Cite (Informal):: MIRAGE: Exploring How Large Language Models Perform in Complex Social Interactive Environments (Cai et al., ACL 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.acl-short.2.pdf

PDF Cite Search Fix data