On the Same Wavelength? Evaluating Pragmatic Reasoning in Language Models across Broad Concepts

Linlu Qiu; Cedegao E. Zhang; Joshua B. Tenenbaum; Yoon Kim; Roger Levy

doi:10.18653/v1/2025.emnlp-main.1008

On the Same Wavelength? Evaluating Pragmatic Reasoning in Language Models across Broad Concepts

Linlu Qiu, Cedegao E. Zhang, Joshua B. Tenenbaum, Yoon Kim, Roger P. Levy

Abstract

Language use is shaped by pragmatics—i.e., reasoning about communicative goals and norms in context. As language models (LMs) are increasingly used as conversational agents, it becomes ever more important to understand their pragmatic reasoning abilities. We propose an evaluation framework derived from *Wavelength*, a popular communication game where a speaker and a listener communicate about a broad range of concepts in a granular manner. We study a range of LMs on both language comprehension and language production using direct and Chain-of-Thought (CoT) prompting, and further explore a Rational Speech Act (RSA) approach to incorporating Bayesian pragmatic reasoning into LM inference. We find that state-of-the-art LMs, but not smaller ones, achieve strong performance on language comprehension, obtaining similar-to-human accuracy and exhibiting high correlations with human judgments even without CoT prompting or RSA. On language production, CoT can outperform direct prompting, and using RSA provides significant improvements over both approaches. Our study helps identify the strengths and limitations in LMs’ pragmatic reasoning abilities and demonstrates the potential for improving them with RSA, opening up future avenues for understanding conceptual representation, language understanding, and social reasoning in LMs and humans.

Anthology ID:: 2025.emnlp-main.1008
Volume:: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Month:: November
Year:: 2025
Address:: Suzhou, China
Editors:: Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng
Venue:: EMNLP
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 19913–19935
Language:
URL:: https://aclanthology.org/2025.emnlp-main.1008/
DOI:: 10.18653/v1/2025.emnlp-main.1008
Bibkey:
Cite (ACL):: Linlu Qiu, Cedegao E. Zhang, Joshua B. Tenenbaum, Yoon Kim, and Roger P. Levy. 2025. On the Same Wavelength? Evaluating Pragmatic Reasoning in Language Models across Broad Concepts. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 19913–19935, Suzhou, China. Association for Computational Linguistics.
Cite (Informal):: On the Same Wavelength? Evaluating Pragmatic Reasoning in Language Models across Broad Concepts (Qiu et al., EMNLP 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.emnlp-main.1008.pdf
Checklist:: 2025.emnlp-main.1008.checklist.pdf

PDF Cite Search Checklist Fix data