Can LLMs Understand Punchlines? LLMs’ Narrative Understanding Evaluation with Short-shorts

Jiashi Cheng, Takehito Utsuro


Abstract
In this study, we constructed a narrative comprehension benchmark using the works of Shinichi Hoshi to examine the extent to which Large Language Models (LLMs) can understand twist endings, or punchlines, in short-short stories. Specifically, story endings were categorized into six types—such as Revelation, Apocalypse, and Sarcasm—and a classification task was designed in which LLMs were prompted with the story text and asked to select the appropriate ending category. We collected human annotations from eight native Japanese speakers to establish a reference benchmark. Experimental comparisons were conducted across multiple LLMs (GPT-4, Claude, Gemini, and Grok), assessing their performance both at the metric level and at the discourse level against human judgments. The results revealed that although certain models approached human performance in specific categories, overall accuracy remained notably lower than the human baseline. Through quantitative and qualitative analyses, this study highlights the challenges LLMs face in capturing narrative subtleties such as irony, implication, and emotional reversal. The proposed benchmark provides a novel framework for evaluating narrative understanding and the deeper semantic reasoning capabilities of LLMs.
Anthology ID:
2026.lrec-1.159
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
2024–2034
Language:
External URL:
https://lrec.elra.info/lrec2026-main-159
DOI:
10.63317/4n2p36736i24
Bibkey:
Cite (ACL):
Jiashi Cheng and Takehito Utsuro. 2026. Can LLMs Understand Punchlines? LLMs’ Narrative Understanding Evaluation with Short-shorts. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 2024–2034, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Can LLMs Understand Punchlines? LLMs’ Narrative Understanding Evaluation with Short-shorts (Cheng & Utsuro, LREC 2026)
Copy Citation: