Jamie Taylor


2025

pdf bib
The More, The Better? A Critical Study of Multimodal Context in Radiology Report Summarization
Mong Yuan Sim | Wei Emma Zhang | Xiang Dai | Biaoyan Fang | Sarbin Ranjitkar | Arjun Burlakoti | Jamie Taylor | Haojie Zhuang
Findings of the Association for Computational Linguistics: EMNLP 2025

The Impression section of a radiology report summarizes critical findings of a radiology report and thus plays a crucial role in communication between radiologists and physicians. Research on radiology report summarization mostly focuses on generating the Impression section by summarizing information from the Findings section, which typically details the radiologist’s observations in the radiology images. Recent work start to explore how to incorporate radiology images as input to multimodal summarization models, with the assumption that it can improve generated summary quality, as it contains richer information. However, the real effectiveness of radiology images remains unclear. To answer this, we conduct a thorough analysis to understand whether current multimodal models can utilize radiology images in summarizing Findings section. Our analysis reveals that current multimodal models often fail to effectively utilize radiology images. For example, masking the image input leads to minimal or no performance drop. Expert annotation study shows that radiology images are unnecessary when they write the Impression section.