ERVQA: A Dataset to Benchmark the Readiness of Large Vision Language Models in Hospital Environments

Sourjyadip Ray, Kushal Gupta, Soumi Kundu, Dr Kasat, Somak Aditya, Pawan Goyal


Abstract
The global shortage of healthcare workers has demanded the development of smart healthcare assistants, which can help monitor and alert healthcare workers when necessary. We examine the healthcare knowledge of existing Large Vision Language Models (LVLMs) via the Visual Question Answering (VQA) task in hospital settings through expert annotated open-ended questions. We introduce the Emergency Room Visual Question Answering (ERVQA) dataset, consisting of <image, question, answer> triplets covering diverse emergency room scenarios, a seminal benchmark for LVLMs. By developing a detailed error taxonomy and analyzing answer trends, we reveal the nuanced nature of the task. We benchmark state-of-the-art open-source and closed LVLMs using traditional and adapted VQA metrics: Entailment Score and CLIPScore Confidence. Analyzing errors across models, we infer trends based on properties like decoder type, model size, and in-context examples. Our findings suggest the ERVQA dataset presents a highly complex task, highlighting the need for specialized, domain-specific solutions.
Anthology ID:
2024.emnlp-main.873
Volume:
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
Month:
November
Year:
2024
Address:
Miami, Florida, USA
Editors:
Yaser Al-Onaizan, Mohit Bansal, Yun-Nung Chen
Venue:
EMNLP
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
15594–15608
Language:
URL:
https://aclanthology.org/2024.emnlp-main.873
DOI:
Bibkey:
Cite (ACL):
Sourjyadip Ray, Kushal Gupta, Soumi Kundu, Dr Kasat, Somak Aditya, and Pawan Goyal. 2024. ERVQA: A Dataset to Benchmark the Readiness of Large Vision Language Models in Hospital Environments. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 15594–15608, Miami, Florida, USA. Association for Computational Linguistics.
Cite (Informal):
ERVQA: A Dataset to Benchmark the Readiness of Large Vision Language Models in Hospital Environments (Ray et al., EMNLP 2024)
Copy Citation:
PDF:
https://aclanthology.org/2024.emnlp-main.873.pdf