hgkai26 at MEDIQA-EVAL 2026: Automated Evaluation of Visual Medical Question Answering Using LLM-as-a-Judge

Haritha Gangavarapu


Abstract
As there is a rise in the use of multimodal large language models (LLMs) for medical response generation, it is necessary to have reliable automated evaluation mechanisms that can assess the quality of model-generated outputs. The MediQA-Eval 2026 shared task focuses on grading AI-generated dermatology and wound care responses using structured human-aligned rubrics. In this work, we explore a zero-shot multimodal LLM-as-a-Judge framework to assess candidate responses across multiple quality dimensions. System performance is evaluated using the official task metrics designed to reflect alignment with human judgments. Our findings provide preliminary insights into the feasibility and limitations of LLM-based evaluators for rubric-guided medical response assessment.
Anthology ID:
2026.clinicalnlp-1.29
Volume:
Proceedings of the 8th Workshop on Clinical Natural Language Processing (Clinical NLP) @ LREC 2026
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Asma Ben Abacha, Steven Bethard, Danielle Bitterman, Tristan Naumann, Kirk Roberts
Venues:
ClinicalNLP | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
257–261
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-clinicalnlp-29
DOI:
10.63317/4n9skmf9rive
Bibkey:
Cite (ACL):
Haritha Gangavarapu. 2026. hgkai26 at MEDIQA-EVAL 2026: Automated Evaluation of Visual Medical Question Answering Using LLM-as-a-Judge. In Proceedings of the 8th Workshop on Clinical Natural Language Processing (Clinical NLP) @ LREC 2026, pages 257–261, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
hgkai26 at MEDIQA-EVAL 2026: Automated Evaluation of Visual Medical Question Answering Using LLM-as-a-Judge (Gangavarapu, ClinicalNLP 2026)
Copy Citation: