SUAT-BMI at MEDIQA-EVAL 2026: An Ensemble Approach to Language Models as Judges for Automatic Rating of Medical Responses

Xinzhe Peng, Liyuan E, Kun Feng, Jielin Li, Yuxuan Tang, Zhao Li


Abstract
The MEDIQA-EVAL 2026 shared task focuses on developing automatic evaluation metrics for LLM-generated responses in dermatology and wound care. While LLMs have shown promise as judge models, the reliability of these metrics remains underexplored. In this work, we study how well judge models can approximate human expert ratings across clinical evaluation criteria. We evaluate multiple approaches, including few-shot prompting, BERT fine-tuning, and retrieval-augmented generation (RAG), and combine them in an ensemble framework. Our method achieves a correlation score of 0.481, ranking first among 41 participating teams. Our results provide insight into the reliability of LLM-based evaluation metrics and highlight their potential for scalable clinical assessment.
Anthology ID:
2026.clinicalnlp-1.2
Volume:
Proceedings of the 8th Workshop on Clinical Natural Language Processing (Clinical NLP) @ LREC 2026
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Asma Ben Abacha, Steven Bethard, Danielle Bitterman, Tristan Naumann, Kirk Roberts
Venues:
ClinicalNLP | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
12–18
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-clinicalnlp-02
DOI:
10.63317/2kdt525sk8is
Bibkey:
Cite (ACL):
Xinzhe Peng, Liyuan E, Kun Feng, Jielin Li, Yuxuan Tang, and Zhao Li. 2026. SUAT-BMI at MEDIQA-EVAL 2026: An Ensemble Approach to Language Models as Judges for Automatic Rating of Medical Responses. In Proceedings of the 8th Workshop on Clinical Natural Language Processing (Clinical NLP) @ LREC 2026, pages 12–18, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
SUAT-BMI at MEDIQA-EVAL 2026: An Ensemble Approach to Language Models as Judges for Automatic Rating of Medical Responses (Peng et al., ClinicalNLP 2026)
Copy Citation: