MTDiag: A Multi-Turn Diagnostic Dataset Towards Clinically Meaningful LLM Evaluation

Pia Chouayfati, Alexander M. Fichtl, Miriam Anschütz, George Doumat, Georg Groh


Abstract
Clinical diagnosis is fundamentally interactive and incremental, yet the dominant paradigm for evaluating Large Language Models (LLMs) in medicine remains static QA benchmarks or template-based dialogues. These benchmarks say little about whether a model can serve as a diagnostic agent in a dynamic clinical encounter, with LLMs showing significant accuracy and reliability degradation in multi-turn settings. To address this issue, we present MTDiag, a large multi-turn diagnostic dialogue dataset constructed from three heterogeneous sources: DDXPlus, MIMIC-IV, and published case reports (AJCR), covering common ED presentations as well as long-tail rare and atypical conditions. All cases are normalized into a canonical PatientVector schema anchored in the most comprehensive and widely-adopted medical knowledge bases (UMLS concept identifiers, with ICD-10 diagnosis codes). We release the PatientVector schema, a UserLM-8B-based utterance-generation pipeline, and the physician-validated dataset that converts structured clinical evidence into natural-language utterances. Importantly, we introduce and motivate clinical knowledge-grounded metrics for evaluating LLMs as diagnostic agents, beyond diagnostic accuracy, for the task of multi-turn differential diagnosis.
Anthology ID:
2026.sigdial-1.58
Volume:
Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Month:
August
Year:
2026
Address:
Atlanta, Georgia, USA
Editors:
Jinho D. Choi, Yun-Nung Chen, Kotaro Funakoshi, Ali Emami
Venue:
SIGDIAL
SIG:
SIGDIAL
Publisher:
Association for Computational Linguistics
Note:
Pages:
821–840
Language:
URL:
https://aclanthology.org/2026.sigdial-1.58/
DOI:
Bibkey:
Cite (ACL):
Pia Chouayfati, Alexander M. Fichtl, Miriam Anschütz, George Doumat, and Georg Groh. 2026. MTDiag: A Multi-Turn Diagnostic Dataset Towards Clinically Meaningful LLM Evaluation. In Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 821–840, Atlanta, Georgia, USA. Association for Computational Linguistics.
Cite (Informal):
MTDiag: A Multi-Turn Diagnostic Dataset Towards Clinically Meaningful LLM Evaluation (Chouayfati et al., SIGDIAL 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.sigdial-1.58.pdf