SFMSS: Service Flow aware Medical Scenario Simulation for Conversational Data Generation

Zhijie Bao; Qingyun Liu; Xuan-Jing Huang (黄萱菁); Zhongyu Wei

doi:10.18653/v1/2025.findings-naacl.259

SFMSS: Service Flow aware Medical Scenario Simulation for Conversational Data Generation

Zhijie Bao, Qingyun Liu, Xuanjing Huang, Zhongyu Wei

Abstract

Medical-specific Large Language Models (LLMs) have demonstrated impressive performance on medical-related exams and tasks. Despite their success in single-turn question and answering, instruction-tuned LLMs often falter in real-world healthcare applications, highlighting a disconnect between existing instruction datasets and practical contexts. To address this issue, we propose Service Flow aware Medical Scenario Simulation (SFMSS), a simulation framework designed for medical conversational data generation. SFMSS employs three key strategies to ensure the quality of the data generation. the use of Authentic Seed Data ensures alignment of real-world distributions. Diverse Patient Simulation enables simulated patients to exhibit distinct communication styles and complex behavioral logic. Service Flow Control ensures that conversations progress in alignment with medical objectives. We construct a dataset targeting on outpatient reception through SFMSS, named SFMSS-CD. Building on this dataset, we develop a model called SFMSS-Nurse. We conduct both automatic and human evaluations, involving 15 users and 15 clinical experts, to assess the effectiveness of SFMSS. The results demonstrate that SFMSS-Nurse outperforms all baselines, including the current state-of-the-art model GPT-4o, and aligns with human preferences and clinical demands.

Anthology ID:: 2025.findings-naacl.259
Volume:: Findings of the Association for Computational Linguistics: NAACL 2025
Month:: April
Year:: 2025
Address:: Albuquerque, New Mexico
Editors:: Luis Chiruzzo, Alan Ritter, Lu Wang
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 4586–4604
Language:
URL:: https://aclanthology.org/2025.findings-naacl.259/
DOI:: 10.18653/v1/2025.findings-naacl.259
Bibkey:
Cite (ACL):: Zhijie Bao, Qingyun Liu, Xuanjing Huang, and Zhongyu Wei. 2025. SFMSS: Service Flow aware Medical Scenario Simulation for Conversational Data Generation. In Findings of the Association for Computational Linguistics: NAACL 2025, pages 4586–4604, Albuquerque, New Mexico. Association for Computational Linguistics.
Cite (Informal):: SFMSS: Service Flow aware Medical Scenario Simulation for Conversational Data Generation (Bao et al., Findings 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.findings-naacl.259.pdf

PDF Cite Search Fix data