Reyhaneh Hosseinpourkhoshkbari

Author directory

2026

Large language models (LLMs) are increasingly used to score clinical communication, but validation is often limited to individual checklist items or total scores. These measures may not capture how communication behaviors occur together within a transcript. We therefore examine communication profiles: recurring combinations of behaviors that characterize different patterns of clinician communication and may support more targeted formative feedback. We analyzed 213 simulated respiratory OSCE transcripts rated on 15 binary Kalamazoo-derived communication items and compared final human ratings with GPT-4o, GPT-o3, and GPT-5.6 Sol. Raw item-level agreement was relatively high overall, but varied substantially across behaviors and was lower for several judgment-intensive items used for profile modeling. Bayesian latent-class analysis identified three stable human-derived profiles. Although the LLM-derived profiles showed broadly similar item-probability patterns, the models frequently assigned individual transcripts to different profiles than the human ratings. These findings show that agreement on individual communication skills does not necessarily translate into agreement on higher-level communication profiles. If LLMs are used to provide profile-based feedback, validation should therefore include profile-level agreement in addition to item-level performance.