David Shin
Author directory2026
Assessing Large Language Model Performance in Post-Certification-Examination Comment Categorization
Huaping Sun | Kristin O’Brien | Colleen Burke Kave | David Shin | Qiao Lin | Jeffrey Marc Lyness
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Huaping Sun | Kristin O’Brien | Colleen Burke Kave | David Shin | Qiao Lin | Jeffrey Marc Lyness
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
This study evaluated an LLM for analyzing 1,406 Neurocritical Care examination comments coded for sentiment and thematic categories. Human raters showed high agreement, whereas LLM-human agreement was moderate. Thematic definitions reduced performance, but human-coded examples improved accuracy. LLMs may support preliminary coding, although human review remains necessary.