David Shin

Author directory

2026

This study evaluated an LLM for analyzing 1,406 Neurocritical Care examination comments coded for sentiment and thematic categories. Human raters showed high agreement, whereas LLM-human agreement was moderate. Thematic definitions reduced performance, but human-coded examples improved accuracy. LLMs may support preliminary coding, although human review remains necessary.