Dogus Darici

Author directory

2026

Comparing LLM-human rater rationales using semantic similarity risks conflating textual proximity with evaluative agreement. We test whether embedding-based similarity reflects qualitative coding distinctions across rationale pairs in medical education. Similarity declined as rationale length differences grew and was less effective at distinguishing whether LLMs preserved the human’s central claim.

2025

We examined how model size, temperature, and prompt style affect Large Language Models’ (LLMs) alignment with human raters in assessing clinical reasoning skills. Model size emerged as a key factor in LLM-human score alignment. Findings reveal both the potential for scalable LLM-raters and the risks of relying on them exclusively.