Minkwon Kim

Author directory

2026

Regenerating LLM Q-matrices ten times per model on TIMSS 2011 items, we find two runs of the same model reassign 61–88% of student mastery profiles. Greedy decoding removes this instability; disagreement with expert judgment (63–76%) survives. Majority voting fixes neither. Report distributions, not single runs.
Search
Co-authors
    Venues
    Fix author