Mubarak O. Mojoyinola
Author directory2026
Validating LLM-Rated Item Features for Explanatory Item Response Models
Ahmed H. Bediwy | Mubarak O. Mojoyinola
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Ahmed H. Bediwy | Mubarak O. Mojoyinola
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Explanatory item response models let test de- velopers anticipate item difficulty from item design, but they depend on subject-matter ex- perts rating every item on every hypothesized feature—a step that is slow, costly, and the practical bottleneck limiting how many items can be modeled. We ask whether large lan- guage models can supply those ratings. Four trained experts and six LLMs independently rated 15 grade-5 mathematics items on six psy- chometric features. We evaluate the LLM rat- ings twice: against the human consensus using quadratic-weighted 𝜅and mixed-effects mod- els, and against 1,452 student responses by us- ing each source’s ratings as the design matrix of a linear logistic test model (LLTM) bench- marked against a Rasch baseline. Claude Opus 4.7 performed best in recovering the item dif- ficulty (r = 0.89), however, the human con- sensus against which agreement is measured performs worst at recovering Rasch difficulty (r = 0.33) compared to the rest of the LLMs. The reason is visible in the expert panel it- self: inter-rater Fleiss’ 𝜅is at or below chance on three of six features, so the consensus is a noisy reference rather than ground truth. We also identify two failure modes (zero-variance and perfect collinearity) that agreement statis- tics cannot detect but make a feature unusable as an LLTM covariate. We argue that agree- ment should be reported alongside criterion- referenced recovery of item parameters, not in place of it.