Validating LLM-Rated Item Features for Explanatory Item Response Models

Ahmed H. Bediwy, Mubarak O. Mojoyinola


Abstract
Explanatory item response models let test de- velopers anticipate item difficulty from item design, but they depend on subject-matter ex- perts rating every item on every hypothesized feature—a step that is slow, costly, and the practical bottleneck limiting how many items can be modeled. We ask whether large lan- guage models can supply those ratings. Four trained experts and six LLMs independently rated 15 grade-5 mathematics items on six psy- chometric features. We evaluate the LLM rat- ings twice: against the human consensus using quadratic-weighted 𝜅and mixed-effects mod- els, and against 1,452 student responses by us- ing each source’s ratings as the design matrix of a linear logistic test model (LLTM) bench- marked against a Rasch baseline. Claude Opus 4.7 performed best in recovering the item dif- ficulty (r = 0.89), however, the human con- sensus against which agreement is measured performs worst at recovering Rasch difficulty (r = 0.33) compared to the rest of the LLMs. The reason is visible in the expert panel it- self: inter-rater Fleiss’ 𝜅is at or below chance on three of six features, so the consensus is a noisy reference rather than ground truth. We also identify two failure modes (zero-variance and perfect collinearity) that agreement statis- tics cannot detect but make a feature unusable as an LLTM covariate. We argue that agree- ment should be reported alongside criterion- referenced recovery of item parameters, not in place of it.
Anthology ID:
2026.aimecon-wip.49
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
382–388
Language:
URL:
https://aclanthology.org/2026.aimecon-wip.49/
DOI:
Bibkey:
Cite (ACL):
Ahmed H. Bediwy and Mubarak O. Mojoyinola. 2026. Validating LLM-Rated Item Features for Explanatory Item Response Models. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress, pages 382–388, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
Validating LLM-Rated Item Features for Explanatory Item Response Models (Bediwy & Mojoyinola, AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-wip.49.pdf