Evaluating Multiple Models for Predicting Item Difficulty in a Principled Assessment Design

Alexandra Lane Perez, Christina Schneider, Sangdon Lim, Garron Gianopulos, Kang Xue


Abstract
Item difficulty should, in theory, be predictable from features derived from RPLDs, task characteristics, and linguistic complexity. This study evaluates the performance of multiple statistical and machine learning models in estimating item difficulty. We found similar results across all four models, the item features selected explain 53%–56% of the variance across all grades.
Anthology ID:
2026.aimecon-wip.20
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
152–158
Language:
URL:
https://aclanthology.org/2026.aimecon-wip.20/
DOI:
Bibkey:
Cite (ACL):
Alexandra Lane Perez, Christina Schneider, Sangdon Lim, Garron Gianopulos, and Kang Xue. 2026. Evaluating Multiple Models for Predicting Item Difficulty in a Principled Assessment Design. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress, pages 152–158, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
Evaluating Multiple Models for Predicting Item Difficulty in a Principled Assessment Design (Perez et al., AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-wip.20.pdf