Predicting Item-to-Range Performance Level Descriptor Matches with Structured Language Model Reasoning

Kang Xue, M. Christina Schneider, Sangdon Lim, Garron Gianopulos, Alexandra Perez


Abstract
Range performance level descriptors (RPLDs) connect assessment items to claims about what students at different achievement levels know and can do. Retrospectively assigning RPLDs to a large item bank is valuable but labor intensive. This work-in-progress study evaluates whether small and large language models can assist expert item-to-RPLD matching for 524 Grade 3–5 English language arts items. We compared direct classification with structured prompts that guide a model to analyze the knowledge, skills, evidence, and cognitive demand required by an item. We also examined self-consistency voting and a rater-informed prompting. Preliminary results indicate that larger hosted models achieved the closest overall performance to humans, although some locally hosted models produced comparable results. Voting improved prediction reliability but did not consistently increase agreement with human scores, whereas training models with human scores generally improved classification accuracy. Model performance also tended to decline as grade level increased.
Anthology ID:
2026.aimecon-wip.23
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
177–183
Language:
URL:
https://aclanthology.org/2026.aimecon-wip.23/
DOI:
Bibkey:
Cite (ACL):
Kang Xue, M. Christina Schneider, Sangdon Lim, Garron Gianopulos, and Alexandra Perez. 2026. Predicting Item-to-Range Performance Level Descriptor Matches with Structured Language Model Reasoning. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress, pages 177–183, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
Predicting Item-to-Range Performance Level Descriptor Matches with Structured Language Model Reasoning (Xue et al., AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-wip.23.pdf