Using Generative AI to Simulate Item Responses by Skills Insight Score Band

Jianshen Chen, Chansoon Lee, Kylie Gorney, Judit Antal


Abstract
This study evaluates GPT-5.4 Thinking simulations of reading-item responses using Skills Insight score-band descriptions and item content. Simulated and empirical item statistics and IRT parameters were compared. Difficulty recovery was strongest, with moderately strong b-parameter correlations, while a and c recovery was weaker, supporting preliminary difficulty evaluation before field testing.
Anthology ID:
2026.aimecon-wip.18
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
140–144
Language:
URL:
https://aclanthology.org/2026.aimecon-wip.18/
DOI:
Bibkey:
Cite (ACL):
Jianshen Chen, Chansoon Lee, Kylie Gorney, and Judit Antal. 2026. Using Generative AI to Simulate Item Responses by Skills Insight Score Band. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress, pages 140–144, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
Using Generative AI to Simulate Item Responses by Skills Insight Score Band (Chen et al., AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-wip.18.pdf