Kylie Gorney

Author directory

2026

This study evaluates GPT-5.4 Thinking simulations of reading-item responses using Skills Insight score-band descriptions and item content. Simulated and empirical item statistics and IRT parameters were compared. Difficulty recovery was strongest, with moderately strong b-parameter correlations, while a and c recovery was weaker, supporting preliminary difficulty evaluation before field testing.