Matthew Gushta
Author directory2026
Using LLM Judges’ Paired Comparisons to Estimate Mathematics Item Difficulty
Lanrong Li | Mohammad A. A. Abulela | Matthew Gushta
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Lanrong Li | Mohammad A. A. Abulela | Matthew Gushta
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Estimating item difficulty without collecting field testing data has been a long sought-after goal in educational measurement. We prompted large language models (LLMs) to compare mathematics items in pairs to estimate their difficulty. Results showed strong correlations between difficulty based on one LLM judge’s paired comparisons and empirical item difficulty.