Using LLM Judges’ Paired Comparisons to Estimate Mathematics Item Difficulty

Lanrong Li, Mohammad A. A. Abulela, Matthew Gushta


Abstract
Estimating item difficulty without collecting field testing data has been a long sought-after goal in educational measurement. We prompted large language models (LLMs) to compare mathematics items in pairs to estimate their difficulty. Results showed strong correlations between difficulty based on one LLM judge’s paired comparisons and empirical item difficulty.
Anthology ID:
2026.aimecon-wip.24
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
184–192
Language:
URL:
https://aclanthology.org/2026.aimecon-wip.24/
DOI:
Bibkey:
Cite (ACL):
Lanrong Li, Mohammad A. A. Abulela, and Matthew Gushta. 2026. Using LLM Judges’ Paired Comparisons to Estimate Mathematics Item Difficulty. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress, pages 184–192, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
Using LLM Judges’ Paired Comparisons to Estimate Mathematics Item Difficulty (Li et al., AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-wip.24.pdf