LLM Difficulty Prediction Across Three L2 Task Types and Two Examination Forms

Meng Lyu, Moses Oluoke Omopekunola, Feng Wang


Abstract
Large language models (LLMs) have been proposed as a substitute for field trials in itemdifficulty estimation, but evidence about where they succeed is usually drawn from a single test form. We compare LLM difficulty predictions across three task types that make different language demands, grammar cloze, semantic cloze and reading comprehension, on two operational Grade 9 English examinations sat by the same cohort in China (1,199 and1,166 examinees). Three LLMs estimated difficulty for 60 MCQ items under the same standardized protocol. On Form 1, grammar cloze had the highest Pearson correlation for all three models; the Fisher-averaged correlations were .81 for grammar cloze and .34 for semantic cloze. On Form 2, reading had the highest correlation for all three models (Fisheraveraged .78), and the grammar–semantic difference was smaller and inconsistent in direction. With ten items per task type, a task-type advantage seen on one form needs testing on other forms before it is read as a property of the model.
Anthology ID:
2026.aimecon-wip.55
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
441–447
Language:
URL:
https://aclanthology.org/2026.aimecon-wip.55/
DOI:
Bibkey:
Cite (ACL):
Meng Lyu, Moses Oluoke Omopekunola, and Feng Wang. 2026. LLM Difficulty Prediction Across Three L2 Task Types and Two Examination Forms. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress, pages 441–447, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
LLM Difficulty Prediction Across Three L2 Task Types and Two Examination Forms (Lyu et al., AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-wip.55.pdf