2026

Large language models (LLMs) have been proposed as a substitute for field trials in itemdifficulty estimation, but evidence about where they succeed is usually drawn from a single test form. We compare LLM difficulty predictions across three task types that make different language demands, grammar cloze, semantic cloze and reading comprehension, on two operational Grade 9 English examinations sat by the same cohort in China (1,199 and1,166 examinees). Three LLMs estimated difficulty for 60 MCQ items under the same standardized protocol. On Form 1, grammar cloze had the highest Pearson correlation for all three models; the Fisher-averaged correlations were .81 for grammar cloze and .34 for semantic cloze. On Form 2, reading had the highest correlation for all three models (Fisheraveraged .78), and the grammar–semantic difference was smaller and inconsistent in direction. With ten items per task type, a task-type advantage seen on one form needs testing on other forms before it is read as a property of the model.