Meng Lyu
Author directory2026
LLM Difficulty Prediction Across Three L2 Task Types and Two Examination Forms
Meng Lyu | Moses Oluoke Omopekunola | Feng Wang
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Meng Lyu | Moses Oluoke Omopekunola | Feng Wang
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Large language models (LLMs) have been proposed as a substitute for field trials in itemdifficulty estimation, but evidence about where they succeed is usually drawn from a single test form. We compare LLM difficulty predictions across three task types that make different language demands, grammar cloze, semantic cloze and reading comprehension, on two operational Grade 9 English examinations sat by the same cohort in China (1,199 and1,166 examinees). Three LLMs estimated difficulty for 60 MCQ items under the same standardized protocol. On Form 1, grammar cloze had the highest Pearson correlation for all three models; the Fisher-averaged correlations were .81 for grammar cloze and .34 for semantic cloze. On Form 2, reading had the highest correlation for all three models (Fisheraveraged .78), and the grammar–semantic difference was smaller and inconsistent in direction. With ten items per task type, a task-type advantage seen on one form needs testing on other forms before it is read as a property of the model.