Automatic Evaluation of Multiple-Choice Items for Reading Comprehension: Effects of Question and Distractor Categories

John S. Y. Lee, Yin Poon, Shunjie Wang, Kai Wah Chu


Abstract
Automatic generation of multiple-choice (MC) items for reading comprehension can support language learning by providing large amounts of practice materials. To enable rapid development of MC generation models, automatic assessment is essential since it is time-consuming to manually evaluate question and distractor quality. Although Text Informativity (TI) has been adopted as an automatic evaluation metric, the ability of Large Language Models (LLMs) to estimate the TI scores of different categories of questions and distractors has not yet been thoroughly analyzed. This paper investigates LLM performance in calculating TI scores for the range of questions and distractors defined in the PIRLS (Progress in International Reading Literacy Study) and STARC (Structured Annotations for Reading Comprehension) frameworks. We show that automatically estimated TI scores may result in systematic preferences for some question and distractor categories, and recommend that TI scores be used for within-category comparisons only.
Anthology ID:
2026.llms4ssh-1.18
Volume:
Proceedings of Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities (LLMs4SSH) @ LREC 2026
Month:
May
Year:
2026
Address:
Palma de Mallorca (Spain)
Editors:
Arturo Montejo-Raez, Cristina Grisot, Joanna Blochowiak, Nikola Ljubešić, Elena Battaner, German Rigau
Venues:
LLMs4SSH | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
170–174
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-llms4ssh-18
DOI:
10.63317/2htm3v3vdbuv
Bibkey:
Cite (ACL):
John S. Y. Lee, Yin Poon, Shunjie Wang, and Kai Wah Chu. 2026. Automatic Evaluation of Multiple-Choice Items for Reading Comprehension: Effects of Question and Distractor Categories. In Proceedings of Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities (LLMs4SSH) @ LREC 2026, pages 170–174, Palma de Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
Automatic Evaluation of Multiple-Choice Items for Reading Comprehension: Effects of Question and Distractor Categories (Lee et al., LLMs4SSH 2026)
Copy Citation: