Responsible use of generative AI when creating reading comprehension questions: Inference matters

Zuowei Wang, Michael Flor, Jiyun Zu, Tenaha O’Reilly, Wanjing Anya Ma


Abstract
AI-generated and expert-created reading comprehension questions can show similar item statistics yet differ in the types of inferences required. This difference stemmed from AI’s failure to follow prompts during an intermediate item generation step. Evaluations of AI-generated items should document prompts and generation steps to identify and mitigate construct-relevant differences.
Anthology ID:
2026.aimecon-wip.27
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
207–210
Language:
URL:
https://aclanthology.org/2026.aimecon-wip.27/
DOI:
Bibkey:
Cite (ACL):
Zuowei Wang, Michael Flor, Jiyun Zu, Tenaha O’Reilly, and Wanjing Anya Ma. 2026. Responsible use of generative AI when creating reading comprehension questions: Inference matters. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress, pages 207–210, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
Responsible use of generative AI when creating reading comprehension questions: Inference matters (Wang et al., AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-wip.27.pdf