Automated Generation and Scoring of Maze Reading Comprehension Assessments

Hatice Kubra Karakis, Walter Leite, Logan Scott, Xinyi Tai, Akihito Kamata


Abstract
This work proposes automated generation and psychometric scoring of Maze comprehension assessments using large language models (LLMs) and a multilevel item response theory (multilevel IRT) framework. Findings show reliable ability estimates across passages and items, offering scalable, curriculum-aligned formative assessment that reduces teacher workload and supports targeted reading instruction.
Anthology ID:
2026.aimecon-sessions.35
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
320–328
Language:
URL:
https://aclanthology.org/2026.aimecon-sessions.35/
DOI:
Bibkey:
Cite (ACL):
Hatice Kubra Karakis, Walter Leite, Logan Scott, Xinyi Tai, and Akihito Kamata. 2026. Automated Generation and Scoring of Maze Reading Comprehension Assessments. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers, pages 320–328, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
Automated Generation and Scoring of Maze Reading Comprehension Assessments (Karakis et al., AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-sessions.35.pdf