Wesley Morris

Author directory

2026

We used confirmatory factor analysis to assess the reliability and construct representation of an LLM-based measurement instrument of language proficiency. LLMs were at least as reliable as human raters and loaded onto the same underlying factor, though analyses indicated a less than perfect alignment between LLM and human raters.
Cloze exercises offer scalable comprehension assessment, but their validity depends on which words are selected as gaps. We compared three automated methods for cloze exercises generated from summaries within an intelligent textbook platform. Conditioning masked language model predictions on the source text (contextuality-plus) produced higher quality and more source-dependent gaps.