Generalizability Theory for Evaluating Fine-Tuned LLMs in Automated Item Generation

Zhifei Li, Won-Chan Lee


Abstract
Automated item generation with large language models (LLMs) is typically evaluated using aggregate accuracy metrics that conflate the quality of the source passage with noise in- troduced by the generation process itself. We address this gap by embedding a prospective i : (p×s×t) Generalizability Theory (G- theory) design into the evaluation of a QLoRA fine-tuned Qwen2.5-7B-Instruct model on the SciQ corpus. Passages (p) serve as the object of measurement; random seeds (s) and prompt templates (t) are fully crossed facets; items are nested within each (p,s,t) cell. Across 9,000 observations and five binary quality met- rics, we find that 77–81% of total variance is attributable to the passage, seed and template main effects are negligible (≤0.02%), and G- coefficients (E 𝜌2) uniformly exceed 0.97 un- der the observed design. D-study projections show that a single seed and template already achieves E 𝜌2 = 0.87, while the observed de- sign (ns = 3, nt = 3, ni = 2) reaches 0.98. Code, data, and R analysis scripts are released to support reproducible psychometric evalua- tion of future item-generation systems.
Anthology ID:
2026.aimecon-sessions.7
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
59–69
Language:
URL:
https://aclanthology.org/2026.aimecon-sessions.7/
DOI:
Bibkey:
Cite (ACL):
Zhifei Li and Won-Chan Lee. 2026. Generalizability Theory for Evaluating Fine-Tuned LLMs in Automated Item Generation. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers, pages 59–69, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
Generalizability Theory for Evaluating Fine-Tuned LLMs in Automated Item Generation (Li & Lee, AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-sessions.7.pdf