Won-Chan Lee

Author directory

2026

This study applied generalizability theory to examine score variation and dependability in AP Chinese speaking tasks completed with and without ChatGPT support. Although ChatGPT-supported tasks were associated with higher scores, the NoGPT condition consistently exhibited higher dependability coefficients. Increasing the numbers of tasks and raters further improved score dependability.
Automated item generation with large language models (LLMs) is typically evaluated using aggregate accuracy metrics that conflate the quality of the source passage with noise in- troduced by the generation process itself. We address this gap by embedding a prospective i : (p×s×t) Generalizability Theory (G- theory) design into the evaluation of a QLoRA fine-tuned Qwen2.5-7B-Instruct model on the SciQ corpus. Passages (p) serve as the object of measurement; random seeds (s) and prompt templates (t) are fully crossed facets; items are nested within each (p,s,t) cell. Across 9,000 observations and five binary quality met- rics, we find that 77–81% of total variance is attributable to the passage, seed and template main effects are negligible (≤0.02%), and G- coefficients (E 𝜌2) uniformly exceed 0.97 un- der the observed design. D-study projections show that a single seed and template already achieves E 𝜌2 = 0.87, while the observed de- sign (ns = 3, nt = 3, ni = 2) reaches 0.98. Code, data, and R analysis scripts are released to support reproducible psychometric evalua- tion of future item-generation systems.