Won-Chan Lee
Author directory2026
Evaluating Score Dependability in ChatGPT-Supported AP Chinese Speaking Tasks
Dan Song | Won-Chan Lee
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Dan Song | Won-Chan Lee
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
This study applied generalizability theory to examine score variation and dependability in AP Chinese speaking tasks completed with and without ChatGPT support. Although ChatGPT-supported tasks were associated with higher scores, the NoGPT condition consistently exhibited higher dependability coefficients. Increasing the numbers of tasks and raters further improved score dependability.
Generalizability Theory for Evaluating Fine-Tuned LLMs in Automated Item Generation
Zhifei Li | Won-Chan Lee
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Zhifei Li | Won-Chan Lee
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Automated item generation with large language models (LLMs) is typically evaluated using aggregate accuracy metrics that conflate the quality of the source passage with noise in- troduced by the generation process itself. We address this gap by embedding a prospective i : (p×s×t) Generalizability Theory (G- theory) design into the evaluation of a QLoRA fine-tuned Qwen2.5-7B-Instruct model on the SciQ corpus. Passages (p) serve as the object of measurement; random seeds (s) and prompt templates (t) are fully crossed facets; items are nested within each (p,s,t) cell. Across 9,000 observations and five binary quality met- rics, we find that 77–81% of total variance is attributable to the passage, seed and template main effects are negligible (≤0.02%), and G- coefficients (E 𝜌2) uniformly exceed 0.97 un- der the observed design. D-study projections show that a single seed and template already achieves E 𝜌2 = 0.87, while the observed de- sign (ns = 3, nt = 3, ni = 2) reaches 0.98. Code, data, and R analysis scripts are released to support reproducible psychometric evalua- tion of future item-generation systems.