YoungKoung Kim

Author directory

2026

This study compares generative models Gemma and Qwen with ModernBERT and Mahalanobis-LOF for invalid-response detection in automated essay scoring. Qwen maintained high recall while reducing false invalid flags on a representative test set and detected some fluent off-topic responses in a matched diagnostic set. Fluent off-topic detection remained difficult.
We present results from reconstructing multiple-choice model (MCM) and three-parameter logistic (3PL) model curves using a fine-tuned multimodal large language model (LLM) based on Qwen3.5. The model is prompted and fine-tuned to replicate choice probabilities across a large training corpus of multiple-choice items containing both image and text stimuli, conditioned on a labeled set of student ability levels. By learning to reproduce the systematic error patterns of students across a discrete range of abilities, the LLM implicitly captures the underlying response probabilities encoded in the 3PL and MCM curves. This allows us to accurately approximate item difficulty on a held-out test set directly from the model’s predicted option probabilities.